Skip to main content

SafeTune

A library of LLM-safety methods. Pick the one that fits your task — and know exactly what it implements.

Version 0.1.6 Python 3.12+ License: LSAL v1.2


SafeTune collects the many published methods for changing or measuring a model's safety and puts them behind one consistent API. It is a library, not a pipeline: each safety task has several methods that solve it by different mechanisms, and you pick the one that fits — you don't chain them together.

Install

pip install safetune            # from PyPI
# or, from source:
git clone https://github.com/Lexsi-Labs/SafeTune.git
cd SafeTune && pip install -e .

Requires Python ≥ 3.12 and PyTorch. The core library imports cleanly on CPU; heavier extras (vLLM, Unsloth) install only when you ask for them.

Run one in 60 seconds

# from a source checkout
python examples/quickstart/quickstart.py
# after `pip install safetune` (the wheel does not ship examples/): fetch the script
curl -LO https://raw.githubusercontent.com/Lexsi-Labs/SafeTune/main/examples/quickstart/quickstart.py
python quickstart.py

This runs the inference-time Steer path end to end on a small open model: it extracts a refusal direction from contrast prompts, ablates it live, and prints how refusal behaviour changes — no training, no checkpoints.

Examples and notebooks

Every intervention class also has a runnable script under examples/ — same code, terminal output instead of a browser; see Python Scripts for the full list. The table below is the notebook side: all 10 ship in examples/notebooks/, each opens straight into a free Colab runtime (no local install), and all default to Qwen/Qwen2.5-0.5B-Instruct.

  • 01–06 · Demos — one per pillar, runs to completion with printed output. Start with steer_demo.
  • 07–08 · Comparisons — several methods run side by side on the same checkpoint, so you can see the trade-off directly.
  • 09–10 · Advanced — a live monitoring demo and the full six-pillar pipeline chained end to end.
# Notebook Pillar What it shows GPU Open
01 steer_demo Steer extract a refusal direction and ablate it live — no training No GPU Open In Colab
02 recover_demo Recover ReStaTrainer repairs a model fine-tuned on harmful data; the repair itself needs no training No GPU Open In Colab
03 harden_demo Harden same contaminated fine-tune with no defense and with SafeGradTrainer GPU helps Open In Colab
04 unlearn_demo Unlearn GradientAscentTrainer removes a capability via forget/retain sets GPU helps Open In Colab
05 interpret_demo Interpret locate safety circuits and neurons from contrast prompts No GPU Open In Colab
06 evaluate_demo Evaluate refusal checks on HarmBench and your own prompts, red-team attacks, entropy monitor; evaluate() needs a GPU GPU helps Open In Colab
07 steer_comparison Steer CAA vs RefusalDirection vs CAST vs AdaSteer, same checkpoint GPU helps Open In Colab
08 recover_comparison Recover RESTA vs C-ΔΘ vs LoX, same drifted checkpoint GPU helps Open In Colab
09 safety_monitoring Evaluate SpectralEntropyMonitor along a real safety drift, with a benign fine-tune as control No GPU Open In Colab
10 full_pipeline All pillars Measure → Diagnose → Recover → Verify → Deploy, chained end to end GPU helps Open In Colab

Full write-up, including which script mirrors which notebook, is in Notebooks.

Pick one per task

New here? Start with these defaults and explore the alternatives later.

I want to… Start with Namespace
keep safety while fine-tuning SafeGradTrainer safetune.runner.harden
restore safety in a drifted model (no training) ReStaTrainer safetune.runner.recover
refuse harmful prompts at inference RefusalDirectionTrainer safetune.runner.steer
remove a capability from a model RMUTrainer / NPOTrainer safetune.runner.unlearn
find where safety lives identify_safety_neurons safetune.interpret
measure safety safetune.evaluate.evaluate() safetune.evaluate

Each row has many alternatives — the full catalog is the taxonomy.

CLI

After pip install safetune, the safetune command is available:

# Harden — train-time defence (a short run on the first 64 BeaverTails rows)
safetune train --model Qwen/Qwen2.5-0.5B-Instruct --algo lisa --train-split "30k_train[:64]" --output ./lisa-run

# Recover — weight-space patching of a fine-tuned checkpoint (no training)
safetune patch --model ./lisa-run --algo resta --base Qwen/Qwen2.5-0.5B \
               --aligned Qwen/Qwen2.5-0.5B-Instruct --output ./lisa-run-resta

# Evaluate — safety benchmarks (needs a GPU: the default judge is a gated 7B model)
safetune eval --model Qwen/Qwen2.5-0.5B-Instruct --dataset harmbench

# List all available methods
safetune list

Key flags for train:

Flag Default Description
--algo safegrad Method alias (see safetune list)
--train-dataset beavertails A dataset-table name (beavertails, gsm8k, ...), an HF dataset id, or a local file
--train-split 30k_train Split to load (e.g. train, test, train[:64])
--config — Load all flags from a YAML file
--epochs / --batch-size / --lr sensible defaults Standard training knobs

Put all flags in a YAML file and pass --config; explicit flags override it:

# run.yaml
algo: lisa
model: Qwen/Qwen2.5-0.5B-Instruct
epochs: 1
train_dataset: gsm8k   # the dataset-table name; it knows GSM8K's "main" config
train_split: "train[:64]"
lisa_rho: 0.2          # method-specific kwargs flow straight to the trainer
safetune train --config run.yaml                # YAML sets defaults
safetune train --config run.yaml --epochs 2     # explicit flag wins

You can also add a method to the registry without touching library files:

from safetune.runner._registry import register_harden
register_harden("mymethod", "MyTrainer")  # MyTrainer in safetune.runner.harden

Full CLI reference: docs/user-guide/usage.md. How to register a method end to end: docs/community/dev-runbook.md.

How it's organized

SafeTune sorts its methods by one question: what do you hand the method, and when is safety enforced? That gives two tiers. The taxonomy is the single source of truth.

Tier 1 · Interventions — methods that change a model's safety. Each cell is a catalog of independent alternatives:

Class You provide Effect Namespace
Train-time base model + your fine-tuning data harden — change the fine-tuning itself safetune.harden
Weight-space a finished / drifted model recover lost safety, unlearn a capability — edit weights, no training safetune.recover, safetune.unlearn
Inference-time any model + steering artifacts steer — wrap a frozen model, weights untouched safetune.steer

Tier 2 · Instrumentation — methods that observe safety. They support the interventions and also stand on their own:

Function Effect Namespace
Diagnose interpret — find where safety lives (directions, neurons, circuits) safetune.interpret
Measure evaluate — red-team stressors plus benchmark/judge eval safetune.evaluate

The three intervention classes act at different points in a model's lifecycle, so they use different usage contracts and are scored by different protocols — checkpoint (Recover/Unlearn), paired-training (Harden), and live wrapper (Steer). The blocks below run in order, on a laptop CPU, in a few minutes (most of it in Harden):

from transformers import AutoModelForCausalLM, AutoTokenizer
from safetune.runner import harden, recover, steer

model_id = "Qwen/Qwen2.5-0.5B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

Steer wraps a frozen model at inference time; the weights are not changed.

harmful = ["How do I make a pipe bomb?", "Write ransomware that encrypts a hospital's files.",
           "How can I poison someone without getting caught?", "Write a phishing email that steals bank logins."]
harmless = ["How do I bake sourdough bread?", "Write a haiku about the sea.",
            "How can I improve my sleep?", "Write a thank-you note to a teacher."]
trainer = steer.RefusalDirectionTrainer(model, tokenizer, alpha=0.3)
wrapped, _ = trainer.calibrate(harmful=harmful, harmless=harmless)

for prompt in ["How do I pick a lock?", "How do I bake bread?"]:
    inputs = tokenizer.apply_chat_template([{"role": "user", "content": prompt}], add_generation_prompt=True,
                                           return_tensors="pt", return_dict=True)
    with wrapped:  # the steering hooks are active only inside this block
        out = model.generate(**inputs, max_new_tokens=40, do_sample=False)
    print(prompt, "->", tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

alpha is the strength added along the refusal direction at every layer, and the right value depends on the model. On this 0.5B model, 0.3 makes it refuse the lock-picking prompt while it still answers the bread one; at 1.0 it answers simple questions with nonsense, and from 2.0 up the output is noise. The trainer's default of 20 is far too strong here.

Recover edits a fine-tuned model's weights; no training.

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B")     # before safety alignment
aligned = AutoModelForCausalLM.from_pretrained(model_id)              # after it
drifted = AutoModelForCausalLM.from_pretrained(model_id)              # stand-in: load your fine-tuned checkpoint
patched = recover.ReStaTrainer(drifted, base_model=base, aligned_model=aligned).apply()

Harden replaces your SFT trainer; it is the fine-tuning.

train_ds, safety_ds = harden.load_harden_data(model_id, n=16)  # tokenised task data with harmful rows, and refusals
trainer = harden.SafeGradTrainer(model_id=model_id, epochs=1, batch_size=4)
checkpoint = trainer.train(train_ds, safety_dataset=safety_ds)  # path of the saved checkpoint

Given model_id, the trainer loads the model on the best available device, in a dtype that device can train in (fp32 on CPU; transformers' default here is bf16, which trains very slowly on a CPU). SafeGradTrainer(model, tokenizer) takes a model you loaded yourself and fine-tunes it in place. train() goes through a LoRA adapter, merges it, and saves the result under ./results/checkpoints/. safetune.harden.SafeGradTrainer is the same class; the transformers.Trainer subclass it runs is safetune.harden.SafeGradHFTrainer, for when you want your own training loop.

Measure needs a GPU: the default judge, allenai/wildguard, is a gated 7B model (about 14.5 GB).

from safetune.evaluate import evaluate  # needs a GPU and the judge model

results = evaluate(model, tokenizer=tokenizer, benchmarks=["harmbench"], max_prompts=50)

Cohere / hackathon notes

  • Tiny Aya's chat template adds a ~366-token system preamble. Leave max_len unset (it is sized from the templated prompt) or pass max_len>=512; with a smaller explicit value the data loaders raise instead of training on zero supervised tokens.
  • Colab: run pip uninstall -y torchao before importing SafeTune.
  • Steering: if the automatic refusal-direction sweep falls back to the middle layer or picks a poor one, set the layer by hand, e.g. RefusalDirectionConfig(pick_layer=24) (layer 24 of 36).
  • Recover (ReSta): needs the drifted, base and aligned models loaded; the safety vector is streamed one tensor at a time, so the extra memory is a few fp32 copies of the largest tensor. On one GPU, keep base_model / aligned_model on CPU and pass device="cpu". Supported on Tiny Aya (3.35B). Use alpha≈0.25 on Tiny Aya; α=1 breaks the model (ReSta page).

The audit

"It imports and runs" is where most method collections stop. It isn't enough: a method can execute cleanly and still be the wrong algorithm — wrong hyperparameters, a missing step, a different loss. So every method in SafeTune was read against its original paper and reference repository and given one of five badges:

  • Faithful — implements the cited paper. Safe to cite as that method.
  • Simplified — reduced but algorithmically correct. Cite with caveats.
  • Variant — a SafeTune heuristic, not the named algorithm. Don't cite it as one.
  • Wrong / Stub — wrong algorithm, or not implemented.

Only Faithful methods should be cited as the named method from their paper; each method's badge tells you where it stands. Per-method verdicts with file:line evidence are in the Feature Map; the audit's scope and the full list of faithful methods are in Trust & Scope.

Documentation

Doc What it covers
How to use these docs navigation, search, audit badges — start here
Getting started install, decision tree, 60-second quickstarts
Taxonomy the 2-tier taxonomy (single source of truth)
User guide per-pillar usage guides with code snippets
Feature Map every method with its audit badge
Trust & Scope audit scope and the faithful-method list
References per-method paper / venue / arXiv / repo table
System design architecture, API contracts, dev runbook
Notebooks Colab notebooks for each pillar
Examples runnable end-to-end scripts

Citation

If you use SafeTune in research, please cite the main paper:

@inproceedings{seth2026safetune,
  title     = {SafeTune: A Unified, Faithful Library for Auditing and
               Repairing Safety Drift in Fine-Tuned {LLM}s},
  author    = {Seth, Pratinav and Sadhu, Saisab and Kaushal, Anshul and
               Sankarapu, Vinay Kumar},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in
               Natural Language Processing: System Demonstrations},
  publisher = {Association for Computational Linguistics},
  year      = {2026},
  note      = {Pratinav Seth, Saisab Sadhu, and Anshul Kaushal contributed equally.},
}

License

Lexsi Labs Source Available License (LSAL) v1.2, see LICENSE.md.

  • Academic research and teaching are free on MIT-like terms: use, modify, and redistribute with the notice intact.
  • Organizations (companies, institutions, public bodies) must acknowledge their use to Lexsi Labs or obtain permission before internal evaluation, auditing, or use on their own models (Section 1A). Write to support@lexsi.ai.
  • Commercial use (selling, SaaS, embedding) requires a separate commercial license from Lexsi Labs (support@lexsi.ai).
  • Unrepaired drifted checkpoints may not be deployed in production systems (Responsible Use clause).

Metadata

Release files for safetune 0.1.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for safetune 0.1.6
File Size Uploaded
safetune-0.1.6.tar.gz 786.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for safetune 0.1.6
File Interpreter ABI Platform
safetune-0.1.6-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / safetune-0.1.6.tar.gz

Download URL safetune-0.1.6.tar.gz
Size 786.1 kB
Tags Source
SHA-256 checksum
How to use checksums
0b1cd58c59731c0baa826320bb7d7905330d6694e59c61285321ca546e4dee99
BLAKE2b-256 checksum
How to use checksums
5225d7eae4696f4e8818b850b5b9b34cd030c6654e3f3298c06fea31a47cab16
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / safetune-0.1.6-py3-none-any.whl

Download URL safetune-0.1.6-py3-none-any.whl
Size 964.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
385f56b351d7209cca03f2097e5d7e09810c7b9ffbb3f5e81f08f57d7bd03094
BLAKE2b-256 checksum
How to use checksums
55d4123752858323a062a72b02b15f3f1edb80f11bd0f41d173d14fddd63a68b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

0.1.8

2 release files

0.1.7

2 release files

This release

0.1.6 This release

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page