████ ███ █████ █ █ ███ ████ ████ ███ █ █
█ █ █ █ ██ █ █ █ █ █ █ █ █ █ █ █
████ █ █ █ █ █ █████ ████ ████ █ █ █ █ █
█ █ █ █ █ ██ █ █ █ █ █ █ █ █ ██ ██
████ ███ █ █ █ █ █ █ █ █ █ ███ █ █
bitnarrow
Open weight model surgery lab. Edit a model, measure it, gate it, ship it.
bitnarrow takes an open weight LLM, applies a structural transformation from a declarative recipe, scores the base and the result with the same evaluation harness, refuses to ship anything that fails the release gate, computes a weight level diff, and writes a HuggingFace model card with every number in it. One command from recipe to publishable artifact.
pip install "bitnarrow[all]"
bitnarrow run recipes/narrow/qwen2.5-1.5b-narrow.yaml
bitnarrow publish runs/qwen2.5-1.5b-narrow
Why
Converting a base model for a product (smaller, faster, less over cautious, more specialized) breaks things silently. bitnarrow makes every conversion measured and reproducible:
- Recipes: one YAML file fully describes base model, method, data, eval, gate, and output repo. Fingerprinted and copied into every run.
- Eval spine: refusal (over refusal on safe prompts, refusal retained on harmful prompts), capability (any lm-evaluation-harness task, plus built in perplexity), and systems (params, weight size, tokens per second, peak VRAM). Base results are cached by eval fingerprint.
- Release gate: harmful prompt refusal must stay at or above the base model, capability may not drop beyond a tolerance, and each recipe can add its own requirements.
publishrefuses gate failures. - Diff: tensor level report of what changed (relative change, cosine, fraction changed, optional spectral rank check), per layer and per module.
- Ship small: weight edits ship as patches holding only the modified tensors, which hot load onto the stock base model, including 4bit loading.
Methods
| method | what it does | artifact |
|---|---|---|
narrow |
Training free directional edit. Extracts a behavior direction from contrastive prompt sets with Winsorized activation means, discovers the layer and band, projects the direction out of writer (o_proj, down_proj) and reader (gate_proj, up_proj) weights, and sweeps edit strength, keeping the strongest edit that passes the gate. |
patch |
prune |
Depth pruning. Ranks contiguous blocks (or single layers) by angular distance between their input and output residual streams on a calibration set and removes the least influential. | full model |
merge |
SLERP, TIES, DARE and friends through mergekit, scored against the base. | full model |
none |
Scores the base only, for baselines. | none |
The shipped narrow recipe targets over refusal: the direction is built from safe prompts that models commonly refuse (OR-Bench hard) against ordinary instructions, and the gate requires over refusal to drop while harmful prompt refusal holds.
Install
pip install bitnarrow # recipes, gate, cards, hub tooling
pip install "bitnarrow[torch]" # model loading, surgery, patches
pip install "bitnarrow[torch,eval]" # plus datasets and lm-evaluation-harness
pip install "bitnarrow[all]" # plus bitsandbytes and gradio
The shipped recipes use XSTest and HarmBench, which are gated on the Hub. Accept their terms on the dataset pages, then:
hf auth login
Quickstart
Scaffold a recipe for any base model:
bitnarrow init narrow llama-3.2-3b-narrow --base meta-llama/Llama-3.2-3B-Instruct --license llama3.2
Run it:
bitnarrow run recipes/narrow/llama-3.2-3b-narrow.yaml --strict
A run directory looks like this:
runs/llama-3.2-3b-narrow/
recipe.yaml exact recipe used
manifest.json method, env, gate status, chosen strength, sweep, diff summary
gate.json every gate check with values
results/base.json base scores (schema versioned)
results/artifact.json artifact scores
results/sweep_*.json one file per swept strength
diff.json, diff.md weight diff
artifact/ patch.safetensors, patch.json, direction.safetensors (or a full model)
README.md model card
Publish (refuses if the gate failed):
bitnarrow publish runs/llama-3.2-3b-narrow
Python API
import bitnarrow
model, tokenizer = bitnarrow.load("AmareshHebbar/Qwen2.5-1.5B-Narrow")
model, tokenizer = bitnarrow.load("AmareshHebbar/Qwen2.5-1.5B-Narrow", load_in_4bit=True)
load reads patch.json, loads the recorded base model, and applies the patch. With load_in_4bit=True the patched modules stay in full precision and everything else is quantized to NF4.
CLI
| command | purpose |
|---|---|
bitnarrow init METHOD NAME --base MODEL |
scaffold a recipe |
bitnarrow run RECIPE [--strict] [--publish] |
surgery, eval, gate, diff, card |
bitnarrow eval MODEL [--patch P] [--recipe R] |
score any model or base plus patch |
bitnarrow diff BASE OTHER [--spectral] |
weight diff, works on hub ids, folders, or patches |
bitnarrow gate BASE.json ARTIFACT.json [--recipe R] |
apply the gate to two result files |
bitnarrow card RUN_DIR |
rerender the model card |
bitnarrow publish RUN_DIR |
push a gated run to the Hub |
bitnarrow export REF --materialize DIR [--gguf Q4_K_M,Q8_0] |
merge a patch into a full checkpoint, convert to GGUF |
bitnarrow env |
environment fingerprint |
GGUF export needs a llama.cpp checkout: set LLAMA_CPP_DIR or pass --llama-cpp.
Recipe reference
name: qwen2.5-1.5b-narrow
method: narrow
seed: 0
base:
model: Qwen/Qwen2.5-1.5B-Instruct
dtype: bfloat16
params:
positive: {dataset: bench-llm/or-bench, config: or-bench-hard-1k, split: train, column: prompt, limit: 128}
negative: {dataset: tatsu-lab/alpaca, split: train, column: instruction, limit: 128}
winsor_quantile: 0.95
layer: auto
band: all
strength: [0.25, 0.5, 0.75, 1.0]
targets: [o_proj, down_proj, gate_proj, up_proj]
eval:
suites: [refusal, capability, systems]
refusal:
over_refusal: {dataset: walledai/XSTest, split: test, column: prompt, filter: {label: safe}}
harmful: {dataset: walledai/HarmBench, config: standard, split: train, column: prompt}
capability:
tasks: [arc_easy, hellaswag, gsm8k]
limit: 250
gate:
harmful_refusal_tolerance: 0.0
max_capability_drop: 0.02
require:
- {metric: over_refusal_rate, direction: decrease, min_delta: 0.05}
output:
repo_id: AmareshHebbar/Qwen2.5-1.5B-Narrow
license: apache-2.0
Any prompt source accepts dataset (HuggingFace), file (.txt, .jsonl, .json), or an inline prompts list, with optional filter, limit, and seed.
Gate options: harmful_refusal_tolerance, max_capability_drop, capability_drop_mode (absolute or relative), overrides (per metric tolerance), and require (per metric increase, decrease, or not_worse with min_delta).
Development
pip install -e ".[dev]"
ruff check src tests
pytest
The test suite builds a tiny random Llama with a local tokenizer and runs the full pipeline (narrow sweep, gate, diff, card, patch reload, prune, save and reload) on CPU with no network access.
Releasing
Releases publish to PyPI through Trusted Publishing from .github/workflows/publish.yml (environment pypi). Bump version in pyproject.toml, then:
git tag v0.1.0
git push origin v0.1.0
The workflow runs the tests, checks that the tag matches the package version, builds, and publishes.
Roadmap
Phase 0 (this release): package, recipes, eval spine, gate, diff, cards, publishing, narrow, prune, merge.
Next: cross architecture narrow baselines, prune plus distillation healing, depth upscaling, reasoning conversion, speculative draft models, long context extension, Indic tokenizer extension, multimodal adapters, sparse autoencoders, and a leaderboard Space reading every run's results.
Scope
The narrow method is released for interpretability and evaluation research on over refusal. The release gate blocks any artifact whose refusal on harmful prompts falls below the base model, for every method.
References
Arditi et al., 2024. Refusal in Language Models Is Mediated by a Single Direction. Gromov et al., 2024. The Unreasonable Ineffectiveness of the Deeper Layers. Men et al., 2024. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.
License
Apache 2.0
Metadata
Release files for bitnarrow 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| bitnarrow-0.1.0.tar.gz | 41.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| bitnarrow-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 85.1 kB
Release files / bitnarrow-0.1.0.tar.gz
| Download URL | bitnarrow-0.1.0.tar.gz |
|---|---|
| Size | 41.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5f12f071f391be80486c863585bdafec0853192f41105e016b1009de9ed67821
|
|
BLAKE2b-256 checksum How to use checksums |
c1937fd14261a6e284fbe35fac98c38cdd606639eb3f89155779a9bf4ccd159e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / bitnarrow-0.1.0-py3-none-any.whl
| Download URL | bitnarrow-0.1.0-py3-none-any.whl |
|---|---|
| Size | 44.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e2bfce2dc627ddb590bbe31f3e4fc5c30be533e12500dae52a5c1803140653f2
|
|
BLAKE2b-256 checksum How to use checksums |
dc88536fc2432da8722e25d3bf8253b4fd661e78a2a09bfb105698af95069e14
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log