Skip to main content

CLI tool to compare model checkpoints — weight deltas, SVD structure, layer drift

Project description

model-diff

See what changed inside any model checkpoint. Weight deltas, SVD structure, spectral analysis, and diagnostic conclusions — in one command.

pip install model-diff
model-diff Qwen/Qwen2.5-7B Qwen/Qwen2.5-7B-Instruct -o report.html

What it does

Compares two model checkpoints (base vs instruct, v1 vs v2, merge A vs merge B) and produces:

  • Per-module metrics: Frobenius norm of weight delta, cosine similarity, sparsity
  • SVD analysis: effective rank, top-k singular value concentration, spectral decay
  • Layer heatmaps: 4 metrics across all layers and module types
  • Diagnostic conclusions: human-readable "diagnosis" — was this surgical SFT or heavy rewriting?

Output formats

Format Flag Description
Terminal (default) Quick summary table with diagnostics
HTML -o report.html Single-file report with embedded plots, heatmaps, SVD spectra, and diagnostic summary
JSON -o report.json Machine-readable, includes diagnostics.profile_tag and diagnostics.summary

Example output (terminal)

model-diff: Qwen/Qwen2.5-7B → Qwen/Qwen2.5-7B-Instruct
Tensors: 283 analyzed, 85 skipped
Total ||ΔW||: 25.44  |  Mean cos_sim: 0.99996  |  Mean eff_rank: 2166

Module                                                    ΔW/W   cos_sim  eff_rank    conc  spars
───────────────────────────────────────────────────────────────────────────────────────────────────
lm_head.weight                                          0.0395  0.99990      1445   0.540  0.002
model.layers.0.self_attn.v_proj.weight                  0.0279  0.99998      1839   0.178  0.002
...

─── Diagnosis ───

Qwen/Qwen2.5-7B-Instruct was surgically fine-tuned with minimal weight changes.
Mean relative change: 0.0120, cosine similarity: 0.99996, total ||ΔW||: 25.44.

  • Surgical fine-tuning
    Mean relative change is very small (0.0120). Typical of careful SFT.

  ▸ Output head is the most changed module
    lm_head change (0.0395) is 3.3× the body average.

  • LayerNorm weights nearly untouched
    LayerNorm mean change is 0.000088 — essentially frozen.

Diagnostic profiles

The diagnostic engine classifies fine-tuning into four profiles based on 7 calibration pairs:

Profile Mean ΔW/W Example
SURGICAL < 0.015 Qwen2.5 family — minimal, targeted changes
STANDARD 0.015–0.05 Llama-3.1-8B, Mistral-7B — typical SFT
HEAVY 0.05–0.12 Llama-3.2-3B — aggressive training, LayerNorm modified
EXTREME > 0.12 Gemma-2-9B — full-rank rewriting, possible continued pre-training

Requirements

  • Python >= 3.9
  • CPU only — no GPU needed
  • ~6.6 GB peak RAM for 7B models (memory-optimized SVD)
  • ~18 minutes per 7B pair including download

CLI options

model-diff MODEL_A MODEL_B [OPTIONS]

  MODEL_A, MODEL_B    HuggingFace model IDs or local paths

Options:
  -o, --output PATH   Output file (.json or .html). Omit for terminal text.
  --top-k INT         Number of top singular values to track [default: 20]
  --top-n INT         Number of modules to show in text output [default: 20]
  --token TEXT        HuggingFace token (or set HF_TOKEN env var)

For AI agents

model-diff produces structured JSON output suitable for programmatic use:

{
  "model_a": "Qwen/Qwen2.5-7B",
  "model_b": "Qwen/Qwen2.5-7B-Instruct",
  "n_tensors": 283,
  "diagnostics": {
    "profile_tag": "surgical",
    "summary": "Qwen2.5-7B-Instruct was surgically fine-tuned...",
    "findings": [
      {
        "category": "magnitude",
        "severity": "info",
        "title": "Surgical fine-tuning",
        "detail": "Mean relative change is very small (0.0120)..."
      }
    ]
  },
  "modules": [...]
}

Use cases for agents:

  • "How was model X fine-tuned?" → run model-diff, read diagnostics.summary
  • "Which fine-tune should I pick?" → compare profile_tags across variants
  • "Did this merge break anything?" → check for unusual patterns (EXTREME profile, LayerNorm modified)

How it works

  1. Downloads safetensors files (not full model) via huggingface_hub
  2. Streams tensor-by-tensor: load → compute delta → SVD → free → next
  3. Never loads both full models simultaneously
  4. Memory-optimized SVD: in-place delta computation, free inputs before SVD phase
  5. Randomized SVD via QR projection for matrices > 8192

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

modeldelta-0.1.0.tar.gz (25.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

modeldelta-0.1.0-py3-none-any.whl (28.3 kB view details)

Uploaded Python 3

File details

Details for the file modeldelta-0.1.0.tar.gz.

File metadata

  • Download URL: modeldelta-0.1.0.tar.gz
  • Upload date:
  • Size: 25.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.3

File hashes

Hashes for modeldelta-0.1.0.tar.gz
Algorithm Hash digest
SHA256 7e93f80af34ff36eb82da2f144581cac74ac3801e324581d6bbe0ee06276e8e2
MD5 c1d539d77428b3137cf15baa66c1f8e4
BLAKE2b-256 24a86ff8dfa6efb172ff4e077c670c6f5ffa0de62abc11a73322dee1901e8386

See more details on using hashes here.

File details

Details for the file modeldelta-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: modeldelta-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 28.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.3

File hashes

Hashes for modeldelta-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 811ec52be26312901a24c3e768bc32980c5f5d30531d44c768b9955cf3df00ac
MD5 7fc8bbfe0b6172dab84ab9810a7bf17e
BLAKE2b-256 45b20bb81a783856b7f679c974234532a51e5b1032ef22bd2003d084b6e0f300

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page