Framework for comparing language model configurations
Project description
lmdiff
Measures how and where two LLM configurations differ — not just whether one scores higher.
Compare language model configurations — not just weights, but weights + context + decoding + adapter + agent — via behavioral distance and multi-level diagnostics.
Why lmdiff?
lm-eval-harness tells you "model A scores 3 points higher than model B on MMLU." That's a scalar.
lmdiff tells you where those 3 points came from: which capabilities shifted, how far the output distribution moved, and whether two different modifications (e.g. a fine-tune vs. a context change) push behavior in the same direction or in opposite directions.
A Configuration is model + context + decoding + adapter + agent scaffold, not just model weights. Same checkpoint with a different system prompt is a different config — and lmdiff can quantify the difference.
Install
pip install lmdiff-kit
The import name is lmdiff; the PyPI distribution is lmdiff-kit (name disambiguation on PyPI).
Development install
mamba create -n lmdiff python=3.12
pip3 install torch torchvision --index-url https://download.pytorch.org/whl/cu130
pip install -e .
cu130 is for RTX 5090 / Blackwell. Pick the CUDA version that matches your GPU.
Command line
# Metric-level comparison (BD, token entropy, token KL)
lmdiff compare gpt2 distilgpt2 --probes v01
# Same, but JSON output to file
lmdiff compare gpt2 distilgpt2 --probes v01 --json --output result.json
# Per-domain capability radar (accuracy + BD per domain)
lmdiff radar gpt2 distilgpt2 --probes v01
# Single-model task evaluation
lmdiff run-task gpt2 --probes v01 --evaluator contains_answer
# List available metrics
lmdiff list-metrics
Python API
from lmdiff import Config, ModelDiff, ProbeSet
from lmdiff.report.terminal import print_report, print_radar
probes = ProbeSet.from_json("lmdiff/probes/v01.json")
md = ModelDiff(
Config(model="gpt2"),
Config(model="distilgpt2"),
probes,
)
# Metric-level comparison
report = md.run(level="output", max_new_tokens=16)
print_report(report)
# Per-domain capability radar
radar_result = md.run_radar(probes=probes, max_new_tokens=16)
print_radar(radar_result)
Example: what lmdiff finds
Llama-2-7B vs YaRN-Llama-2-7b-128k on short prompts:
- BD = 1.03 nats — significant distributional shift even on prompts well within the original 4k context.
- TokenEntropy delta ≈ 0 — the distributions shifted direction, not spread; YaRN didn't make the model more or less uncertain on average.
- TokenKL = 0.35 — single-step distributions are similar, but multi-step generation diverges much further.
- Stopping behavior changed — YaRN learned to emit EOS after short answers; base Llama-2 keeps generating. Invisible to perplexity benchmarks; obvious in BD on generated continuations.
The point: same parameter count, similar single-step KL, but the generation behavior is meaningfully different — and the kind of difference is what lmdiff surfaces.
What gets measured
Three output-level metrics:
- BehavioralDistance — symmetric, self-entropy-baseline-subtracted cross-entropy distance. BPB-normalized when tokenizers differ.
- TokenEntropy — mean per-token next-token entropy delta, A vs B.
- TokenKL — symmetric KL divergence over full vocab.
CapabilityRadar adds per-domain accuracy + BD breakdown across math/knowledge/code (or any multi-domain probe set).
All return structured results with per-probe breakdowns in .details.
Configuration abstraction
A Config is more than a model name:
Config(
model="gpt2",
system_prompt="You are concise.",
context=[{"role": "user", "content": "..."}],
decode={"strategy": "sample", "temperature": 0.7},
name="gpt2-concise",
)
Same weights + different context/decoding = different config = measurable behavioral difference.
JSON output
All results serialize to deterministic JSON with schema_version for forward compatibility:
from lmdiff.report.json_report import to_json, write_json
write_json(report, "output.json")
Status
Phase 1 shipped — published to PyPI as lmdiff-kit v0.1.0. Working: BehavioralDistance, TokenEntropy, TokenKL, CapabilityRadar, CLI, JSON reports, Python API.
Phase 2 in progress: Change Geometry — treating behavioral changes as vectors with direction/magnitude/cosine similarity across multiple variants.
Not yet: representation/trajectory/causal metrics, HTML/LaTeX reports, viz. See CLAUDE.md for the full roadmap.
Development
pytest # fast tests (mocks only)
pytest -m slow -o "addopts=" # includes gpt2/distilgpt2 E2E
Architecture rules, implementation order, and coding conventions live in CLAUDE.md.
License
MIT — see LICENSE.
Citation
Paper forthcoming.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file lmdiff_kit-0.1.1.tar.gz.
File metadata
- Download URL: lmdiff_kit-0.1.1.tar.gz
- Upload date:
- Size: 48.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8ceaf3775b7fd225b0a5885f8f34e4c65af1478a6db163bed17ada2dc2ce41d1
|
|
| MD5 |
16a92764d34465a6ed00219ea5926bf4
|
|
| BLAKE2b-256 |
8fc752163de152bec0c575a89853c5152d38f0af323921d329ab6bcb9d2b74d0
|
Provenance
The following attestation bundles were made for lmdiff_kit-0.1.1.tar.gz:
Publisher:
publish.yml on MaiqiVerse/lmdiff
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
lmdiff_kit-0.1.1.tar.gz -
Subject digest:
8ceaf3775b7fd225b0a5885f8f34e4c65af1478a6db163bed17ada2dc2ce41d1 - Sigstore transparency entry: 1341624010
- Sigstore integration time:
-
Permalink:
MaiqiVerse/lmdiff@5b2559daa01dce239e66c6dfd2573071ad2e3e98 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/MaiqiVerse
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5b2559daa01dce239e66c6dfd2573071ad2e3e98 -
Trigger Event:
release
-
Statement type:
File details
Details for the file lmdiff_kit-0.1.1-py3-none-any.whl.
File metadata
- Download URL: lmdiff_kit-0.1.1-py3-none-any.whl
- Upload date:
- Size: 33.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8f49e99e1fb0f178d7a57a3df3b289d17ede982b4e690ad4e75456dcc3fec451
|
|
| MD5 |
66d38606f70be41b57a2663f77df8349
|
|
| BLAKE2b-256 |
ed5dc0bfa780a605413a5634cf6343c61eb06eeb7831a395dada73d42ab96723
|
Provenance
The following attestation bundles were made for lmdiff_kit-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on MaiqiVerse/lmdiff
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
lmdiff_kit-0.1.1-py3-none-any.whl -
Subject digest:
8f49e99e1fb0f178d7a57a3df3b289d17ede982b4e690ad4e75456dcc3fec451 - Sigstore transparency entry: 1341624011
- Sigstore integration time:
-
Permalink:
MaiqiVerse/lmdiff@5b2559daa01dce239e66c6dfd2573071ad2e3e98 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/MaiqiVerse
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5b2559daa01dce239e66c6dfd2573071ad2e3e98 -
Trigger Event:
release
-
Statement type: