MergeLens
MergeLens is an experimental inspection toolkit for homologous LLM checkpoints. It reports exact tensor coverage and weight, spectral, task-vector, and optional activation-similarity signals; highlights tensors worth inspecting; and proposes a rule-based MergeKit starting configuration. It is intended for model-merging researchers and engineers who need inspectable evidence before spending compute on candidate merges.
MergeLens does not establish downstream merged-model quality, capability retention, or the best merge method. Its aggregate score and thresholds are hand-specified, unvalidated heuristics. Post-merge behavioural evaluation remains necessary.
Run a local software demonstration with tiny synthetic safetensors:
pip install -e .
pip install -e '.[report]'
python examples/synthetic_demo.py --output-dir demo-output
The example is a software and known-answer demonstration, not scientific validation of merge outcomes.
Install
pip install mergelens
pip install 'mergelens[report]' # offline HTML reports
pip install 'mergelens[mcp]' # MCP server
Compare checkpoints
mergelens compare model_a/ model_b/
mergelens compare finetune_a/ finetune_b/ --base shared_base/
mergelens compare model_a/ model_b/ --metric cosine_similarity --metric l2_distance --json result.json
mergelens compare model_a/ model_b/ --report report.html
The result exposes:
- the reference and candidate identity for every pair-tensor row;
- separately attributable
candidate_set_metricsfor sign/TSV andactivation_metricsfor CKA; - total tensors and parameters, missing names, shape/dtype issues, and exact comparable coverage;
- architecture metadata and known structural incompatibilities;
- a status and reason for every computed, skipped, unavailable, failed, or resource-limited signal;
- raw parameter-weighted heuristic components, with partial availability reducing component weight by parameter coverage;
- a machine-readable
validation_status: heuristic_unvalidated; - pair-bounded tensor inspection regions; and
- an illustrative or parser-validated MergeKit starting configuration.
Structurally unsupported comparisons retain their raw coverage and measurements but suppress aggregate scoring.
Python API:
from mergelens import compare_models
result = compare_models(["model_a/", "model_b/"])
print(result.coverage[0].parameter_coverage_reference)
print(result.mci.score) # float or None when suppressed
print(result.mci.risk_tier)
print(result.mci.validation_status) # heuristic_unvalidated
for row in result.tensor_metrics:
print(row.reference_model, row.candidate_model, row.tensor_name, row.cosine_similarity)
for row in result.candidate_set_metrics:
print(row.base_model, row.candidate_models, row.tensor_name, row.sign_disagreement_rate)
for row in result.activation_metrics:
print(row.comparison_id, row.activation_layer, row.cka_similarity, row.warnings)
for signal in result.metric_availability:
print(signal.metric, signal.status.value, signal.reason)
Nine underlying diagnostic signals
The composite heuristic is not counted as a separate diagnostic signal.
| Signal | Level | Direct object | Default | Composite |
|---|---|---|---|---|
| Cosine similarity | Weight | Exact-shape flattened tensor alignment | Yes | Yes |
| Normalized L2 distance | Weight | Difference relative to average tensor norm | Yes | No; displayed raw |
| Weight-distribution divergence | Weight, experimental | Directional softmax transform of flattened weights | No; explicit selection only | No |
| Spectral overlap | Weight | Leading left-singular-subspace overlap for matrices | Yes, resource bounded | Yes |
| Effective-rank ratio | Weight | Ratio of entropy-derived effective ranks | Yes, resource bounded | Yes |
| Sign disagreement | Candidate set task vector | Pairwise sign mismatch; zero/nonzero counts as mismatch | Yes when a shared base and at least two candidates exist | No |
| TSV interference | Candidate set task vector | Pairwise numerical-rank right-subspace overlap | Yes when a shared base and at least two candidates exist | No |
| Task-vector energy | Pair tensor task vector | Fraction of spectral energy in retained leading values | Yes with an explicit base | No |
| Linear CKA | Activation layer, optional | Exact activation-layer observations with calibration and feature-width provenance | Only when supplied | No |
All SVD-backed signals use a conservative full-decomposition resource policy and retain only numerical-rank directions. Full-ambient subspaces are reported as uninformative. A metric skipped by the resource policy is reported as resource_limit_skipped; it is not silently converted into a plausible number.
MergeKit configuration diagnosis
mergelens diagnose merge.yaml --json diagnosis.json
Diagnosis honours checkpoint references, an explicit task-vector base, and finite non-negative scalar weights only for top-level full-model inputs. It discloses ignored or unsupported semantics such as slice assembly, gradients, tokenizer remapping, chat templates, and method-specific merge execution. Unknown merge methods fail closed instead of becoming linear.
Generated configurations follow current MergeKit model/parameter placement. They are marked schema_validated only when the installed MergeKit parser accepted them; otherwise they are marked illustrative.
Reports and MCP
mergelens[report] produces one HTML file with Plotly JavaScript embedded. Charts group by explicit comparison ID and preserve missing metric values.
The MCP server exposes seven tools: compare_models, diagnose_merge, get_conflict_zones, suggest_strategy, generate_report, explain_layer, and get_compatibility_score.
{
"mcpServers": {
"mergelens": {
"command": "mergelens",
"args": ["serve"]
}
}
}
Memory and reproducibility
Safetensors are memory-mapped and aligned tensor groups are consumed lazily. No exact peak-memory multiplier is claimed: runtime memory also includes float32 conversions, task vectors, bounded SVD workspaces, activation tensors, result rows, report data, and framework overhead.
See limitations, validation status, migration guidance, and the changelog before interpreting results.
Development
python -m pip install -e '.[dev,all]'
ruff check .
ruff format --check .
pytest -q
mypy src/mergelens
python -m build
Supported Python versions are 3.10, 3.11, and 3.12.
References
- MergeKit documentation and schema: arcee-ai/mergekit
- Kornblith et al., “Similarity of Neural Network Representations Revisited”: arXiv:1905.00414
- Yadav et al., “TIES-Merging”: arXiv:2306.01708
- Gargiulo et al., “Task Singular Vectors”: arXiv:2412.00081
License
Apache-2.0. See the license.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mergelens-2.0.0.tar.gz.
File metadata
- Download URL: mergelens-2.0.0.tar.gz
- Upload date:
- Size: 75.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
01595c598ef0cee61250af23e47fe8e621c6ded6b082ea1f842e62e240d967db
|
|
| MD5 |
59e6c3df2d34ce8a92bc7a312d7fec98
|
|
| BLAKE2b-256 |
16cc7a9447712818effbb16858f5e6017cbd570582372097b471461916e71569
|
File details
Details for the file mergelens-2.0.0-py3-none-any.whl.
File metadata
- Download URL: mergelens-2.0.0-py3-none-any.whl
- Upload date:
- Size: 62.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cb39e0c614b9477266a363c207c9dca36e0741e7fcac871377bf5de00a542ae4
|
|
| MD5 |
c88ffd65b9feb3fdb335bfaa0b3d3b9a
|
|
| BLAKE2b-256 |
50427f217933b0c938d5c83fb3e13f2b3ef86314e7367ab2b1c05bf280c84de3
|