This release is a pre-release and may not be stable for production use.
ReproLLM
Make LLM experiments reproducible.
Status: alpha. PyPI 0.1.1 supports Audit Level 0/1 and init. The development
checkout adds lock, runtime capture, Level 2 consistency checks and semantic
diff. The 0.2/0.3/0.4 releases still await their validation gates.
A reproducibility linter, experiment recorder, lockfile system, and drift detector for LLM research. It records the LLM-specific state that other tools ignore — model revision, tokenizer and chat-template hashes, prompt hashes, generation parameters, LLM-as-a-Judge configuration, and the pinnability of closed-source API models — and tells you why two runs differ.
ReproLLM answers two questions:
- Does this experiment contain enough information for someone to understand, rebuild, and compare it later?
- Why is this run different from that run?
Core principle: LLM discovers. Rules decide. Runtime verifies. Known reproducibility requirements are checked by deterministic rules. Runtime capture records what actually happened, independent of what was declared. An optional, opt-in LLM step only proposes candidates for project-specific parameters; a candidate takes effect only after you explicitly accept it.
Quick start (development checkout)
$ pip install . # inside a checkout of ReproLLM main
$ cd your-llm-experiment
$ reprollm audit .
ReproLLM audit · level 0 · profiles: core
WARNING (3)
! env.llm_critical_deps_pinned vllm is used but not pinned to an exact version (suggestion: vllm==<version>)
requirements.txt (declared as >=0.10)
fix: Pin vllm exactly, e.g. `vllm==<version>`.
…
6 passed · 0 suppressed · 1 skipped
Result: FAIL (3 warning)
Detected profiles: evaluation (medium), inference (high) — run: reprollm init --profiles evaluation,inference
$ reprollm init --profiles evaluation,inference # or plain `reprollm init`
Created reprollm.yaml (profiles: evaluation, inference; 7 required fields to fill)
Next: fill the TODO fields, then run `reprollm audit .`
$ $EDITOR reprollm.yaml # fill the TODOs
$ reprollm audit . # now at level 1: model/dataset/generation gaps
$ reprollm lock . # resolve revisions and hash declared inputs
$ reprollm audit . # now at level 2: verify locked identities and files
$ reprollm lock . --check # network-free manifest/project-rule freshness check
$ reprollm run -- python eval.py --temperature 0.0
$ reprollm runs list
$ reprollm runs show <run-id-or-unique-prefix>
$ reprollm audit . # compare the latest run against declarations and lock
$ reprollm diff <run-a-id> <run-b-id> --fail-on HIGH
Without any configuration ReproLLM audits your repository at Level 0 (code
state, dependency pins, secret files, detected experiment types). With a
reprollm.yaml manifest it audits at Level 1 (what your experiment is
missing to be rebuildable). The output above is real output from an evaluation
repository — nothing is fabricated.
The lock records uncertainty explicitly. For example, the API judge in the FastChat dogfooding manifest has this identity (other fields omitted):
models:
judge:
provider: openai
id: gpt-4o-2024-08-06
pinnability: snapshot_alias
revision:
value: null
source: provider_no_pinning
confidence: unresolved
Read the lockfile guide for provenance, offline mode, authentication, file changes, and the distinction between exact revisions, provider snapshot labels, and mutable aliases.
Diff identifies changes by their experiment meaning. This excerpt is from the
tested hf_vllm_eval fixture, which captures a small python -c pass command
before and after changing a prompt and the temperature; it is not a model run:
prompts
prompts.system.sha256 sha256:1012fd0edf9e → sha256:f3773a23d712 [HIGH]
prompts.system.size_bytes 162 → 180 [MEDIUM]
files
files.prompts/system.txt.sha256 sha256:1012fd0edf9e → sha256:f3773a23d712 [HIGH]
generation
generation.temperature 1.0 → 0.7 [HIGH]
a has inconsistent sources (see audit); b has inconsistent sources (see audit)
The full report concludes: Highest drift: HIGH (3 changes). These runs are not directly comparable. Read the diff guide for input formats,
severity policy, source conflicts, profile overrides and CI exit thresholds.
What it checks
ReproLLM's deterministic rules cover repository and environment state plus the LLM-specific declarations that commonly disappear from experiment records: model/provider identity, dtype and quantization, dataset split and preprocessing, prompt sources, generation and backend settings, evaluation metrics, judge configuration, training hyperparameters, and privacy assumptions. The complete rule catalog records each rule's severity, level, fix, and profile membership; the profile catalog shows inheritance, required fields, severity overrides, and detection signals.
The manifest guide explains model and dataset roles, generation versus judge settings, implementation references, execution bindings, and how to review repository-wide detections in multi-workflow codebases.
Level 2 verifies lock freshness and compares file hashes, observed generation parameters, model identities, LLM package versions and accepted custom fields against the latest valid run. The runtime guide explains bindings, environment policies, redacted snapshots and how to share selected records.
Commands
| Command | Status | Purpose |
|---|---|---|
reprollm audit |
usable (Level 0/1/2 on main) | deterministic reproducibility audit |
reprollm init |
usable | create reprollm.yaml from detected experiment profiles |
reprollm doctor |
usable | environment diagnostics |
reprollm profiles list/show |
usable | inspect the seven built-in profiles |
reprollm lock |
usable on main, planned for 0.2.0 | resolve models/datasets/prompts into a reviewable reprollm.lock |
reprollm run -- CMD |
usable in development, planned for 0.3.0 | execute a command and record runtime truth |
reprollm runs list/show |
usable in development, planned for 0.3.0 | inspect saved runtime evidence |
reprollm diff A B |
usable in development, planned for 0.4.0 | semantic drift between two runs or lockfiles |
reprollm export |
usable | generate a REPRODUCIBILITY.md for your paper artifact |
reprollm rules add/list |
usable | accept repository-specific requirements |
ReproLLM is CLI-first, local-first, and collects no telemetry. The only network calls are
revision resolution against provider APIs (lock), an opt-in LLM endpoint
(discover --experimental), an opt-in doctor --check-network, and version verification
you explicitly request (lock --verify-api).
Roadmap
The architecture and the full Beta specification are frozen in the repository:
-
docs/plan/00_architecture_and_decisions.md— product definition, boundaries, and the decision register (D-01 … D-42) -
docs/plan/01_specification.md— CLI contract, schemas, rule catalog, redaction policy, diff semantics -
docs/index.md— documentation index -
docs/adoption.md— monthly adoption metrics (updated from week 1)
Milestones: M1 foundation → M2 audit core + init → M3 rules + profiles → M4 lock →
M5 run + redaction → M6 diff → M7 export/discover → M8 Beta (0.5.0).
Contributing
See CONTRIBUTING.md. The repository is developed in the open under Apache-2.0.
License
Metadata
Release files for reprollm 0.5.0a1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| reprollm-0.5.0a1.tar.gz | 270.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| reprollm-0.5.0a1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 447.4 kB
Release files / reprollm-0.5.0a1.tar.gz
| Download URL | reprollm-0.5.0a1.tar.gz |
|---|---|
| Size | 270.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fe31743482d75f748770dd706929e1b5586460078415aacb74dda66e06db4266
|
|
BLAKE2b-256 checksum How to use checksums |
bd16f3b2ce63e9bc595f2e98dd96e294875ca2fdb9316d7056754f4e0296bb58
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / reprollm-0.5.0a1-py3-none-any.whl
| Download URL | reprollm-0.5.0a1-py3-none-any.whl |
|---|---|
| Size | 176.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0c0a7cdf53150615fb5170553cdbcd5b5ed94e592a36b1682205a394ce45cb84
|
|
BLAKE2b-256 checksum How to use checksums |
8d5bc870352130e08e9bbeefa0543fdc39e04f220ce18be9daa88664b37b1790
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log