Skip to main content

ReproLLM

CI PyPI License: Apache-2.0

Make LLM experiments reproducible.

Status: alpha. PyPI 0.1.1 supports Audit Level 0/1 and init. The development checkout adds lock, runtime capture, Level 2 consistency checks and semantic diff. The 0.2/0.3/0.4 releases still await their validation gates.

A reproducibility linter, experiment recorder, lockfile system, and drift detector for LLM research. It records the LLM-specific state that other tools ignore — model revision, tokenizer and chat-template hashes, prompt hashes, generation parameters, LLM-as-a-Judge configuration, and the pinnability of closed-source API models — and tells you why two runs differ.

ReproLLM answers two questions:

  1. Does this experiment contain enough information for someone to understand, rebuild, and compare it later?
  2. Why is this run different from that run?

Core principle: LLM discovers. Rules decide. Runtime verifies. Known reproducibility requirements are checked by deterministic rules. Runtime capture records what actually happened, independent of what was declared. An optional, opt-in LLM step only proposes candidates for project-specific parameters; a candidate takes effect only after you explicitly accept it.

Quick start (development checkout)

$ pip install .             # inside a checkout of ReproLLM main
$ cd your-llm-experiment
$ reprollm audit .
ReproLLM audit · level 0 · profiles: core

WARNING (3)
  ! env.llm_critical_deps_pinned      vllm is used but not pinned to an exact version (suggestion: vllm==<version>)
      requirements.txt (declared as >=0.10)
      fix: Pin vllm exactly, e.g. `vllm==<version>`.
  …

6 passed · 0 suppressed · 1 skipped
Result: FAIL (3 warning)

Detected profiles: evaluation (medium), inference (high) — run: reprollm init --profiles evaluation,inference

$ reprollm init --profiles evaluation,inference   # or plain `reprollm init`
Created reprollm.yaml (profiles: evaluation, inference; 7 required fields to fill)
Next: fill the TODO fields, then run `reprollm audit .`

$ $EDITOR reprollm.yaml     # fill the TODOs
$ reprollm audit .          # now at level 1: model/dataset/generation gaps
$ reprollm lock .           # resolve revisions and hash declared inputs
$ reprollm audit .          # now at level 2: verify locked identities and files
$ reprollm lock . --check   # network-free manifest/project-rule freshness check
$ reprollm run -- python eval.py --temperature 0.0
$ reprollm runs list
$ reprollm runs show <run-id-or-unique-prefix>
$ reprollm audit .          # compare the latest run against declarations and lock
$ reprollm diff <run-a-id> <run-b-id> --fail-on HIGH

Without any configuration ReproLLM audits your repository at Level 0 (code state, dependency pins, secret files, detected experiment types). With a reprollm.yaml manifest it audits at Level 1 (what your experiment is missing to be rebuildable). The output above is real output from an evaluation repository — nothing is fabricated.

The lock records uncertainty explicitly. For example, the API judge in the FastChat dogfooding manifest has this identity (other fields omitted):

models:
  judge:
    provider: openai
    id: gpt-4o-2024-08-06
    pinnability: snapshot_alias
    revision:
      value: null
      source: provider_no_pinning
      confidence: unresolved

Read the lockfile guide for provenance, offline mode, authentication, file changes, and the distinction between exact revisions, provider snapshot labels, and mutable aliases.

Diff identifies changes by their experiment meaning. This excerpt is from the tested hf_vllm_eval fixture, which captures a small python -c pass command before and after changing a prompt and the temperature; it is not a model run:

prompts
  prompts.system.sha256  sha256:1012fd0edf9e → sha256:f3773a23d712  [HIGH]
  prompts.system.size_bytes  162 → 180  [MEDIUM]

files
  files.prompts/system.txt.sha256  sha256:1012fd0edf9e → sha256:f3773a23d712  [HIGH]

generation
  generation.temperature  1.0 → 0.7  [HIGH]
    a has inconsistent sources (see audit); b has inconsistent sources (see audit)

The full report concludes: Highest drift: HIGH (3 changes). These runs are not directly comparable. Read the diff guide for input formats, severity policy, source conflicts, profile overrides and CI exit thresholds.

What it checks

ReproLLM's deterministic rules cover repository and environment state plus the LLM-specific declarations that commonly disappear from experiment records: model/provider identity, dtype and quantization, dataset split and preprocessing, prompt sources, generation and backend settings, evaluation metrics, judge configuration, training hyperparameters, and privacy assumptions. The complete rule catalog records each rule's severity, level, fix, and profile membership; the profile catalog shows inheritance, required fields, severity overrides, and detection signals.

The manifest guide explains model and dataset roles, generation versus judge settings, implementation references, execution bindings, and how to review repository-wide detections in multi-workflow codebases.

Level 2 verifies lock freshness and compares file hashes, observed generation parameters, model identities, LLM package versions and accepted custom fields against the latest valid run. The runtime guide explains bindings, environment policies, redacted snapshots and how to share selected records.

Commands

Command Status Purpose
reprollm audit usable (Level 0/1/2 on main) deterministic reproducibility audit
reprollm init usable create reprollm.yaml from detected experiment profiles
reprollm doctor usable environment diagnostics
reprollm profiles list/show usable inspect the seven built-in profiles
reprollm lock usable on main, planned for 0.2.0 resolve models/datasets/prompts into a reviewable reprollm.lock
reprollm run -- CMD usable in development, planned for 0.3.0 execute a command and record runtime truth
reprollm runs list/show usable in development, planned for 0.3.0 inspect saved runtime evidence
reprollm diff A B usable in development, planned for 0.4.0 semantic drift between two runs or lockfiles
reprollm export 0.5.0 generate a REPRODUCIBILITY.md for your paper artifact

ReproLLM is CLI-first, local-first, and collects no telemetry. The only network calls are revision resolution against provider APIs (lock), an opt-in LLM endpoint (discover --experimental), an opt-in doctor --check-network, and version verification you explicitly request (lock --verify-api).

Roadmap

The architecture and the full Beta specification are frozen in the repository:

Milestones: M1 foundation → M2 audit core + init → M3 rules + profiles → M4 lock → M5 run + redaction → M6 diff → M7 export/discover → M8 Beta (0.5.0).

Contributing

See CONTRIBUTING.md. The repository is developed in the open under Apache-2.0.

License

Apache-2.0

Metadata

Release files for reprollm 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for reprollm 0.4.0
File Size Uploaded
reprollm-0.4.0.tar.gz 253.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for reprollm 0.4.0
File Interpreter ABI Platform
reprollm-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 410.4 kB

Release files / reprollm-0.4.0.tar.gz

Download URL reprollm-0.4.0.tar.gz
Size 253.2 kB
Tags Source
SHA-256 checksum
How to use checksums
6e00d2e8b4f487fd29dc6b1c967a8ad68b5bd372aedf68edf15123fe3dc01add
BLAKE2b-256 checksum
How to use checksums
90f27c7bfb901c099fa3557647eb6d8d4527a7ea3c104afd88af22c60ff56c11
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / reprollm-0.4.0-py3-none-any.whl

Download URL reprollm-0.4.0-py3-none-any.whl
Size 157.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
418ada06cc047a2357fdac782face26049fba7063327c88340098226421fcc25
BLAKE2b-256 checksum
How to use checksums
340c5fe1f79f8016fad90a5358014e91def34a3a8194db305cbeb9ab8e462f65
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.1

2 release files

0.6.0

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

This release

0.4.0 This release

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page