Skip to main content

ReproLLM

CI PyPI License: Apache-2.0

Make LLM experiments reproducible.

Status: beta (0.5.0). The full workflow — audit, init, lock, run, diff, export, rules, and experimental discover — is usable.

A reproducibility linter, experiment recorder, lockfile system, and drift detector for LLM research. It records the LLM-specific state that other tools ignore — model revision, tokenizer and chat-template hashes, prompt hashes, generation parameters, LLM-as-a-Judge configuration, and the pinnability of closed-source API models — and tells you why two runs differ.

ReproLLM answers two questions:

  1. Does this experiment contain enough information for someone to understand, rebuild, and compare it later?
  2. Why is this run different from that run?

Core principle: LLM discovers. Rules decide. Runtime verifies. Known reproducibility requirements are checked by deterministic rules. Runtime capture records what actually happened, independent of what was declared. An optional, opt-in LLM step only proposes candidates for project-specific parameters; a candidate takes effect only after you explicitly accept it.

Quick start (development checkout)

$ pip install .             # inside a checkout of ReproLLM main
$ cd your-llm-experiment
$ reprollm audit .
ReproLLM audit · level 0 · profiles: core

WARNING (3)
  ! env.llm_critical_deps_pinned      vllm is used but not pinned to an exact version (suggestion: vllm==<version>)
      requirements.txt (declared as >=0.10)
      fix: Pin vllm exactly, e.g. `vllm==<version>`.
  …

6 passed · 0 suppressed · 1 skipped
Result: FAIL (3 warning)

Detected profiles: evaluation (medium), inference (high) — run: reprollm init --profiles evaluation,inference

$ reprollm init --profiles evaluation,inference   # or plain `reprollm init`
Created reprollm.yaml (profiles: evaluation, inference; 7 required fields to fill)
Next: fill the TODO fields, then run `reprollm audit .`

$ $EDITOR reprollm.yaml     # fill the TODOs
$ reprollm audit .          # now at level 1: model/dataset/generation gaps
$ reprollm lock .           # resolve revisions and hash declared inputs
$ reprollm audit .          # now at level 2: verify locked identities and files
$ reprollm lock . --check   # network-free manifest/project-rule freshness check
$ reprollm run -- python eval.py --temperature 0.0
$ reprollm runs list
$ reprollm runs show <run-id-or-unique-prefix>
$ reprollm audit .          # compare the latest run against declarations and lock
$ reprollm diff <run-a-id> <run-b-id> --fail-on HIGH

Without any configuration ReproLLM audits your repository at Level 0 (code state, dependency pins, secret files, detected experiment types). With a reprollm.yaml manifest it audits at Level 1 (what your experiment is missing to be rebuildable). The output above is real output from an evaluation repository — nothing is fabricated.

The lock records uncertainty explicitly. For example, the API judge in the FastChat dogfooding manifest has this identity (other fields omitted):

models:
  judge:
    provider: openai
    id: gpt-4o-2024-08-06
    pinnability: snapshot_alias
    revision:
      value: null
      source: provider_no_pinning
      confidence: unresolved

Read the lockfile guide for provenance, offline mode, authentication, file changes, and the distinction between exact revisions, provider snapshot labels, and mutable aliases.

Diff identifies changes by their experiment meaning. This excerpt is from the tested hf_vllm_eval fixture, which captures a small python -c pass command before and after changing a prompt and the temperature; it is not a model run:

prompts
  prompts.system.sha256  sha256:1012fd0edf9e → sha256:f3773a23d712  [HIGH]
  prompts.system.size_bytes  162 → 180  [MEDIUM]

files
  files.prompts/system.txt.sha256  sha256:1012fd0edf9e → sha256:f3773a23d712  [HIGH]

generation
  generation.temperature  1.0 → 0.7  [HIGH]
    a has inconsistent sources (see audit); b has inconsistent sources (see audit)

The full report concludes: Highest drift: HIGH (3 changes). These runs are not directly comparable. Read the diff guide for input formats, severity policy, source conflicts, profile overrides and CI exit thresholds.

What it checks

ReproLLM's deterministic rules cover repository and environment state plus the LLM-specific declarations that commonly disappear from experiment records: model/provider identity, dtype and quantization, dataset split and preprocessing, prompt sources, generation and backend settings, evaluation metrics, judge configuration, training hyperparameters, and privacy assumptions. The complete rule catalog records each rule's severity, level, fix, and profile membership; the profile catalog shows inheritance, required fields, severity overrides, and detection signals.

The manifest guide explains model and dataset roles, generation versus judge settings, implementation references, execution bindings, and how to review repository-wide detections in multi-workflow codebases.

Level 2 verifies lock freshness and compares file hashes, observed generation parameters, model identities, LLM package versions and accepted custom fields against the latest valid run. The runtime guide explains bindings, environment policies, redacted snapshots and how to share selected records.

Use in CI (one line)

- uses: EnumaElish123/reprollm-action@v1

Findings appear as inline PR annotations; no reprollm.yaml needed for Level 0. See the Action repo.

The 30-second demo

$ cd your-llm-experiment
$ reprollm audit .
ReproLLM audit · level 0 · profiles: core

WARNING (3)
  ! env.llm_critical_deps_pinned      vllm is used but not pinned...

Detected profiles: evaluation (medium), inference (high)
  — run: reprollm init --profiles evaluation,inference

Commands

Command Purpose
reprollm audit deterministic reproducibility audit (Level 0/1/2)
reprollm init scaffold reprollm.yaml from detected experiment profiles
reprollm lock resolve models/datasets/prompts into a reviewable reprollm.lock
reprollm run -- CMD execute a command and record runtime truth
reprollm runs list/show inspect saved runtime evidence
reprollm diff A B semantic drift between two runs or lockfiles
reprollm export generate REPRODUCIBILITY.md for your paper artifact
reprollm rules add/list/accept/ignore repository-specific requirements
reprollm discover --experimental opt-in LLM candidate discovery
reprollm doctor environment diagnostics
reprollm profiles list/show inspect the seven built-in profiles

Docs: Quick start · Concepts · Why ReproLLM · CLI reference · Reading a diff · FAQ ReproLLM is CLI-first, local-first, and collects no telemetry. The only network calls are revision resolution against provider APIs (lock), an opt-in LLM endpoint (discover --experimental), an opt-in doctor --check-network, and version verification you explicitly request (lock --verify-api).

Roadmap

The architecture and the full Beta specification are frozen in the repository:

Milestones: M1 foundation → M2 audit core + init → M3 rules + profiles → M4 lock → M5 run + redaction → M6 diff → M7 export/discover → M8 Beta (0.5.0).

Contributing

See CONTRIBUTING.md. The repository is developed in the open under Apache-2.0.

License

Apache-2.0

Metadata

Release files for reprollm 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for reprollm 0.6.0
File Size Uploaded
reprollm-0.6.0.tar.gz 370.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for reprollm 0.6.0
File Interpreter ABI Platform
reprollm-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 555.5 kB

Release files / reprollm-0.6.0.tar.gz

Download URL reprollm-0.6.0.tar.gz
Size 370.8 kB
Tags Source
SHA-256 checksum
How to use checksums
2e2746ac3bcc1ec5d1173ad576968aa65c223c060153eb6d4398f7bd115a2108
BLAKE2b-256 checksum
How to use checksums
c5ca6a08f2b16106306a89d15d1c9d0cbad8ac51a4bf034ae50f55d5d047bcfd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / reprollm-0.6.0-py3-none-any.whl

Download URL reprollm-0.6.0-py3-none-any.whl
Size 184.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a49a3d36301c6fb537f15a38b39e576eff370d05c404fa8caf596936f503ac8f
BLAKE2b-256 checksum
How to use checksums
b1c8c46f13cb7183491190ebed95198335dfbf09435389b11bbe4581f3dc0f1d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.1

2 release files

This release

0.6.0 This release

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page