Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

ReproLLM

CI PyPI License: Apache-2.0

Make LLM experiments reproducible.

Status: alpha. PyPI 0.1.1 supports Audit Level 0/1 and init. The development checkout adds lock, runtime capture, Level 2 consistency checks and semantic diff. The 0.2/0.3/0.4 releases still await their validation gates.

A reproducibility linter, experiment recorder, lockfile system, and drift detector for LLM research. It records the LLM-specific state that other tools ignore — model revision, tokenizer and chat-template hashes, prompt hashes, generation parameters, LLM-as-a-Judge configuration, and the pinnability of closed-source API models — and tells you why two runs differ.

ReproLLM answers two questions:

  1. Does this experiment contain enough information for someone to understand, rebuild, and compare it later?
  2. Why is this run different from that run?

Core principle: LLM discovers. Rules decide. Runtime verifies. Known reproducibility requirements are checked by deterministic rules. Runtime capture records what actually happened, independent of what was declared. An optional, opt-in LLM step only proposes candidates for project-specific parameters; a candidate takes effect only after you explicitly accept it.

Quick start (development checkout)

$ pip install .             # inside a checkout of ReproLLM main
$ cd your-llm-experiment
$ reprollm audit .
ReproLLM audit · level 0 · profiles: core

WARNING (3)
  ! env.llm_critical_deps_pinned      vllm is used but not pinned to an exact version (suggestion: vllm==<version>)
      requirements.txt (declared as >=0.10)
      fix: Pin vllm exactly, e.g. `vllm==<version>`.
  …

6 passed · 0 suppressed · 1 skipped
Result: FAIL (3 warning)

Detected profiles: evaluation (medium), inference (high) — run: reprollm init --profiles evaluation,inference

$ reprollm init --profiles evaluation,inference   # or plain `reprollm init`
Created reprollm.yaml (profiles: evaluation, inference; 7 required fields to fill)
Next: fill the TODO fields, then run `reprollm audit .`

$ $EDITOR reprollm.yaml     # fill the TODOs
$ reprollm audit .          # now at level 1: model/dataset/generation gaps
$ reprollm lock .           # resolve revisions and hash declared inputs
$ reprollm audit .          # now at level 2: verify locked identities and files
$ reprollm lock . --check   # network-free manifest/project-rule freshness check
$ reprollm run -- python eval.py --temperature 0.0
$ reprollm runs list
$ reprollm runs show <run-id-or-unique-prefix>
$ reprollm audit .          # compare the latest run against declarations and lock
$ reprollm diff <run-a-id> <run-b-id> --fail-on HIGH

Without any configuration ReproLLM audits your repository at Level 0 (code state, dependency pins, secret files, detected experiment types). With a reprollm.yaml manifest it audits at Level 1 (what your experiment is missing to be rebuildable). The output above is real output from an evaluation repository — nothing is fabricated.

The lock records uncertainty explicitly. For example, the API judge in the FastChat dogfooding manifest has this identity (other fields omitted):

models:
  judge:
    provider: openai
    id: gpt-4o-2024-08-06
    pinnability: snapshot_alias
    revision:
      value: null
      source: provider_no_pinning
      confidence: unresolved

Read the lockfile guide for provenance, offline mode, authentication, file changes, and the distinction between exact revisions, provider snapshot labels, and mutable aliases.

Diff identifies changes by their experiment meaning. This excerpt is from the tested hf_vllm_eval fixture, which captures a small python -c pass command before and after changing a prompt and the temperature; it is not a model run:

prompts
  prompts.system.sha256  sha256:1012fd0edf9e → sha256:f3773a23d712  [HIGH]
  prompts.system.size_bytes  162 → 180  [MEDIUM]

files
  files.prompts/system.txt.sha256  sha256:1012fd0edf9e → sha256:f3773a23d712  [HIGH]

generation
  generation.temperature  1.0 → 0.7  [HIGH]
    a has inconsistent sources (see audit); b has inconsistent sources (see audit)

The full report concludes: Highest drift: HIGH (3 changes). These runs are not directly comparable. Read the diff guide for input formats, severity policy, source conflicts, profile overrides and CI exit thresholds.

What it checks

ReproLLM's deterministic rules cover repository and environment state plus the LLM-specific declarations that commonly disappear from experiment records: model/provider identity, dtype and quantization, dataset split and preprocessing, prompt sources, generation and backend settings, evaluation metrics, judge configuration, training hyperparameters, and privacy assumptions. The complete rule catalog records each rule's severity, level, fix, and profile membership; the profile catalog shows inheritance, required fields, severity overrides, and detection signals.

The manifest guide explains model and dataset roles, generation versus judge settings, implementation references, execution bindings, and how to review repository-wide detections in multi-workflow codebases.

Level 2 verifies lock freshness and compares file hashes, observed generation parameters, model identities, LLM package versions and accepted custom fields against the latest valid run. The runtime guide explains bindings, environment policies, redacted snapshots and how to share selected records.

Commands

Command Status Purpose
reprollm audit usable (Level 0/1/2 on main) deterministic reproducibility audit
reprollm init usable create reprollm.yaml from detected experiment profiles
reprollm doctor usable environment diagnostics
reprollm profiles list/show usable inspect the seven built-in profiles
reprollm lock usable on main, planned for 0.2.0 resolve models/datasets/prompts into a reviewable reprollm.lock
reprollm run -- CMD usable in development, planned for 0.3.0 execute a command and record runtime truth
reprollm runs list/show usable in development, planned for 0.3.0 inspect saved runtime evidence
reprollm diff A B usable in development, planned for 0.4.0 semantic drift between two runs or lockfiles
reprollm export usable generate a REPRODUCIBILITY.md for your paper artifact
reprollm rules add/list usable accept repository-specific requirements

ReproLLM is CLI-first, local-first, and collects no telemetry. The only network calls are revision resolution against provider APIs (lock), an opt-in LLM endpoint (discover --experimental), an opt-in doctor --check-network, and version verification you explicitly request (lock --verify-api).

Roadmap

The architecture and the full Beta specification are frozen in the repository:

Milestones: M1 foundation → M2 audit core + init → M3 rules + profiles → M4 lock → M5 run + redaction → M6 diff → M7 export/discover → M8 Beta (0.5.0).

Contributing

See CONTRIBUTING.md. The repository is developed in the open under Apache-2.0.

License

Apache-2.0

Metadata

Release files for reprollm 0.5.0a1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for reprollm 0.5.0a1
File Size Uploaded
reprollm-0.5.0a1.tar.gz 270.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for reprollm 0.5.0a1
File Interpreter ABI Platform
reprollm-0.5.0a1-py3-none-any.whl Python 3 none any Details

Total release size: 447.4 kB

Release files / reprollm-0.5.0a1.tar.gz

Download URL reprollm-0.5.0a1.tar.gz
Size 270.8 kB
Tags Source
SHA-256 checksum
How to use checksums
fe31743482d75f748770dd706929e1b5586460078415aacb74dda66e06db4266
BLAKE2b-256 checksum
How to use checksums
bd16f3b2ce63e9bc595f2e98dd96e294875ca2fdb9316d7056754f4e0296bb58
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / reprollm-0.5.0a1-py3-none-any.whl

Download URL reprollm-0.5.0a1-py3-none-any.whl
Size 176.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0c0a7cdf53150615fb5170553cdbcd5b5ed94e592a36b1682205a394ce45cb84
BLAKE2b-256 checksum
How to use checksums
8d5bc870352130e08e9bbeefa0543fdc39e04f220ce18be9daa88664b37b1790
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.1

2 release files

0.6.0

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

This release

0.5.0a1 This release

2 release files

0.4.0

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page