Measured optimization of the text layer around a language model: versioned artifacts, held-out evaluation, leakage detection, and a complete audit record.
Project description
recension
Measured optimization of the text layer around a language model (prompts, context templates, skill and instruction files) with the rigor normally reserved for weight training: a held-out objective, a baseline, versioned artifacts, and a complete audit trail.
The name comes from textual criticism. A recension is the revision of a text by collating variant readings and keeping the best-supported one. That is the loop this library runs: propose multiple candidate edits, test each against held-out evidence, commit only what measurably improves, and record why.
Why
The usual way to improve a prompt is to edit it, eyeball a few outputs, and ship. There is no held-out measurement, no record of why a change was made, and no defense against overfitting to the handful of cases you inspected. recension replaces that loop with a measured one:
- No edit is accepted without a held-out score that beats the incumbent. Failures are diagnosed on a train split; acceptance happens only on a validation split, and can require the gain to be statistically significant rather than above an epsilon. An optional locked test split gives an unbiased final estimate.
- Every accepted version carries provenance: the failures that motivated it, the diagnosis, every sibling candidate considered (with scores), and the diff. A reviewer who didn't run the optimization can reconstruct every decision.
- One number doesn't hide regressions. Optional per-slice scores, non-regression guard objectives, and a token-cost ledger surface what an aggregate averages away.
- Leakage is checked, not assumed away. Heuristics flag candidates that embed validation content or show implausible validation gains.
- Records are built to be acted on. Gate a prompt in CI with
recension check, detect tampering withrecension verify, and share a standalone HTML audit withrecension report. - Compute is a dial. Candidates per round, rounds, diagnosis depth, and a hard ceiling on model calls are all caller-controlled.
Real-world use cases
- Production classification and extraction (support-ticket triage, invoice fields, moderation): improve a labeling prompt on labeled data with measured, regression-safe edits.
- RAG context templates: tune how retrieved chunks are assembled into the prompt with the model held fixed, so the metric move is attributable to the text.
- Agent and skill instructions: optimize longer instruction files judged by an
LLMJudgerubric when there is no gold answer. - Governance and audit: ship a replayable, tamper-evident
RunRecordfor every prompt change; gate merges in CI withrecension check, and hand reviewers a standalone HTML audit withrecension report.
Full write-ups, plus a "how it works" walkthrough, are on the documentation site.
Prior art, honestly
DSPy and GEPA own the optimization mechanics this library's ReflectiveOptimizer performs; if you want state-of-the-art prompt optimization algorithms, look there. recension's contribution is the measurement and governance shell around a text artifact: versioned artifacts with provenance, leakage detection, the complete audit record, and budgeted update-time compute. That delegation is a real seam, not a promise: the Proposer protocol lets an external engine supply the candidate edits while recension keeps owning the artifact, the measurement, and the record (see the Bring your own optimizer guide).
Install
pip install recension # core: zero provider dependencies
pip install "recension[anthropic]" # Anthropic backend
pip install "recension[openai]" # OpenAI (and OpenAI-compatible servers)
pip install "recension[gemini]" # Google Gemini
# Ollama (local models) needs no extra: it uses the standard library
Python 3.12+. The core (and the whole test suite) runs against a deterministic MockModel with no API key and no network.
Quickstart
from recension import (
Budget, EvalSet, ExactMatch, MockModel, ReflectiveOptimizer, TextArtifact,
)
artifact = TextArtifact.from_text("Label the sentiment of the message.")
# Held-out examples, split into train (for diagnosis) and validation (for
# acceptance). Load from a JSONL file with EvalSet.from_jsonl(path) instead.
evalset = EvalSet.from_records([
{"id": "t1", "input": "Absolutely love this", "expected": "positive", "split": "train"},
{"id": "t2", "input": "Broke after a day", "expected": "negative", "split": "train"},
{"id": "v1", "input": "Terrible support", "expected": "negative", "split": "validation"},
{"id": "v2", "input": "Exceeded expectations", "expected": "positive", "split": "validation"},
])
optimizer = ReflectiveOptimizer(
artifact=artifact,
evalset=evalset,
objective=ExactMatch(),
model=MockModel(), # offline mock; see below for the real backend
budget=Budget(candidates_per_round=4, rounds=3, max_model_calls=200),
seed=7,
)
record = optimizer.run()
print(record.summary()) # baseline → accepted versions → final score
record.save("run_record.json") # the complete audit artifact
To run against a real model, install the matching extra, set the provider's key in your environment, and pass the backend. The core never imports a provider, so each is kept off the top-level import:
from recension.models.anthropic import AnthropicModel # ANTHROPIC_API_KEY
from recension.models.openai import OpenAIModel # OPENAI_API_KEY (also OpenAI-compatible servers via base_url)
from recension.models.gemini import GeminiModel # GEMINI_API_KEY / GOOGLE_API_KEY
from recension.models.ollama import OllamaModel # local Ollama server, no key
optimizer = ReflectiveOptimizer(..., model=OpenAIModel(model="gpt-4o-mini"))
Any object satisfying the small Model protocol works, so bringing your own provider is a ~30-line adapter (MockModel is a worked template). API keys are read from the environment only, never from code or config. One honesty note: deterministic, reproducible runs are a guarantee of MockModel only; hosted APIs treat seed as a best-effort hint or ignore it.
CLI
recension run --config run.yaml # execute an optimization, write the record
recension show run_record.json # baseline, accepted diffs, score progression, integrity
recension diff run_record.json vA vB # diff between two artifact versions
recension check --config run.yaml --baseline run_record.json # CI guard: exit non-zero on regression
recension verify run_record.json # detect tampering (content-addressed version chain)
recension report run_record.json -o report.html # standalone HTML audit page
A runnable, fully commented config and dataset live in examples/cli/ (recension run --config examples/cli/run.yaml); the config schema and a GitHub Actions recipe for recension check are in the CLI guide.
Documentation
Full docs, API reference, and three worked examples (each reproducible offline against MockModel): https://anthonynystrom.github.io/recension/
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file recension-0.6.0.tar.gz.
File metadata
- Download URL: recension-0.6.0.tar.gz
- Upload date:
- Size: 65.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a3196774dc982be4aee763b08761907ffada95ca763deaaa058c3f30c489ab73
|
|
| MD5 |
2a0b17584f594bdaecef0822d3f735f3
|
|
| BLAKE2b-256 |
44fdd48bbcd5ed50e22730096cf48575626f3fe2a8f25e8af20694a7f42ba0a6
|
Provenance
The following attestation bundles were made for recension-0.6.0.tar.gz:
Publisher:
publish.yml on AnthonyNystrom/recension
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
recension-0.6.0.tar.gz -
Subject digest:
a3196774dc982be4aee763b08761907ffada95ca763deaaa058c3f30c489ab73 - Sigstore transparency entry: 1791905697
- Sigstore integration time:
-
Permalink:
AnthonyNystrom/recension@48c6f40d6f2709daaa5fa9196771ebd2f6e20659 -
Branch / Tag:
refs/tags/v0.6.0 - Owner: https://github.com/AnthonyNystrom
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@48c6f40d6f2709daaa5fa9196771ebd2f6e20659 -
Trigger Event:
release
-
Statement type:
File details
Details for the file recension-0.6.0-py3-none-any.whl.
File metadata
- Download URL: recension-0.6.0-py3-none-any.whl
- Upload date:
- Size: 57.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5ad878120e29fb215dde689d4c093d271d9bb06bd4f735ab70afc11ef8b1a945
|
|
| MD5 |
2e324c8b9d53bdeb6399b5c2dbb1fd03
|
|
| BLAKE2b-256 |
eacc9355cdc3530a28e5fc0462b39a4c7354922e6bb142af156a64ece55002ac
|
Provenance
The following attestation bundles were made for recension-0.6.0-py3-none-any.whl:
Publisher:
publish.yml on AnthonyNystrom/recension
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
recension-0.6.0-py3-none-any.whl -
Subject digest:
5ad878120e29fb215dde689d4c093d271d9bb06bd4f735ab70afc11ef8b1a945 - Sigstore transparency entry: 1791905797
- Sigstore integration time:
-
Permalink:
AnthonyNystrom/recension@48c6f40d6f2709daaa5fa9196771ebd2f6e20659 -
Branch / Tag:
refs/tags/v0.6.0 - Owner: https://github.com/AnthonyNystrom
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@48c6f40d6f2709daaa5fa9196771ebd2f6e20659 -
Trigger Event:
release
-
Statement type: