Skip to main content
CJE Logo

CJE — Causal Judge Evaluation

LLM-judge scores are cheap and plentiful, but their scale can differ materially from the outcome you actually care about. In the paper's Chatbot Arena benchmark, naive 95% intervals around raw judge-score means had 0% coverage. CJE calibrates a judge against a small sample of ground-truth labels, evaluates policies on fresh responses, and reports uncertainty and diagnostics under explicit sampling and transport assumptions.

arXiv Dataset Open In Colab Docs Python Tests License PyPI Downloads

60 seconds

pip install cje-eval

Rather delegate? Point your coding agent at the bundled agent skill and it handles everything below — data reshaping, calibration, diagnostics.

You need three things: responses from each policy on a shared prompt set, a score for every response from one fixed LLM judge, and ground-truth labels (oracle_label) on a randomly sampled slice you can afford — human ratings, expert review, or a downstream KPI (stratify the sample by judge score to cover the range). Each record is one judged response: {"prompt_id", "judge_score", "oracle_label" (optional)} — CJE calls these fresh draws: the responses you sampled from each policy for this eval, as opposed to logged production traffic. Any bounded judge and oracle scales work (0–1, 0–100, Likert), and they don't need to match each other.

from cje import analyze_dataset

# Two policies, gpt-5.6 vs fable-5, each answered the same 20 prompts.
# A separate fixed judge model scored all 40 responses; human raters
# labeled a random half of gpt-5.6's (None = not labeled).
judge_scores = {
    "gpt-5.6": [0.62, 0.68, 0.72, 0.76, 0.79, 0.83, 0.85, 0.88, 0.91, 0.95,
                0.64, 0.69, 0.73, 0.77, 0.80, 0.84, 0.87, 0.89, 0.92, 0.94],
    "fable-5": [0.70, 0.74, 0.75, 0.78, 0.81, 0.83, 0.86, 0.90, 0.93, 0.94,
                0.72, 0.76, 0.79, 0.80, 0.84, 0.85, 0.88, 0.89, 0.91, 0.95],
}
human_labels = [0.55, 0.60, 0.70, 0.74, 0.75, 0.80, 0.90, 0.92, 0.88, 0.97,
                None, None, None, None, None, None, None, None, None, None]

# gpt-5.6's labeled slice calibrates the judge for BOTH policies. Reusing
# that map for fable-5 is an assumption; the output flags it as
# "residual transport NOT_CHECKED" until a held-out probe audit grades it.
draws = {
    "gpt-5.6": [
        {"prompt_id": f"q{i:02d}", "judge_score": s, "oracle_label": y}
        for i, (s, y) in enumerate(zip(judge_scores["gpt-5.6"], human_labels))
    ],
    "fable-5": [
        {"prompt_id": f"q{i:02d}", "judge_score": s}
        for i, s in enumerate(judge_scores["fable-5"])
    ],
}
results = analyze_dataset(fresh_draws_data=draws)
print(results.summary())
CJE Estimation Results (method: calibrated_direct)
  fable-5  0.824  95% CI [0.766, 0.882]
  gpt-5.6  0.786  95% CI [0.706, 0.866]
Best by point estimate: fable-5
Limitations: residual transport NOT_CHECKED
Status: warning

Every policy gets a calibrated estimate and a confidence interval — including fable-5, which has no labels of its own. The Limitations line is the guardrails talking: CJE hands you the estimate but never lets an unchecked assumption pass silently (how to clear it). The intervals account for evaluation sampling and the finite label budget; interpreting them still depends on the sampling design and shared-calibration assumptions.

→ Runnable Colab with real data · Full docs

Use CJE from your AI agent

You don't have to learn the API yourself. skills/cje/ teaches a coding agent the full workflow — reshape your eval data, drive the labeling loop, calibrate, compare, respect the refusal gates. It's plain Markdown: any agent can use it, and agents with skill support load it natively. Or just paste:

Read https://raw.githubusercontent.com/cimo-labs/cje/main/skills/cje/SKILL.md,
then use CJE to compare the policies in my eval data.

Is CJE the right tool?

Your situation Use
Rank/compare policies using an LLM judge, with some ground-truth labels CJE
One dataset, labels sampled from it, want a CI on its mean Either works — this is prediction-powered inference (PPI); CJE's calibrated_mean_ci is the same primitive with the diagnostics built in
Evaluate many policies without labeling under each CJE — labels pool across policies; audit that reuse with held-out probes before relying on it
Predict how a specific response will score Not CJE — per-item prediction (conformal methods)
Off-policy estimates from logs only (importance weighting / doubly robust) pip install "cje-eval==0.3.*" — the frozen OPE line; this library is Direct-mode only (see Why Direct mode only?)

How it works

  1. Calibrate: learn the judge → oracle mapping on the labeled slice (isotonic, two-stage when needed; mean-preserving by construction; cross-fitted).
  2. Evaluate: score every policy's fresh responses through the calibrated judge and compare policies on the same prompts.
  3. Diagnose: automatically report scalar score-range support, and optionally run a held-out residual equivalence audit with a predeclared practical margin. These answer different questions and are reported separately.

Confidence intervals include finite-label calibration uncertainty on supported inference paths. Their interpretation still depends on the oracle sampling design, shared-calibration assumptions, and any transport claims being made.

CJE forest plot showing calibrated policy estimates with confidence intervals
Calibrated estimates with 95% CIs under the experiment's stated sampling and calibration assumptions

Validation on real ground truth

  • HealthBench (physician labels, n=29,511): two LLM judges were overconfident by 24.5 and 13.0 points and disagreed with each other by up to 73 points on specific criteria categories. Calibrated on 5% physician labels (~1,400 records), both converged to the physician ground truth. Read the full audit →
  • Chatbot Arena (4,961 prompts, 5 policies): 99% pairwise ranking accuracy at a 5% oracle fraction — 14× cheaper than labeling everything, with ~95% CI coverage vs 0% for naive judge-score CIs. An adversarial policy that fools the judge is correctly flagged by the transport audit. Paper →

Guardrails: claims CJE refuses to make

Diagnostics never act silently — every estimate ships with its limitations attached.

Score-support badge (automatic). Each policy gets a scalar badge checking whether its judge scores extrapolate beyond the labeled score range. When most scores land outside it, the estimate carries REFUSE-LEVEL:

REFUSE-LEVEL for policy 'candidate': 88.3% of fresh-draw judge scores fall
outside the oracle calibration range [0.161, 0.595]. Do not report level
(absolute) claims for this policy from this fit. Collect oracle labels covering
the missing score range.

The badge checks scalar support only — it does not test mean residual bias, covariate shift, or ranking validity.

Residual transport audit (opt-in). Reusing a calibration map on another policy, time period, or domain is an assumption. Grade it with held-out oracle probes that were not used to fit the calibrator, plus a predeclared practical margin:

from cje import TransportAuditConfig

transport = TransportAuditConfig(
    probes_by_policy={"fable-5": held_out_probe_rows},  # same record shape as draws, oracle_label filled
    delta_max_by_policy={"fable-5": 0.03},  # OUTPUT units (units of results.estimates)
)
results = analyze_dataset(fresh_draws_data=draws, transport=transport)
print(results.metadata["transport_audits"]["fable-5"]["status"])

PASS requires the simultaneous residual CI to lie wholly inside [-delta_max, +delta_max]; wholly outside is FAIL; overlap is INCONCLUSIVE; omitting the margin is NOT_GRADED. Fewer than 20 effective clusters withholds PASS but can still grade FAIL — a policy cannot escape a FAIL by supplying too small a probe. Policies without probes stay NOT_CHECKED. Only an observed FAIL hard-flags a policy; every other unresolved state remains visible as a limitation without suppressing the estimate. For an already fitted calibrator, the array primitive transport_audit(probe_scores, probe_labels, results.calibrator, delta_max=...) runs the same audit directly.

Reliability-aware winner. results.best_policy() demotes a gate-flagged argmax to the best gate-passing policy (the default, reliable_only=True), and the demotion is loud — the flagged raw winner stays visible with its limitations (reliable_only=False returns the raw argmax, marked flagged):

Best by point estimate: candidate
Limitations: flagged by the reliability gates; residual transport NOT_CHECKED
Best reliable policy: baseline — raw argmax candidate was flagged (boundary:
88.3% of judge scores outside the oracle calibration range); pass
reliable_only=False for the raw argmax

The array API

calibrated_mean_ci is the library's bottom layer: a ppi_py-style primitive — plain NumPy arrays in, calibrated mean and confidence interval out. Reach for it when you have one sample of judge scores with ground-truth labels on a random slice; use analyze_dataset for multi-policy comparisons. The interval accounts for both sampling noise and the finite label budget (prompt-cluster-robust variance plus a delete-one-oracle-fold jackknife; t interval with Welch–Satterthwaite effective df); inference="bootstrap" switches to refit-bootstrap percentile intervals.

import numpy as np
from cje import calibrated_mean_ci

rng = np.random.default_rng(0)
scores = rng.uniform(size=400)                      # judge scores for every sample
labels = np.full(400, np.nan)                       # NaN = unlabeled
labeled = rng.choice(400, size=100, replace=False)  # oracle slice (25%)
labels[labeled] = np.clip(scores[labeled] + rng.normal(0, 0.1, size=100), 0, 1)

result = calibrated_mean_ci(scores, labels)
print(result.summary())
Calibrated mean: 0.5316 (SE 0.0174, CI [0.4974, 0.5659], n=400, n_oracle=100, cluster_robust)

When partial oracle coverage requires calibration, result.calibrator predicts in the same public judge and oracle units supplied by the caller; complete oracle coverage returns the direct oracle mean with result.calibrator is None. Grade any fitted calibrator's reuse on an independent probe with transport_audit(..., delta_max=<practical margin>); result.diagnostics["boundary_card"] carries the separate scalar score-support badge when calibration is fitted.

Documentation

Resource Description
Interactive Tutorial Walk through a complete example in Colab — no setup required
Agent Skill Teach any coding agent to run CJE correctly
CJE in 3 Minutes Video: why raw judge scores mislead and how CJE fixes it
Technical Walkthrough Video: calibration, evaluation, and transport auditing pipeline
Operational Playbook End-to-end runbook: audits, drift correction, label budgeting
Migration Guide Upgrading from 0.5.x or earlier: what changed and how to adapt
Planning Notebook Optimize your evaluation budget with pilot data
Full Docs Installation, assumptions, API reference, research notes

Bridges: Already running evals in Promptfoo, TruLens, LangSmith, or OpenCompass? Convert those outputs into CJE format with one command.

Module deep dives: Calibration · Diagnostics · Estimators · Interface/API · Data formats

Why Direct mode only (no IPS/DR)?

CJE is Direct-mode only: fresh draws, calibrated judge, audits. There is no off-policy machinery — no importance-sampling or doubly-robust estimators (calibrated-ips, dr-cpo, mrdr, tmle, stacked-dr), teacher forcing, SIMCal weight stabilization, or overlap diagnostics. Our own paper's results drove that design: for realistic LLM policy pairs, importance weighting failed even when ESS looked healthy (target-typicality coverage 0.19–0.49, far below the 0.70 gate), and the best DR stack merely matched Direct mode's accuracy at ~12× the compute. Direct mode is what the evidence supports, so it is the whole product.

  • Need IPS/DR from logged propensities? Pin the frozen OPE line: pip install "cje-eval==0.3.*" (maintained on the 0.3.x branch; docs at the v0.3.0 tag; requires Python <=3.12 — on 3.13 use a 3.12 env for OPE).
  • Have old logged data with judge_score + oracle_label? It works as the calibration source: analyze_dataset(fresh_draws_dir=..., calibration_data_path="logged.jsonl").
  • OPE entry points raise migration errors that say exactly this.

Full version history in the CHANGELOG.

Development

git clone https://github.com/cimo-labs/cje.git
cd cje && poetry install && make test

Citation

If you use CJE in your research, please cite:

@misc{landesberg2025causaljudgeevaluationcalibrated,
  title={Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems},
  author={Eddie Landesberg and Manjari Narayan},
  year={2025},
  eprint={2512.11150},
  archivePrefix={arXiv},
  primaryClass={stat.ME},
  url={https://arxiv.org/abs/2512.11150},
}

License

MIT — See LICENSE for details.

Metadata

Release files for cje-eval 0.7.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cje-eval 0.7.2
File Size Uploaded
cje_eval-0.7.2.tar.gz 207.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cje-eval 0.7.2
File Interpreter ABI Platform
cje_eval-0.7.2-py3-none-any.whl Python 3 none any Details

Total release size: 433.0 kB

Release files / cje_eval-0.7.2.tar.gz

Download URL cje_eval-0.7.2.tar.gz
Size 207.3 kB
Tags Source
SHA-256 checksum
How to use checksums
f7a506f77b9608d10e18188f7fa867422cbc48ef8d124a3a6b6ca941de7a727f
BLAKE2b-256 checksum
How to use checksums
1678ae8a2571da45fa72d94851d3fd30f28be8666faf5c2ab4b2bf5cf12c49b4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log

Release files / cje_eval-0.7.2-py3-none-any.whl

Download URL cje_eval-0.7.2-py3-none-any.whl
Size 225.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
06891059a8d3ed87734ea4d32368e3e127e92b40bc479e846c322ec2e2418080
BLAKE2b-256 checksum
How to use checksums
7462187189b21332db54a52e5fcb0d1a8b14991db4ba6470465d60c564002499
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page