🚦 CIGate
Gate your CI/CD on the confidence interval, not the vibes.
Eval-gated CI/CD for AI products: block a merge when answer quality statistically regresses — per failure-mode, with the LLM-judge's bias corrected for, under a cost budget.
The problem (a real, expensive one)
Teams shipping LLM and agent products change prompts, retrieval, tools, and model versions constantly — and those changes have non-local effects: fixing one answer silently breaks ten others. The industry default is "vibes-based" shipping, and it produces silent quality regressions that reach production and cost real money.
The system-design case study this implements opens with the canonical disaster: a well-meaning prompt tweak quietly degraded answer quality on contract questions, an LLM-judge metric drifted 11 points undetected, and the company lost a $4M renewal before anyone noticed.
The gap nobody fills
The popular eval tools — Promptfoo, Braintrust, Langfuse, DeepEval — gate on raw LLM-judge scores. But in-domain LLM-judge accuracy is only ~75–88%, so the judge's observed pass rate is biased. Gating on it means you either:
- over-block (false alarms → developers route around the gate), or
- under-block (real regressions still ship).
CIGate gates on the bias-corrected pass rate's confidence-interval lower bound, per failure-mode axis. That's the part the case study makes its centerpiece and the mainstream tools skip. (The "CI" pun is the whole thesis: gate your Continuous Integration on a Confidence Interval.)
What it looks like on a PR
A one-line prompt change ships answer_v2 ("be helpful and complete, citations
optional"). CIGate runs on the PR and posts this — then blocks the merge:
❌ CIGate: merge blocked — 2 axis regression(s)
prompt=answer_v2·judge=mock·sample=60/300·cost=$0.00
Axis Raw judge Corrected 95% CI Baseline Δ Verdict 🔴 hallucination45.0% 37.0% [11.1%, 59.0%] 98.9% −61.9 pp REGRESSED 🟢 retrieval_miss85.0% 91.1% [69.8%, 100%] 100% −8.9 pp ok 🔴 citation_error61.7% 65.5% [47.6%, 82.3%] 100% −34.5 pp REGRESSED 🟢 refusal71.7% 76.5% [55.9%, 91.9%] 93.2% −16.7 pp ok 🟢 format_violation100% 95.0% [88.2%, 100%] 98.9% −3.9 pp ok
The regression is isolated to the two axes the change actually hurt — a single
composite score would have hidden it. A clean change goes green and merges. Full samples:
docs/samples/.
How it works
flowchart LR
PR[Pull request] --> RUN[Run SUT over a<br/>stratified golden-set sample]
RUN --> CODE[Code evaluators<br/>schema · citations · retrieval]
RUN --> JUDGE[LLM judge<br/>hallucination · refusal · ...]
CODE --> DET[Per-axis detector]
JUDGE --> DET
CAL[Human-labeled<br/>calibration set] --> CORR
DET --> CORR[Statistical correction<br/>Rogan–Gladen + CI]
CORR --> GATE{Drop vs main<br/>baseline > tolerance?}
GATE -- yes --> BLOCK[🔴 block + per-axis report]
GATE -- no --> PASS[🟢 merge allowed]
- Sample the golden set, stratified so every failure-mode axis is represented (cost control — a per-PR run touches a fraction, nightly runs the full set).
- Score each case two ways: cheap deterministic code checks (citations, schema, retrieval) and an LLM-as-judge for subjective axes.
- Correct the judge's bias: using its sensitivity/specificity measured on a human-labeled calibration set, recover the true pass rate with a confidence interval. Deterministic axes skip correction (they're unbiased) and use an exact binomial interval.
- Gate per axis with a one-sided two-sample drop test vs the committed
mainbaseline (Bonferroni-corrected across axes): block only when we're confident the drop exceeds tolerance. Identical builds never false-block, regardless of CI width.
The statistical core (the part that matters)
Raw judge pass rate p_obs is biased. With judge sensitivity (TPR) and specificity
(TNR) measured on a labeled calibration set, the Rogan–Gladen estimator recovers the
true rate:
p_obs + TNR − 1
θ̂ = clip( ───────────────── , 0, 1 )
TPR + TNR − 1
The confidence interval uses the adjusted-Wald delta method (Lee, Zeng et al.,
arXiv:2511.21140), combining all three uncertainty
sources — evaluated-sample, sensitivity, and specificity — with correct ~95% coverage
even on small calibration sets. We cross-check it against the
judgy library in the test suite, and the
implementation reproduces the paper's worked example exactly. See
docs/METHODOLOGY.md.
If the judge is no better than chance (TPR + TNR ≤ 1) or the CI is too wide, CIGate
refuses to gate that axis rather than guess.
Try it in 60 seconds ($0, offline)
Everything runs in a deterministic mock mode — no API key, no spend — which is also what powers the test suite and the demo CI.
git clone https://github.com/awesome-pro/cigate && cd cigate
pip install -e ".[dev]"
bash scripts/try_local.sh # one command: blocks a bad build, passes a good one
Or run the steps yourself:
cigate baseline --promote # establish a 'good' baseline (full run)
BUILD_FLAVOR=regressed cigate gate # → blocks: hallucination + citation_error red
BUILD_FLAVOR=good cigate gate # → passes: all axes within tolerance
pytest -q # 28 tests, all green, $0
Want the real thing? CIGate runs cross-provider so no model grades its own output — the product runs on OpenAI, the judge is Claude:
cp .env.example .env # add OPENAI_API_KEY + ANTHROPIC_API_KEY
bash scripts/real_eval.sh # ~2 min, live progress → docs/results/
→ See docs/RESULTS.md for a real run: the Claude judge calibrated
against 166 CUAD expert labels, a real GPT regression blocked for $0.39 (citation
pass-rate collapsed 0.94 → 0.24 corrected), a safe change passed for $0.32.
Two datasets: synthetic + real
synthetic_contract— a generated contract/insurance support set (300 cases, 50 policy docs). Fully controlled; powers the deterministic demo and tests.cuad_real— built from CUAD (real commercial contracts with expert clause annotations, CC BY 4.0). Its headline use: the LLM judge is calibrated against real human expert labels, so the correction's confusion matrix is measured from real bias, not assumed. (evalconfig_cuad.yaml)
Use it on your own product
CIGate is product-agnostic. Point evalconfig.yaml at any callable
(question) -> SUTOutput and bring your own golden set:
sut: "yourapp.bot:answer" # module:callable
goldenset: "goldensets/yours.yaml"
axes: [hallucination, citation_error, ...]
gate: { tolerance: 0.02, confidence_level: 0.95 }
Then drop the GitHub Action into your pipeline (see .github/):
- uses: awesome-pro/cigate@v0.1
with:
config: evalconfig.yaml
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }} # omit -> $0 mock mode
What's in here
| Path | What |
|---|---|
src/cigate/stats.py |
Rogan–Gladen + adjusted-Wald CI — the correction core |
src/cigate/gate.py |
per-axis two-sample drop test vs baseline |
src/cigate/{runner,evaluators,calibrate}.py |
eval execution, code+judge scoring, drift |
src/refbot/ |
the demo RAG bot (BM25 + Claude/mock generator) |
.github/ |
composite Action + PR / nightly workflows |
dashboard/app.py |
Streamlit dashboard (per-axis, calibration, live gate) |
docs/ |
architecture, methodology, auditor pack, article, demo script |
More
- 🏗 Architecture · 📐 Methodology · 🧾 Auditor pack sample · 🎬 Demo script
- Built from the "Eval-Gated CI/CD" system-design case study. Grounded in the evals work of Hamel Husain, Shreya Shankar, and Eugene Yan.
License
MIT. CUAD data under CC BY 4.0 — see data/cuad/ATTRIBUTION.md.
Metadata
Release files for cigate 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cigate-0.2.0.tar.gz | 1.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cigate-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.2 MB
Release files / cigate-0.2.0.tar.gz
| Download URL | cigate-0.2.0.tar.gz |
|---|---|
| Size | 1.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f363906a9c90812aef55a1101d8edd767fd7276b11512f511dc5fe2428cf4e2b
|
|
BLAKE2b-256 checksum How to use checksums |
620ce56a006d335f0bcad532157c636498c87cc7580f1176c5a5f572bc6490e7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 19, 2026.
Transparency logRelease files / cigate-0.2.0-py3-none-any.whl
| Download URL | cigate-0.2.0-py3-none-any.whl |
|---|---|
| Size | 69.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9c01d74bef2be05394ab2084bfef03355da24783d7cc6937b6b2d3669893f60a
|
|
BLAKE2b-256 checksum How to use checksums |
9b334fe757f3c1e3ba04d1f12ee345382a62292e1427ffe8f3589bc8ff326be8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 19, 2026.
Transparency log