clef-evals
Calibration-first evaluation toolkit for Cloudflare Clef decision models. Judge cheap, audit confidence.
Most eval harnesses stop at accuracy. clef-evals also asks whether Clef's probabilities mean what they say. It computes Expected Calibration Error and Brier score over your own datasets, then turns both into a CI gate. A model that is right but overconfident fails your build before it fails your users.
Install
pip install clef-evals
export CLEF_ACCOUNT_ID=your_account_id
export CLEF_API_TOKEN=your_api_token
Requires Python 3.10+. The only runtime dependency is httpx.
Quick start
from clef_evals import ClefJudge
judge = ClefJudge() # config from environment
result = judge.evaluate([
{"state": "Email: I need a refund", "instructions": "Which team?",
"criteria": {"billing": "Payments, invoices, refunds",
"technical": "Bugs and outages",
"sales": "Plans and upgrades"},
"gold": "billing"},
{"state": "Checkout is down for everyone", "instructions": "Is this urgent?",
"gold": True}, # binary items use a boolean gold
])
print(result.summary())
samples=2 failures=0
accuracy=1.0000
ece=0.0700
brier=0.0049 brier_multiclass=0.0082
latency_ms p50=210.1 p95=238.6 p99=238.6
input_tokens=240 output_tokens=16
cost: $0.000058 total | $0.028800 per 1k calls
Async fan-out with a bounded semaphore:
import asyncio
from clef_evals import AsyncClefJudge
judge = AsyncClefJudge() # httpx.AsyncClient under the hood
result = asyncio.run(judge.evaluate(eval_set, concurrency=8))
Single decisions with the full probability distribution:
decision = judge.judge_choice(
"Email: charged twice", "Which team?",
{"billing": "Payments, invoices, refunds", "technical": "Bugs and outages"},
)
print(decision.choice, decision.probabilities) # billing {'billing': 0.93, ...}
p_yes = judge.judge_binary("Checkout is down", "Is this urgent?") # 0.97
Architecture
Animated version: docs/pipeline.svg · Showcase video: brag-output/brag.mp4 (rendered by brag-output/render_video.py, no stock assets) · Walkthrough notebook: research/clef_calibration_walkthrough.ipynb
┌─────────────────────────────────────────────────────────┐
│ your CI / your code │
└──────────┬─────────────────────────────────┬────────────┘
│ │
┌───────▼────────┐ ┌─────────▼─────────┐
│ ClefJudge │ │ clef-eval CLI │
│ sync + Async │ │ run / gate │
└───────┬────────┘ └─────────┬─────────┘
│ │
┌───────▼─────────────────────────────────▼─────────┐
│ ClefClient │
│ retries · exponential backoff + jitter │
│ timeouts · Retry-After · structured errors │
└───────┬───────────────────────────────────────────┘
│ HTTPS POST /accounts/{id}/ai/run/@cf/cloudflare/clef
┌───────▼───────────────────────────────────────────┐
│ Cloudflare Workers AI (clef 27B / clef-flash 9B) │
│ state + typed questions -> probabilities │
└───────┬───────────────────────────────────────────┘
│ per-option probabilities + usage
┌───────▼───────────────────────────────────────────┐
│ metrics │
│ accuracy · ECE · Brier · Brier-multiclass │
│ latency p50/p95/p99 · cost per 1k calls │
└───────┬───────────────────────────────────────────┘
│ EvalResult JSON
┌───────▼───────────────────────────────────────────┐
│ regression-gate GitHub Action │
│ fresh run vs committed baseline │
└───────────────────────────────────────────────────┘
CLI
# evaluate a dataset (JSON array or JSONL), human summary
clef-eval run evals/data/support_routing.jsonl
# machine-readable, save artifact
clef-eval run evals/data/support_routing.jsonl --json --output results/run.json
# CI gate: fail the build when quality or calibration regress
clef-eval run evals/data/support_routing.jsonl \
--min-accuracy 0.90 --max-ece 0.15
Exit codes: 0 gate passed · 1 gate failed · 2 config or dataset error.
Benchmarks
Published reference (Cloudflare's Decision Index 0.2.1)
Numbers below are Cloudflare's published measurements on their
infrastructure (model card,
blog), not measurements
made with this toolkit. Full table committed at
evals/results/published_reference.json.
| Benchmark | Clef | Clef-flash | Jev | Laya |
|---|---|---|---|---|
| BFCL · case exact | 98.5 | 98.8 | 95.8 | 38.1 |
| BANKING77 · macro-F1 | 94.2 | 90.9 | 79.7 | 14.3 |
| CLINC150+OOS · macro-F1 | 97.4 | 66.8 | 89.3 | 3.2 |
| When2Call · accuracy | 72.4 | 65.6 | 81.0 | 11.9 |
| ForecastBench · Brier (↓) | 13.9 | 10.6 | 17.4 | 41.1 |
| Median latency · ms | 209.3 | 38.8 | 524.1 | 5.8 |
| p95 latency · ms | 238.6 | 122.4 | 536.0 | 222.5 |
Our runs
| Dataset | Model | n | accuracy | ECE | Brier | p50 / p95 / p99 (ms) | $/1k calls |
|---|---|---|---|---|---|---|---|
| support_routing | clef | pending first live run | |||||
| support_routing | clef-flash | pending first live run |
Reproduce and add your numbers (needs CLEF_ACCOUNT_ID/CLEF_API_TOKEN):
make eval # both models, all datasets
python evals/run_eval.py --model @cf/cloudflare/clef-flash --concurrency 8
make test-integration # pytest against the real API
Cost model: published price is $0.24 per M input tokens
(e.g. ~120 input-token calls ≈ $0.029 per 1k calls). Output-token pricing
is not published by Cloudflare; output_tokens is reported but not priced.
clef vs laya
| Clef (Workers AI) | Clef-flash | Laya | |
|---|---|---|---|
| Type | 27B decision model (hosted) | 9B decision model (hosted) | decision model (open weights) |
| Context window | 65,536 tokens | 65,536 | 32,768 |
| Vision / images | yes | yes | no |
| Median latency | 209.3 ms | 38.8 ms | 5.8 ms |
| p95 latency | 238.6 ms | 122.4 ms | 222.5 ms |
| Quality (BFCL / BANKING77 / CLINC150) | 98.5 / 94.2 / 97.4 | 98.8 / 90.9 / 66.8 | 38.1 / 14.3 / 3.2 |
| Calibration (ForecastBench Brier, ↓) | 13.9 | 10.6 | 41.1 |
| Cost | $0.24 / M input tokens (hosted) | $0.24 / M input tokens | self-hosted (your GPUs) |
Reading: Laya wins raw latency. Clef wins quality and calibration by large margins, with clef-flash as the fast middle ground. For routing and gating workloads, miscalibrated confidence is what breaks automation. That is exactly what this toolkit measures on your data.
CI gate (reusable GitHub Action)
Commit a baseline JSON (any EvalResult.to_dict() output), then gate PRs:
- uses: jorgealizola/clef-evals/.github/actions/regression-gate@main
with:
current: results/run.json
baseline: results/baselines/support-routing-clef.json
metrics: |
accuracy:min:0.03
ece:max
latency_p95:max:50
accuracy:min = may not drop more than tolerance; ece:max = may not grow.
Pure Python at gate time: no credentials, no network.
Kaggle kernel
Reproduce the benchmark on Kaggle's free CPU runtime (verified push/status/output loop):
# add secrets CLEF_ACCOUNT_ID / CLEF_API_TOKEN on kaggle.com first
KAGGLE_API_TOKEN=... python -m kaggle kernels push -p kaggle-kernel
KAGGLE_API_TOKEN=... python -m kaggle kernels status gjusev/clef-evals-benchmark
KAGGLE_API_TOKEN=... python -m kaggle kernels output gjusev/clef-evals-benchmark -p out/
Error handling
Every failure is a typed exception under ClefError:
ClefError
├── ConfigurationError missing/invalid env (reports ALL problems at once)
├── ClefAPIError API refused the request
│ ├── ClefAuthError 401/403 (not retried)
│ ├── ClefRateLimitError 429 (retried, honors Retry-After)
│ └── ClefServerError 5xx (retried)
├── ClefResponseError body does not match the Clef schema
├── ClefTimeoutError retried
└── ClefNetworkError DNS / connection (retried)
Retries default to max_retries=2 with exponential backoff + jitter; every
error carries message and log-safe details. The library logs to the
clef_evals logger. It never prints and never logs your token.
Limitations (honest section)
- Calibration metrics audit, they don't fix. ECE/Brier tell you how much to trust Clef's probabilities on your distribution; they don't recalibrate them. Use the reported confidence accordingly (or calibrate downstream).
- Cost model covers input tokens only. Cloudflare publishes $0.24/M input
tokens but no output-token price for Clef at the time of writing.
output_tokensis reported so you can price it the day it appears. - Mixed-type datasets blend confidence semantics. Choice items use
P(chosen option); binary items usemax(p, 1−p). ECE/Brier over a mixed set pool both. Prefer per-type runs when the distinction matters. - Latency numbers are client-side (includes your network RTT to Cloudflare). Do not compare them 1:1 with Cloudflare's published infra-side medians.
- Fail-soft evaluation. Items that error after retries are excluded from
metrics and counted in
result.failures. The CLI gate fails on any failure, but direct library users should checkfailuresor risk silent drift. - Local inference is out of scope for most machines. Clef is a 27B model with a custom joint-schema head (reference hardware: a single H200; weights ~55 GB fp16). No GGUF/vLLM-quantized path is published. These benchmarks target the hosted Workers AI API.
- v0.x API. Expect small breaking changes before 1.0; the v0.1 names
ClefEvalResult,ece,brier_scoreremain importable.
Development
make install # editable install with dev extras
make test # pytest with coverage (>90% enforced); integration tests excluded
make test-integration # real-API tests: needs CLEF_ACCOUNT_ID / CLEF_API_TOKEN
make lint # ruff
make build # wheel + sdist
Project layout: src/clef_evals/ (config, client, models, metrics, judge, cli) ·
tests/ (unit + fixtures with real API shapes) · evals/ (datasets, runner,
committed results) · scripts/check_regression.py + .github/actions/regression-gate/
(CI gate) · kaggle-kernel/ (cloud reproduction) · research/ (calibration
walkthrough notebook) · docs/ + brag-output/ (visuals, video renderer).
License
Apache 2.0, see LICENSE. Clef models are Apache 2.0 on Hugging Face.
Metadata
Release files for clef-evals 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| clef_evals-0.2.0.tar.gz | 246.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| clef_evals-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 275.9 kB
Release files / clef_evals-0.2.0.tar.gz
| Download URL | clef_evals-0.2.0.tar.gz |
|---|---|
| Size | 246.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6e7493e312d727ac87da47022aa5a50a9a4917459fc09f76f91dc4f0096be6cb
|
|
BLAKE2b-256 checksum How to use checksums |
d519ebac358e82a09096b78931474c382e2b4fd60cccdf36c34bce995ab44bfc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / clef_evals-0.2.0-py3-none-any.whl
| Download URL | clef_evals-0.2.0-py3-none-any.whl |
|---|---|
| Size | 29.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9824bc466220594f7126f09cafceabb5ce9758d0d72d713e424e534103b43397
|
|
BLAKE2b-256 checksum How to use checksums |
0a7c8197c4d0b40a3917ef3355bc838f4d77ab05ee45aeb87ab0861ce3aec890
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|