Route prompts with Cloudflare's Clef decision model.
Send each prompt to the affordable model when it is enough, and to the frontier model when it matters.
Get started · How it works · Benchmarks · clef vs laya · Limitations
Quickstart
pip install clef-router
export CLEF_ACCOUNT_ID=your_cloudflare_account_id
export CLEF_API_TOKEN=your_cloudflare_api_token
Run it as a proxy and point any OpenAI client at it:
clef-router --port 8000
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
completion = client.chat.completions.create(
model="auto", # the router chooses; your value is ignored
messages=[{"role": "user", "content": "Say hi in three words"}],
)
print(completion.choices[0].message.content) # the routing decision as JSON
print(completion.extensions["clef"]["tier"]) # "cheap" or "frontier"
The proxy is a decision service. It returns the routing decision (tier, confidence, reason, token usage) in an OpenAI-shaped response, and your client calls the chosen model, exactly like laya-router's decision headers.
Prefer in-process? The library and the CLI do the same thing without HTTP:
from clef_router import ClefRouter
with ClefRouter() as router:
routing = router.route("Design a rate limiter")
print(routing.tier) # "frontier"
print(routing.reason) # "team selected frontier"
print(routing.decision.usage.input_tokens)
clef-route "Design a rate limiter"
# tier: frontier
# reason: team selected frontier
# model: clef-flash
Async clients, custom question sets, and raw decisions are one import away:
from clef_router import AsyncClefRouter
async with AsyncClefRouter() as router:
decision = await router.decide({"invoice_total": 1250, "due_days": 45})
How it works
your app (OpenAI SDK) your app (library / CLI)
| |
v v
+---------------------+ +----------------+
| clef-router proxy | | clef_router |
| POST /v1/chat/... | | ClefRouter |
| POST /v1/decide | | |
+----------+----------+ +-------+--------+
| |
+------------+-------------+
v
+--------------------------------------+
| Clef decision model (one call) |
| questions: team, urgency |
| answers: probabilities per option |
+------------------+-------------------+
|
team=cheap & conf>=0.45 --> CHEAP tier
otherwise (low confidence, urgent,
unknown answer)----------> FRONTIER tier
One forward pass scores every option of every question, so a decision costs a single small request. The policy is fail-safe on purpose:
- When the
teamanswer names a tier, that tier wins, unless its confidence is belowmin_confidence(default 0.45), which escalates to frontier. - With no usable
teamanswer, theurgencyprobability routes: 0.5 or higher goes frontier. - Anything the router cannot parse escalates. It never silently picks the cheap tier without evidence.
Animated explainer
Open docs/how-it-works.html in a browser. Static version: docs/how-it-works.svg.
Benchmarks
Two different things get measured, and mixing them up is the oldest trick in routing marketing. We keep them separate.
What we measured (this repo, evals/run_eval.py, committed results):
| Metric | Result | What it is |
|---|---|---|
| Policy accuracy | 100% (n=42) | Does the tier policy map labeled decisions to the right tier. Fixtures committed in evals/data/. |
| Escalation precision | 1.00 | Of the frontier escalations, how many were justified. 0 over-escalations, 0 under-routes. |
| Decision overhead (library) | p50 0.12 ms / p95 0.26 ms / p99 0.47 ms | Client-side cost of parsing and policy. Excludes the network call. |
| Cost per 1k decisions | $0.033 | Mean ~137 input tokens at Cloudflare's published $0.24 per million input tokens. |
| Kaggle rerun (Linux, clean box) | accuracy 1.00, p50 0.34 ms | Same dataset and code, executed by the public kernel gjusev/clef-router-evals; log and JSON committed in evals/results/. |
What Cloudflare measured (Decision Index 0.2.1, from the model card; that suite scores the decision model itself, not this router):
| Benchmark | Clef | Clef-flash | Laya |
|---|---|---|---|
| BFCL (accuracy) | 98.5 | 98.8 | 38.1 |
| MMLU (accuracy) | 90.3 | 91.8 | 30.7 |
| Median latency (ms) | 209.3 | 38.8 | 5.8 |
| p95 latency (ms) | 238.6 | 122.4 | 222.5 |
Reproduce our numbers offline (no credentials, seconds to run):
python evals/run_eval.py # writes evals/results/routing-eval-mock.json
python evals/run_eval.py --mode api # real end-to-end numbers; needs credentials
The repository ships a regression gate so a routing change cannot silently drop accuracy:
- uses: ./.github/actions/regression-gate
with:
current: evals/results/current.json
baseline: evals/results/baselines/routing-eval-mock.json
metrics: |
metrics.accuracy:max
metrics.escalations.escalation_precision:max:0.05
clef vs laya
| laya-router | clef-router | |
|---|---|---|
| Decision quality (BFCL) | 38.1 | 98.8 (clef-flash) |
| Median decision latency | 5.8 ms | 38.8 ms (hosted flash) / 209.3 ms (full) |
| Decision overhead (local library) | n/a, proxy-local | 0.12 ms |
| Hosting | local CPU | Cloudflare Workers AI |
| Routing cost | $0 | ~$0.03 per 1k decisions |
| Context | 32k | 64k tokens |
| Answer production | proxy swaps the model upstream | returns the decision; you call the model |
Honest reading: laya wins on latency because it runs on your machine. clef-router wins on decision quality, which shows up as fewer wrong routes on genuinely hard prompts. If your workload is latency-sensitive at the routing step and mostly easy prompts, laya is a great fit. If misroutes cost more than 30 ms, use clef-router.
Context compaction
clef_router.compactor uses the same one-pass trick for RAG: one Clef
request scores the relevance of up to 64 retrieved documents, then a token
budget decides what survives. Documents are kept verbatim or cut with a
reason; nothing is rewritten.
from clef_router.compactor import compact
result = compact(
"invoice processing rules",
documents=retrieved_chunks,
budget=2048,
)
print(result.stats.savings_pct) # e.g. 61.3
print(result.kept[0].score) # 0..3 relevance
print(result.cut[2].reason) # "score 0 below min_score 1"
Framework adapters ship as extras (pip install "clef-router[compactor]"):
from clef_router.compactor.langchain import ClefDocumentCompressor # LangChain
from clef_router.compactor.llamaindex import ClefNodePostprocessor # LlamaIndex
Configuration
Environment variables, with CLEF_* winning over CLOUDFLARE_* fallbacks:
| Variable | Default | Purpose |
|---|---|---|
CLEF_ACCOUNT_ID |
(required) | Cloudflare account. Falls back to CLOUDFLARE_ACCOUNT_ID. |
CLEF_API_TOKEN |
(required) | API token with Workers AI permissions. Falls back to CLOUDFLARE_API_TOKEN. |
CLEF_MODEL |
clef-flash |
clef or clef-flash. |
CLEF_MIN_CONFIDENCE style options |
0.45 |
Set min_confidence in code; 0 disables the gate. |
CLEF_TIMEOUT |
30 |
Per-request timeout in seconds. |
CLEF_MAX_RETRIES |
2 |
Retry attempts for timeouts, network errors, 429 and transient 5xx, with exponential backoff and jitter; honors Retry-After. |
CLEF_LOG_LEVEL |
INFO |
Standard Python level names. |
CLEF_HOST / CLEF_PORT |
127.0.0.1:8000 |
Proxy bind address. |
Retries are structured: timeouts and connection errors retry, HTTP 401/403
raise ClefAuthError immediately, 429 raises ClefRateLimitError with the
server's Retry-After once the budget is spent. Every HTTP-derived error
carries the status code, the Cloudflare error code, and the cf-ray request
id when present.
Limitations
Stated plainly, because routing libraries that hide these waste your time.
- The proxy decides; it does not complete. It never forwards your prompt to a chat model. Clients that want the answer call the chosen model themselves.
- No streaming.
POST /v1/chat/completionsis request/response. SSE pass-through is on the roadmap. - The committed accuracy number is policy accuracy on 42 labeled fixtures,
not a model benchmark. Use
--mode apiwith real credentials for end-to-end numbers on your own workload, or run the Decision Index suite for model quality. - The tier mapping reads the
teamandurgencyquestion ids. Custom question sets work withdecide();route()still expects those ids. - Latency depends on your network to Cloudflare. The 38.8 ms median is Cloudflare's measurement; add your round trip.
- Images are accepted by the API (
decide(), max 4) but the routing question set is text-only today. - Running
Cloudflare/clefweights locally is possible (Apache-2.0) but the model card lists a single H200 as the tested environment; the Kaggle kernel inevals/kaggle-kernel/measures policy overhead on CPU instead.
Reproduce on Kaggle
evals/kaggle-kernel/ holds a kernel that clones this repository and reruns
the offline evaluation on clean Kaggle hardware, writing
/kaggle/working/routing-eval.json:
KAGGLE_API_TOKEN=... python -m kaggle kernels push -p evals/kaggle-kernel
Development
git clone https://github.com/Gjusev/clef-router.git && cd clef-router
pip install -e ".[dev]"
make test # offline suite, integration deselected
make lint # ruff
make eval # regenerate evals/results/routing-eval-mock.json
make assets # regenerate logo, social card, diagrams, demo video
pytest -m integration # live API tests; needs credentials
src/clef_router/ client (sync+async), config, errors, models, transport
src/clef_router/ server (FastAPI proxy), compat (in-process OpenAI)
src/clef_router/ compactor (core + LangChain/LlamaIndex adapters)
evals/ dataset, runner, committed results, Kaggle kernel
scripts/ asset generation, regression check CLI
tests/ offline suite; @pytest.mark.integration opts into the live API
License
Apache-2.0. See LICENSE. Clef is a Cloudflare model; its weights carry the same license on the model page.
Metadata
Release files for clef-router 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| clef_router-0.2.0.tar.gz | 423.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| clef_router-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 470.3 kB
Release files / clef_router-0.2.0.tar.gz
| Download URL | clef_router-0.2.0.tar.gz |
|---|---|
| Size | 423.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2d3955648a3ec6e77e33150a8a2896ff587a66d062642ebaba30448e86760535
|
|
BLAKE2b-256 checksum How to use checksums |
74976f42047f108bf66dc19d428641a89e5675c765a71d9e72f881c733dea8ff
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / clef_router-0.2.0-py3-none-any.whl
| Download URL | clef_router-0.2.0-py3-none-any.whl |
|---|---|
| Size | 46.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cf24d2e8c6be672017f854463c5640e22b0068327c4d0d086261754e37e05290
|
|
BLAKE2b-256 checksum How to use checksums |
2bc385bf3074dd82739849d3e98d4c2f4f159aee0a3b970007ff05d362097c9f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|