Skip to main content

clef-router sends each prompt to the cheap or the frontier model

Route prompts with Cloudflare's Clef decision model.
Send each prompt to the affordable model when it is enough, and to the frontier model when it matters.

CI PyPI Python 3.10+ Apache 2.0 license

Get started · How it works · Benchmarks · clef vs laya · Limitations


Quickstart

pip install clef-router
export CLEF_ACCOUNT_ID=your_cloudflare_account_id
export CLEF_API_TOKEN=your_cloudflare_api_token

Run it as a proxy and point any OpenAI client at it:

clef-router --port 8000
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")

completion = client.chat.completions.create(
    model="auto",  # the router chooses; your value is ignored
    messages=[{"role": "user", "content": "Say hi in three words"}],
)
print(completion.choices[0].message.content)  # the routing decision as JSON
print(completion.extensions["clef"]["tier"])  # "cheap" or "frontier"

The proxy is a decision service. It returns the routing decision (tier, confidence, reason, token usage) in an OpenAI-shaped response, and your client calls the chosen model, exactly like laya-router's decision headers.

Prefer in-process? The library and the CLI do the same thing without HTTP:

from clef_router import ClefRouter

with ClefRouter() as router:
    routing = router.route("Design a rate limiter")
print(routing.tier)            # "frontier"
print(routing.reason)          # "team selected frontier"
print(routing.decision.usage.input_tokens)
clef-route "Design a rate limiter"
# tier:       frontier
# reason:     team selected frontier
# model:      clef-flash

Async clients, custom question sets, and raw decisions are one import away:

from clef_router import AsyncClefRouter

async with AsyncClefRouter() as router:
    decision = await router.decide({"invoice_total": 1250, "due_days": 45})

How it works

   your app (OpenAI SDK)          your app (library / CLI)
          |                               |
          v                               v
   +---------------------+      +----------------+
   | clef-router proxy   |      | clef_router    |
   | POST /v1/chat/...   |      | ClefRouter     |
   | POST /v1/decide     |      |                |
   +----------+----------+      +-------+--------+
              |                          |
              +------------+-------------+
                           v
        +--------------------------------------+
        |  Clef decision model  (one call)     |
        |  questions: team, urgency            |
        |  answers:   probabilities per option |
        +------------------+-------------------+
                           |
        team=cheap & conf>=0.45 --> CHEAP tier
        otherwise (low confidence, urgent,
        unknown answer)----------> FRONTIER tier

One forward pass scores every option of every question, so a decision costs a single small request. The policy is fail-safe on purpose:

  1. When the team answer names a tier, that tier wins, unless its confidence is below min_confidence (default 0.45), which escalates to frontier.
  2. With no usable team answer, the urgency probability routes: 0.5 or higher goes frontier.
  3. Anything the router cannot parse escalates. It never silently picks the cheap tier without evidence.
Animated explainer

Open docs/how-it-works.html in a browser. Static version: docs/how-it-works.svg.

Benchmarks

Two different things get measured, and mixing them up is the oldest trick in routing marketing. We keep them separate.

What we measured (this repo, evals/run_eval.py, committed results):

Metric Result What it is
Policy accuracy 100% (n=42) Does the tier policy map labeled decisions to the right tier. Fixtures committed in evals/data/.
Escalation precision 1.00 Of the frontier escalations, how many were justified. 0 over-escalations, 0 under-routes.
Decision overhead (library) p50 0.12 ms / p95 0.26 ms / p99 0.47 ms Client-side cost of parsing and policy. Excludes the network call.
Cost per 1k decisions $0.078 measured ($0.033 estimated from fixtures) Real mean is 324 input tokens per decision at Cloudflare's published $0.24 per million input tokens. The fixture estimate undershot; the GPU rerun below measured the real token counts.
Kaggle rerun (Linux, clean box) accuracy 1.00, p50 0.34 ms Same dataset and code, executed by the public kernel gjusev/clef-router-evals; log and JSON committed in evals/results/.
Real model, local weights (Kaggle T4) accuracy 92.9% (39/42), warm p50 506 ms clef-flash weights in 4-bit on one T4, same prompts and policy, no Cloudflare API. The 3 misses: two genuinely ambiguous prompts routed cheap ("write a sonnet", "what about the other approach") and one safe over-escalation of a Python question. Full trace in evals/results/routing-gpu-eval.json.

What Cloudflare measured (Decision Index 0.2.1, from the model card; that suite scores the decision model itself, not this router):

Benchmark Clef Clef-flash Laya
BFCL (accuracy) 98.5 98.8 38.1
MMLU (accuracy) 90.3 91.8 30.7
Median latency (ms) 209.3 38.8 5.8
p95 latency (ms) 238.6 122.4 222.5

Reproduce our numbers offline (no credentials, seconds to run):

python evals/run_eval.py                          # writes evals/results/routing-eval-mock.json
python evals/run_eval.py --mode api               # real end-to-end numbers; needs credentials

The repository ships a regression gate so a routing change cannot silently drop accuracy:

- uses: ./.github/actions/regression-gate
  with:
    current: evals/results/current.json
    baseline: evals/results/baselines/routing-eval-mock.json
    metrics: |
      metrics.accuracy:max
      metrics.escalations.escalation_precision:max:0.05

clef vs laya

laya-router clef-router
Decision quality (BFCL) 38.1 98.8 (clef-flash)
Median decision latency 5.8 ms 38.8 ms (hosted flash) / 209.3 ms (full)
Decision overhead (local library) n/a, proxy-local 0.12 ms
Hosting local CPU Cloudflare Workers AI
Routing cost $0 ~$0.03 per 1k decisions
Context 32k 64k tokens
Answer production proxy swaps the model upstream returns the decision; you call the model

Honest reading: laya wins on latency because it runs on your machine. clef-router wins on decision quality, which shows up as fewer wrong routes on genuinely hard prompts. If your workload is latency-sensitive at the routing step and mostly easy prompts, laya is a great fit. If misroutes cost more than 30 ms, use clef-router.

Context compaction

clef_router.compactor uses the same one-pass trick for RAG: one Clef request scores the relevance of up to 64 retrieved documents, then a token budget decides what survives. Documents are kept verbatim or cut with a reason; nothing is rewritten.

from clef_router.compactor import compact

result = compact(
    "invoice processing rules",
    documents=retrieved_chunks,
    budget=2048,
)
print(result.stats.savings_pct)   # e.g. 61.3
print(result.kept[0].score)       # 0..3 relevance
print(result.cut[2].reason)       # "score 0 below min_score 1"

Framework adapters ship as extras (pip install "clef-router[compactor]"):

from clef_router.compactor.langchain import ClefDocumentCompressor   # LangChain
from clef_router.compactor.llamaindex import ClefNodePostprocessor   # LlamaIndex

Configuration

Environment variables, with CLEF_* winning over CLOUDFLARE_* fallbacks:

Variable Default Purpose
CLEF_ACCOUNT_ID (required) Cloudflare account. Falls back to CLOUDFLARE_ACCOUNT_ID.
CLEF_API_TOKEN (required) API token with Workers AI permissions. Falls back to CLOUDFLARE_API_TOKEN.
CLEF_MODEL clef-flash clef or clef-flash.
CLEF_MIN_CONFIDENCE style options 0.45 Set min_confidence in code; 0 disables the gate.
CLEF_TIMEOUT 30 Per-request timeout in seconds.
CLEF_MAX_RETRIES 2 Retry attempts for timeouts, network errors, 429 and transient 5xx, with exponential backoff and jitter; honors Retry-After.
CLEF_LOG_LEVEL INFO Standard Python level names.
CLEF_HOST / CLEF_PORT 127.0.0.1:8000 Proxy bind address.

Retries are structured: timeouts and connection errors retry, HTTP 401/403 raise ClefAuthError immediately, 429 raises ClefRateLimitError with the server's Retry-After once the budget is spent. Every HTTP-derived error carries the status code, the Cloudflare error code, and the cf-ray request id when present.

Operate the proxy

docker build -t clef-router .
docker run -p 8000:8000 -e CLEF_ACCOUNT_ID=... -e CLEF_API_TOKEN=... clef-router
  • "stream": true works: the decision arrives as OpenAI-shaped SSE chunks (role, decision JSON, stop, [DONE]), and the clef extension rides in the first chunk.
  • GET /metrics serves Prometheus text: clef_router_requests_total by tier and status, plus a clef_router_routing_seconds histogram.
  • CLEF_DECISION_LOG=/path/decisions.jsonl appends one JSON line per decision (tier, confidence, reason, latency, token usage, prompt preview) for post-hoc calibration of the confidence gate.

Limitations

Stated plainly, because routing libraries that hide these waste your time.

  • The proxy decides; it does not complete. It never forwards your prompt to a chat model. Clients that want the answer call the chosen model themselves.
  • The committed policy number is accuracy on 42 labeled fixtures. The GPU rerun measures the real model on the same 42 prompts (92.9%), which is still a small set: neither number is a Decision-Index-grade benchmark. Use --mode api with real credentials to measure your own workload.
  • The tier mapping reads the team and urgency question ids. Custom question sets work with decide(); route() still expects those ids.
  • Latency depends on your network to Cloudflare. The 38.8 ms median is Cloudflare's measurement; add your round trip.
  • Images are accepted by the API (decide(), max 4) but the routing question set is text-only today.
  • Running Cloudflare/clef-flash weights locally is possible (Apache-2.0, 9.4B) and is exactly what evals/kaggle-kernel-gpu/ does on one T4 in 4-bit; expect ~500 ms per decision there versus 38.8 ms median on Cloudflare's unquantized datacenter GPUs.

Reproduce on Kaggle

Two public kernels rerun the evaluation on Kaggle hardware:

  • evals/kaggle-kernel-gpu/ — GPU (T4): downloads the real clef-flash weights, routes every labeled prompt through the model, and writes routing-gpu-eval.json. Needs a GPU-verified account; about 15 minutes and a slice of your 30 h weekly quota.
  • evals/kaggle-kernel/ — CPU: replays the committed fixtures through the policy pipeline in about two minutes, no GPU or credentials.
KAGGLE_API_TOKEN=... python -m kaggle kernels push -p evals/kaggle-kernel-gpu

Development

git clone https://github.com/Gjusev/clef-router.git && cd clef-router
pip install -e ".[dev]"
make test          # offline suite, integration deselected
make lint          # ruff
make eval          # regenerate evals/results/routing-eval-mock.json
make assets        # regenerate logo, social card, diagrams, demo video
pytest -m integration   # live API tests; needs credentials
src/clef_router/     client (sync+async), config, errors, models, transport
src/clef_router/     server (FastAPI proxy), compat (in-process OpenAI)
src/clef_router/     compactor (core + LangChain/LlamaIndex adapters)
evals/               dataset, runner, committed results, Kaggle kernel
scripts/             asset generation, regression check CLI
tests/               offline suite; @pytest.mark.integration opts into the live API

License

Apache-2.0. See LICENSE. Clef is a Cloudflare model; its weights carry the same license on the model page.

Metadata

Release files for clef-router 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for clef-router 0.3.0
File Size Uploaded
clef_router-0.3.0.tar.gz 438.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for clef-router 0.3.0
File Interpreter ABI Platform
clef_router-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 489.2 kB

Release files / clef_router-0.3.0.tar.gz

Download URL clef_router-0.3.0.tar.gz
Size 438.9 kB
Tags Source
SHA-256 checksum
How to use checksums
5e5d755f03db1606377d2d41f7aeb70101c9ea9c914a9d588dc4e0fa36100d3f
BLAKE2b-256 checksum
How to use checksums
bc89a936a2fdf1ab6ada14c6a61059e94e7bda0823524f81148282126bbc5437
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / clef_router-0.3.0-py3-none-any.whl

Download URL clef_router-0.3.0-py3-none-any.whl
Size 50.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2c527a1e85529cb5c7f20a076c97cbc9c32fba5ba6735d1c00506b75530f269f
BLAKE2b-256 checksum
How to use checksums
846e1e0159b769d5363678aa5fd2cc16d1f377c238d1ec6fb0090211542efc72
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page