Skip to main content
clef-compactor logo: retrieved chunks pass through a relevance gate and only the useful ones continue

clef-compactor

Keep the evidence. Cut the noise.

Query-aware RAG context compaction using Cloudflare's Clef. Score a whole retrieval batch against the query in one API call, keep the chunks worth sending to your LLM.

CI PyPI Python 3.10+ License Types

Quick start · How it works · Measured results · clef vs laya · Limitations

clef-compactor demo: five retrieved chunks are scored, the irrelevant one is cut with a reason, and the kept three form the context with 34% fewer tokens

▶ Watch the 12-second demo · landing page · animated diagram

Your retriever returns 12 chunks. Your LLM reads all of them and you pay for all of them, including the ones about the Eiffel Tower when the user asked about refunds. clef-compactor asks Clef one question per chunk ("is this needed to answer the query?"), then fills a token budget with the best answers.

Chunks are kept verbatim or removed with a recorded reason. The library never rewrites text, so citations stay auditable.

How it works

                       clef-compactor
                       ───────────────
 query ──────────────► │             │
                       │  build      │
 chunks ─────────────► │  1 noul     │        ┌─────────────────────────┐
 [c1][c2]...[cN]       │  question   │──────► │  Cloudflare Clef API    │
                       │  per chunk  │        │  P(relevant) per chunk  │
                       │  (≤64/req)  │◄───────┴─────────────────────────┘
                       │             │
                       │  rank:      │
                       │  score desc │        c1 P=0.93  ──► keep
                       │  tokens asc │        c3 P=0.85  ──► keep
                       │  index asc  │        c5 P=0.42  ──► cut (budget)
                       │             │        c2 P=0.05  ──► cut (irrelevant)
                       │  fill       │
                       │  budget     │──────► CompactResult
                       └─────────────┘        kept / dropped / scores
                                              tokens, usage, cost estimate

One API call per 64 chunks. Dropped chunks carry a reason: irrelevant (below the threshold) or budget_exhausted (did not fit). Both are dataclasses, so you can log them or show them to a user.

Quick start

pip install clef-compactor
export CLEF_ACCOUNT_ID=your_account_id
export CLEF_API_TOKEN=your_api_token
from clef_compactor import ClefCompactor

compactor = ClefCompactor()
result = compactor.compact(
    "What is the refund policy?",
    retrieved_chunks,
    token_budget=1000,
)

result.kept_texts()          # surviving chunks, ranked best first, text unchanged
result.dropped[0].drop_reason  # DropReason.IRRELEVANT or BUDGET_EXHAUSTED
result.saved_fraction        # e.g. 0.71 -> 71% fewer context tokens
result.cost_estimate         # USD, input tokens at Cloudflare's $0.24/M

Async

from clef_compactor import AsyncClefCompactor

compactor = AsyncClefCompactor(model="clef-flash")   # faster, slightly less precise
result = await compactor.compact(query, chunks, token_budget=800)
await compactor.aclose()

CLI

clef-compact -q "refund policy" -d "chunk one" "chunk two" --budget 1000
clef-compact -q "..." -d @chunk1.txt @chunk2.txt --json --model clef-flash

A runnable walkthrough lives in examples/demo.ipynb; it executes offline against a canned API.

OpenAI-compatible endpoint

Any OpenAI SDK can talk to a compactor. handle_chat_completions takes an OpenAI request body and returns an OpenAI response:

from fastapi import FastAPI
from clef_compactor import ClefCompactor
from clef_compactor.compat.openai import handle_chat_completions

app = FastAPI()
compactor = ClefCompactor()

@app.post("/v1/chat/completions")
async def chat_completions(payload: dict):
    return handle_chat_completions(payload, compactor=compactor)

Request shape: put the chunks in clef.chunks, or pass context as non-user messages. The response carries the compacted context in choices[0].message.content and before/after token counts in usage.

LangChain

pip install "clef-compactor[langchain]"
from clef_compactor.integrations.langchain import ClefDocumentCompressor

compressor = ClefDocumentCompressor(token_budget=800)
compressed = compressor.compress_documents(docs, query="refund policy?")

LlamaIndex

pip install "clef-compactor[llamaindex]"
from clef_compactor.integrations.llamaindex import ClefNodePostprocessor

query_engine = RetrieverQueryEngine(
    retriever=base_retriever,
    node_postprocessors=[ClefNodePostprocessor(token_budget=800)],
)

Measured results

Three measurement sources, each labelled for what it actually proves.

Real model, modest hardware — open-weights Cloudflare/clef-flash (9B), float16 sharded across 2x Kaggle T4, 20 cases, 104 chunks. Produced by python evals/run_eval.py --mode local; results committed in evals/results/results-local-t4x2.json and reproducible from the clef-compactor-evals kernel.

metric value
chunk accuracy 0.712
kept precision 0.771
relevant recall 0.746
kept F1 0.758
context tokens saved 38.4%
scoring latency p50 / p95 1,274 / 1,407 ms
cost per 1k calls $0.00 (self-hosted weights)

Pipeline validation — deterministic simulated scorer, same dataset, seed 20261001 (python evals/run_eval.py). Proves the ranking and budget machinery, not model quality: chunk accuracy 0.990, kept F1 0.992, 34.0% tokens saved.

Hosted API — pending credentials; --mode live produces it.

Reading the real numbers plainly: on a T4 the 9B model makes the right keep/drop call 71% of the time, keeps three quarters of the relevant chunks, and still removes 38% of the context tokens. Local latency is dominated by the modest GPU and the torch fallback for Qwen3.5's linear attention (the fast-path kernels were not installed in the kernel); the hosted endpoint reports 38.8 ms median for clef-flash on Cloudflare's own hardware.

clef vs laya

Published numbers from Cloudflare's blog post "Introducing Clef":

benchmark clef clef-flash laya
BFCL, case exact 98.47 98.76 38.13
ToolRet, nDCG@10 69.19 66.43 12.69
API-Bank accuracy 91.93 93.11 11.41
median latency, ms 209.3 38.8 5.8
p95 latency, ms 238.6 122.4 222.5
context window 65,536 65,536 32k (reported)

Trade-off in plain terms: laya is faster (5.8 ms median because it runs locally), clef is far more accurate on decision benchmarks, and clef-flash is the middle path. clef-compactor adds one network round trip per 64 chunks on top of the model latency, and it batches, so a 12-chunk batch is one call.

Configuration

variable default meaning
CLEF_ACCOUNT_ID (or CLOUDFLARE_ACCOUNT_ID) required Cloudflare account id
CLEF_API_TOKEN (or CLOUDFLARE_API_TOKEN) required token with Workers AI run permission
CLEF_MODEL clef clef or clef-flash
CLEF_BASE_URL https://api.cloudflare.com/client/v4 override for AI Gateway or tests
CLEF_TIMEOUT 60 per-request timeout, seconds
CLEF_MAX_RETRIES 2 retries for 408/429/5xx and network errors
CLEF_LOG_LEVEL WARNING stdlib level for the clef_compactor logger

Errors are structured: everything derives from ClefError, so one except catches auth (401/403), rate limits (429 with retry_after), server errors, timeouts and malformed responses. Retries use exponential backoff with jitter and honor Retry-After.

Limitations

  • Real-model numbers are from a 9B model on a T4 pair, not from the hosted endpoint. The hosted API (and the 27B model, which needs ~54 GB) may score better. Measuring the hosted endpoint is one command away: --mode live with credentials.
  • Latency. clef's median decision latency is 209 ms (38.8 ms for clef-flash) plus network. laya keeps the whole job local at 5.8 ms. If you need sub-10 ms compaction on every request, see laya-compactor.
  • Token counting is an estimate. The default cl100k_base encoding is a close proxy, not Cloudflare's tokenizer. Budgets are enforced on the estimate.
  • 64 questions per request. Larger batches are split into multiple calls; a 200-chunk batch costs 4 calls and pays network latency 4 times.
  • Previews, not full chunks, are scored by default. chunk_preview_chars is 512 to keep requests small. Raise it toward the 64k window when chunks carry context deep into their body.
  • Relevance is binary under the hood. Clef answers P(relevant); there is no graded "supporting vs essential" signal, and the threshold (default 0.5) is a blunt but predictable cut.

Status

v0.2.0. The API surface (ClefCompactor, AsyncClefCompactor, Settings, the exception hierarchy) is settling but not frozen. The evals gate (--min-accuracy) runs on every push, and publish happens through GitHub releases with PyPI trusted publishing.

License

Apache 2.0. Clef itself is open source on Hugging Face under the same license.

Metadata

Release files for clef-compactor 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for clef-compactor 0.2.1
File Size Uploaded
clef_compactor-0.2.1.tar.gz 317.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for clef-compactor 0.2.1
File Interpreter ABI Platform
clef_compactor-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 352.1 kB

Release files / clef_compactor-0.2.1.tar.gz

Download URL clef_compactor-0.2.1.tar.gz
Size 317.3 kB
Tags Source
SHA-256 checksum
How to use checksums
f4503ca7e3f922c99b961566e42ac892d40bb303de74186944cf6d0cf1425853
BLAKE2b-256 checksum
How to use checksums
55cd1507c3d6595c9e38fe77d7203bac51ec2ba49715060ca49c49ec2cbfcbf6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / clef_compactor-0.2.1-py3-none-any.whl

Download URL clef_compactor-0.2.1-py3-none-any.whl
Size 34.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1f18e880b064913417bff25cc1055461a9593977f6a7180ea03f309f4ea5d0c1
BLAKE2b-256 checksum
How to use checksums
047bf8e74999feae9a9a8ecf11de0dd06af6a810d4589741ac801c87b9085699
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page