hypercompress
Query-aware, meaning-first context compression for LLM applications. Same objective as a commercial learned-model compressor — cut input tokens before they hit the model, keep the meaning — with a different set of engineering bets:
| the commercial baseline | hypercompress | |
|---|---|---|
| Core | trained 200KB policy model | transparent BM25 + structural boosts |
| Latency | ~60ms claimed | ~1ms p50, ~2ms p95 (measured, below) |
| Cost | free 5M tok/mo, then $1/M | $0 forever, self-hosted |
| Privacy | context sent to their API (or local pip) | never leaves your process |
| Explainability | black box | every kept/dropped block carries reasons |
| Weak-signal behavior | unknown | declines to compress (risk=high), caller fails open |
| API | hosted /api/v1/compress |
wire-compatible clone, self-hosted |
The honest caveat: a well-trained learned scorer can beat lexical scoring on paraphrase-heavy queries (where the question shares few exact words with the evidence). That is the one axis we don't claim to win — run the included head-to-head harness on your traffic and let the data decide.
Benchmark (reproduce: python benchmarks/run_benchmark.py)
60 synthetic cases, ~50 near-topic distractor paragraphs each, evidence placed at a random position, retention = ALL answer spans present verbatim in the compressed output:
| System | Evidence retention | Tokens saved (mean) | p50 latency |
|---|---|---|---|
| hypercompress-adaptive | 100% | 78.3% | 1.05ms |
| hypercompress-fixed@0.30 | 100% | 54.0% | 1.24ms |
| tail-keep@0.30 (baseline) | 35.0% | 70.0% | ~0ms |
| random-drop@0.30 (baseline) | 28.3% | 70.6% | ~0ms |
| head-keep@0.30 (baseline) | 21.7% | 70.0% | ~0ms |
Read these numbers for what they are: the questions share topic vocabulary
and exact identifiers (codes, names, figures) with their evidence — the
realistic case for support/RAG/incident queries, and exactly where lexical
scoring shines. Paraphrase-only queries will score lower; the risk signal
is designed to catch that (weak coverage → high → caller sends the
original context).
vs LLMLingua / LongLLMLingua (measured, reproduce with benchmarks/vs_llmlingua.py)
24 standard + 16 paraphrase cases, all systems local on the same CPU ("safe savings" = savings on cases where every answer span survived):
| Suite | System | Retention | Saved | Safe savings | p50 |
|---|---|---|---|---|---|
| Standard | hypercompress hybrid | 100% | 75.5% | 75.5% | 1.6ms |
| Standard | LLMLingua-2 @0.33 | 20.8% | 68.2% | 48.3% | 2.0s |
| Standard | LongLLMLingua (gpt2) | 100% | 63.5% | 63.5% | 8.8s |
| Paraphrase | hypercompress hybrid | 100% | 73.2% | 73.2% | 8.5s |
| Paraphrase | LLMLingua-2 @0.33 | 0% | 68.3% | 0% | 1.7s |
| Paraphrase | LongLLMLingua (gpt2) | 100% | 63.4% | 63.4% | 9.2s |
The hybrid wins BOTH suites on safe savings with equal-or-better retention: a lexical fast path (~1.6ms) serves confidently-answerable queries, and a false-confidence guard (top-block score spike detection) routes everything else to a LongLLMLingua-style neural stage run at a harder rate (0.25 vs their 0.33) whose over-compression is rescued by a fingerprint sweep — re-adding identifier-bearing paragraphs the neural stage dropped. Credit where due: the fallback uses the llmlingua package itself; the wins come from the routing, the harder rate, and the sweep. Caveats: their backbone here is gpt2 (CPU constraint) — the ACL'24 paper used LLaMA-7B, which would score higher; and these are synthetic suites — validate on your own traffic (benchmarks/my_data_benchmark.py).
from hypercompress import compress_context_hybrid # pip install ".[neural]"
result = compress_context_hybrid(context, question) # 1ms fast path, neural only on declines
Install
pip install . # core: zero dependencies
pip install ".[server]" # + self-hosted API (FastAPI/uvicorn)
pip install ".[mcp]" # + MCP server for coding agents
pip install ".[exact-tokens]" # + tiktoken for exact token counts
Library (drop-in for common compression-client conventions)
from hypercompress import compress_context
result = compress_context(context, question) # adaptive mode
result = compress_context(context, question, 0.3) # fixed 30% budget
result.compressed_text # send this to the LLM
result.tokens_saved_pct # e.g. 78.3
result.compression_risk # "low" | "medium" | "high" -> fall back if high
result.kept_blocks # audit trail: every block, score, reasons
Message-list form (system prompt + latest user message always verbatim):
from hypercompress import compress_for_turn
messages = compress_for_turn(messages)
Self-hosted API (hosted-API compatible)
uvicorn hypercompress.server:app --port 8765
# optional auth: export HYPERCOMPRESS_API_KEY=hc_your_key
curl -X POST localhost:8765/api/v1/compress \
-H 'content-type: application/json' \
-d '{"context":"...long context...","query":"what failed?"}'
Response schema matches the commercial baseline (compressed_text, original_tokens,
kept_tokens, tokens_saved_pct, important_kept_pct, compression_risk,
kept_blocks, dropped_blocks, policy_name), and the same auth headers
(X-API-Key / Authorization: Bearer) are accepted — existing the commercial baseline
client code migrates by changing one URL.
Integrations (full the commercial baseline parity)
OpenAI — hypercompress/wrappers/openai_wrapper.py
from openai import OpenAI
from hypercompress.wrappers.openai_wrapper import HyperCompressOpenAI
client = HyperCompressOpenAI(OpenAI())
client.chat.completions.create(model="gpt-4o-mini", messages=msgs)
client.stats.tokens_saved_pct # verified savings, not vendor claims
Anthropic — hypercompress/wrappers/anthropic_wrapper.py
from hypercompress.wrappers.anthropic_wrapper import HyperCompressAnthropic
client = HyperCompressAnthropic(anthropic.Anthropic())
LangChain — hypercompress/wrappers/langchain_hook.py
from hypercompress.wrappers.langchain_hook import compress_lc_messages
chain.invoke(compress_lc_messages(messages))
Express / Next.js — js/src/index.js
const { expressMiddleware } = require("hypercompress");
app.post("/chat", expressMiddleware(), handler); // compresses req.body.messages
Vercel AI SDK — provider-agnostic
const { wrapGenerateText } = require("hypercompress");
const gen = wrapGenerateText(generateText);
await gen({ model, messages });
MCP (Claude Code / Cursor / Codex / Windsurf)
pip install ".[mcp]"
claude mcp add hypercompress -- python -m hypercompress.mcp_server
Exposes compress_context and compress_file tools so agents can pull
query-relevant slices of big files instead of whole files.
Guardrails (identical across every integration)
- System prompts and the latest user message are never compressed.
- Fail open — errors, timeouts, and
compression_risk == "high"all send the original context. Compression must never break an answer. - Elisions are marked with
[…]so the model knows content was removed. - Savings are measured and logged locally; nothing here asks you to trust a marketing number.
How it works
splitter.py cuts context into structure-aware blocks (code fences atomic,
markdown sections, chat turns, log runs with ERROR lines isolated).
scoring.py ranks blocks with Okapi BM25 plus exact-identifier boosts
(error codes, numbers, dotted names), bigram matches, log severity, and chat
recency; headers inherit their best child's score so surviving sections keep
their titles. core.py selects adaptively (keep while marginal relevance is
meaningful) or under a fixed budget, stitches ±1 neighbor blocks for local
coherence, reassembles in original order with […] gap markers, and computes
compression_risk from measured query-term coverage.
Tests
python -m pytest tests/ # 14 tests: core behavior, wire compat, wrappers
Author
Built by Natarajan Venkatasubramaniam (Natarajan.Venkatasubramaniam@wissen.com).
License
MIT © 2026 Natarajan Venkatasubramaniam. See LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hypercompress-0.3.2.tar.gz.
File metadata
- Download URL: hypercompress-0.3.2.tar.gz
- Upload date:
- Size: 32.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d0039e1aadb740c54ba4d5df4f128ac2bab1a11465c6e8f542415d550a8e9dcd
|
|
| MD5 |
7fdecdffced3e7eb969976a3e4d4fa71
|
|
| BLAKE2b-256 |
1db1e967c8d990df5e1913d7218c1729916575ef9daa8113ba74c8e69730654e
|
File details
Details for the file hypercompress-0.3.2-py3-none-any.whl.
File metadata
- Download URL: hypercompress-0.3.2-py3-none-any.whl
- Upload date:
- Size: 29.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dc2329867c01a1c761981201293b1c831f3cee0094f54a112bbd078ec702eb20
|
|
| MD5 |
8466ad4b6ef225306155e1be628ef7e9
|
|
| BLAKE2b-256 |
20838005221b24ce0979fd1467905a848db9cf21d7735fe7e8dd66108ac26374
|