ContextPilot
Cut your LLM API costs 60-80% with one line of code.
ContextPilot is a Python middleware library that compresses LLM context before each API call. It wraps OpenAI and Anthropic SDKs, runs a compression pipeline, and falls back to the original payload if quality drops. Works as a Python library, a local proxy, an MCP server, or a CLI migration tool.
Website: contextpilot.org | PyPI: contextpilot-ai
How it works
Every API call goes through four steps:
- Analyze each message block for staleness, redundancy, relevance, and density
- Compress by summarizing history, deduplicating system prompts, pruning irrelevant RAG chunks, and stripping structural noise
- Quality gate checks the predicted score. If it drops below the threshold (default 72/100), the original payload goes out instead
- Forward the optimized (or original) payload to the provider; the response comes back unchanged
No prompt content ever leaves your machine. Telemetry is numerical metadata only.
Before you install it against real code
Compression runs in your own process, in memory. Your prompts and responses go to the LLM provider you're already using and nowhere else, not to us, not to any third party. The local event log is metadata only (token counts, latency, quality scores, no prompt text) and it stays on your machine unless you deliberately opt in to a hosted dashboard, which doesn't exist yet, so there's currently nothing to opt into even if you tried. See Privacy below, SECURITY.md for the full data-handling policy, and docs/limitations.md for an honest list of tradeoffs and what this tool doesn't do.
Benchmarks
Measured on realistic production conversation patterns. Each scenario uses actual repetition patterns developers encounter: accumulated context, repeated RAG chunks, repeated error traces, multi-agent handoffs.
| Scenario | Tokens | Reduction | Quality | Latency |
|---|---|---|---|---|
| AI coding assistant, 25 turns, growing project context | 5,810 → 1,118 | 80.8% | 82.8/100 | 10ms |
| RAG chatbot, 18 turns, 5 retrieved chunks per query | 4,980 → 1,034 | 79.2% | 83.4/100 | 9ms |
| Multi-agent code review, 4 agents x 6 rounds | 19,619 → 4,049 | 79.4% | 83.9/100 | 22ms |
| Production debugging, 20 turns, repeated tracebacks | 3,814 → 928 | 75.7% | 82.4/100 | 9ms |
| LangChain tool agent, 15 turns, 3 tool outputs/turn | 5,368 → 1,278 | 76.2% | 83.7/100 | 8ms |
| Document Q&A, 16 turns, full spec prepended each query | 4,561 → 1,110 | 75.7% | 83.9/100 | 8ms |
The quality gate skips compression whenever quality drops below threshold. In all 6 scenarios above, quality held at 82-84/100, well above the default 72.
Cost at scale (most impactful scenario: multi-agent on Claude Opus):
| Volume | Without ContextPilot | With ContextPilot | Monthly saving |
|---|---|---|---|
| 100 calls/day | $29/day | $6/day | $701/mo |
| 1,000 calls/day | $294/day | $61/day | $7,006/mo |
| 10,000 calls/day | $2,943/day | $607/day | $70,065/mo |
Run python benchmarks/benchmark_readme.py to reproduce locally. These benchmarks top out around 20K tokens per conversation; see docs/limitations.md for where the performance budget is and isn't independently verified yet.
Integration surfaces
| Surface | Entry point | Best for |
|---|---|---|
| Python library | contextpilot.wrap(client) |
Backend apps, RAG pipelines, agents |
| Proxy (service) | contextpilot service install |
Claude Code, GPT Codex, Aider, always on |
| Proxy (manual) | contextpilot proxy --port 8432 |
Temporary sessions or per-project use |
| MCP server | claude mcp add contextpilot -- contextpilot mcp |
Claude Desktop, Claude Code |
| CLI migration | contextpilot migrate ./src/ |
Existing codebases with 50+ LLM calls |
If you're using Claude Code, Codex CLI, or another agent that already does its own session-level context compaction, ContextPilot is complementary, not a replacement: it trims the payload of each individual API call, while the coding tool manages the overall conversation.
Quick Start
Python library
pip install contextpilot-ai
OpenAI:
import contextpilot
from openai import OpenAI
client = contextpilot.wrap(OpenAI())
response = client.chat.completions.create(
model="gpt-4o",
messages=messages # compressed transparently
)
Anthropic:
import contextpilot
from anthropic import Anthropic
client = contextpilot.wrap(Anthropic())
response = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=messages
)
That's the full integration. No other code changes required.
Proxy for Claude Code, GPT Codex, and Aider
The proxy intercepts every request from your AI coding tool and compresses it before it reaches the provider.
Recommended: install as a background service
One command. Runs automatically on every login, no terminal to keep open.
pipx install "contextpilot-ai[proxy]"
contextpilot service install
That's it. ContextPilot will:
- Start silently on login (Windows Task Scheduler / macOS launchd / Linux systemd)
- Set
ANTHROPIC_BASE_URLpermanently in your environment - Restart automatically if it ever crashes
- Compress every Claude Code, GPT Codex, and Aider request with zero ongoing effort
Restart VS Code (or open a new terminal) once to pick up the environment variable.
contextpilot service status # confirm it's running
contextpilot service uninstall # remove if you ever want to stop
Manual: start per session
Useful for temporary use or when you only want compression for a specific project:
# Terminal 1 (keep this open)
contextpilot proxy --port 8432
# Terminal 2 (set the env var, then use your tool normally)
export ANTHROPIC_BASE_URL=http://localhost:8432 # Linux / macOS
$env:ANTHROPIC_BASE_URL = "http://localhost:8432" # Windows PowerShell
# OpenAI SDK / GPT Codex / Aider
export OPENAI_BASE_URL=http://localhost:8432/v1
python -m contextpilot proxy --port 8432 works as a fallback if contextpilot is not in your PATH.
MCP Server for Claude Desktop and Claude Code
Register once:
claude mcp add contextpilot -- contextpilot mcp
Restart Claude Code (or reload the VS Code window). ContextPilot appears as a connected MCP server. Claude will call optimize_context when processing large contexts, include contextpilot.wrap() in any LLM code it generates for you, and report savings on request via the contextpilot://savings resource.
To verify: ask Claude Code "What MCP tools do you have available?" and you should see optimize_context and optimize_llm_code.
CLI Migration for existing codebases
# Preview what would change
contextpilot migrate ./src/ --dry-run
# Rewrite files in place
contextpilot migrate ./src/ --apply
Uses AST parsing (not regex) to find every OpenAI() and Anthropic() instantiation and wrap it with contextpilot.wrap(). Designed for codebases with 50+ LLM calls where manual refactoring isn't practical.
Savings Report
contextpilot report
Reads the local event log (~/.contextpilot/events.jsonl) and shows token savings, compression ratio, quality scores, and estimated cost saved. No dashboard required.
ContextPilot · Savings Report
────────────────────────────────────────
Calls logged : 142
Token reduction
████████████░░░░░░░░░░░░░░░░ 37.3% saved
284,391 → 178,203 (saved 106,188 tokens)
Quality avg : 91.4 / 100 ✓
Fallback rate : 8/142 (5.6%)
Est. cost saved: ~$0.5309
Log: /home/user/.contextpilot/events.jsonl
Agent Memory Middleware
Compress inter-agent context handoffs in LangChain, CrewAI, and AutoGen pipelines that otherwise multiply tokens 5-30x:
from contextpilot.middleware import AgentMemory
memory = AgentMemory(
compression_level="aggressive",
preserve_keys=["final_answer", "tool_outputs"],
)
compressed = memory.compress_handoff(agent_a.run(task))
result = agent_b.run(task, context=compressed)
Configuration
Drop a contextpilot.yaml in your project root:
compression:
level: balanced # conservative | balanced | aggressive
quality_threshold: 72 # fallback to original if score drops below this
history_window: 6 # keep last N turns verbatim
rag_relevance_min: 0.15 # drop RAG chunks below this relevance score
shadow_testing:
enabled: false
sample_rate: 0.05 # fraction of calls sent both compressed and uncompressed
telemetry:
enabled: true
# The two fields below are reserved for the future hosted dashboard.
# That service doesn't exist yet, so setting api_key today has no effect,
# nothing gets sent anywhere. Local logging to ~/.contextpilot/events.jsonl
# always works and needs neither of these.
endpoint: https://api.contextpilot.org/v1/telemetry
api_key: ${CONTEXTPILOT_API_KEY}
Environment variable overrides: CONTEXTPILOT_COMPRESSION_LEVEL, CONTEXTPILOT_QUALITY_THRESHOLD, CONTEXTPILOT_API_KEY.
Privacy
Telemetry sends numerical metadata only: token counts, latency, quality scores, model IDs, timestamps. No prompt content, no response content, no PII ever leaves your environment. This is an architectural guarantee, not a policy: compression runs in-process, and the telemetry schema has no field for content, so there's nothing to accidentally send even if that changed.
Local logging (~/.contextpilot/events.jsonl) is on by default and never leaves your machine. Remote sync to a hosted dashboard is opt-in only, requires an explicit API key, and today that endpoint isn't live yet, so enabling it is a no-op rather than a silent data leak.
See SECURITY.md for the full data handling policy, proxy trust model, and vulnerability reporting process, and docs/limitations.md for what this tool doesn't do yet.
Installation
Library (inside a project)
pip install contextpilot-ai # core library
pip install "contextpilot-ai[proxy]" # + proxy server (starlette, uvicorn)
pip install "contextpilot-ai[openai]" # + openai SDK
pip install "contextpilot-ai[anthropic]" # + anthropic SDK
pip install "contextpilot-ai[mcp]" # + MCP server
pip install "contextpilot-ai[all]" # everything
CLI / proxy (recommended: pipx)
pipx installs CLI tools in isolated environments and wires them into your PATH automatically, no virtualenv activation needed in new terminals:
pipx install "contextpilot-ai[proxy,mcp]"
Without pipx:
pip install "contextpilot-ai[proxy,mcp]"
If contextpilot is not recognized after install, use the module form:
python -m contextpilot service install
python -m contextpilot proxy --port 8432
python -m contextpilot mcp
Contributing
See CONTRIBUTING.md.
License
MIT, see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file contextpilot_ai-0.3.1.tar.gz.
File metadata
- Download URL: contextpilot_ai-0.3.1.tar.gz
- Upload date:
- Size: 90.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ede4acbb19b46b8fc6761744f94f9bc702d76b3e1a5bad9e0fae9e2402b1a55f
|
|
| MD5 |
9245df1734e32915142830b37777e9cb
|
|
| BLAKE2b-256 |
54be1e6b8a0392e22b4a4d465c4033d2fa426cdc7595e0368ff362baabc7aa8a
|
Provenance
The following attestation bundles were made for contextpilot_ai-0.3.1.tar.gz:
Publisher:
release.yml on msousa202/ContextPilot
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
contextpilot_ai-0.3.1.tar.gz -
Subject digest:
ede4acbb19b46b8fc6761744f94f9bc702d76b3e1a5bad9e0fae9e2402b1a55f - Sigstore transparency entry: 2256399340
- Sigstore integration time:
-
Permalink:
msousa202/ContextPilot@63ff844e47001b9023a3b40468787639c6e7d561 -
Branch / Tag:
refs/tags/v0.3.1 - Owner: https://github.com/msousa202
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@63ff844e47001b9023a3b40468787639c6e7d561 -
Trigger Event:
push
-
Statement type:
File details
Details for the file contextpilot_ai-0.3.1-py3-none-any.whl.
File metadata
- Download URL: contextpilot_ai-0.3.1-py3-none-any.whl
- Upload date:
- Size: 49.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
948cef31a6e45e1ab44444964b664d3823cba1370362d66ddbe18f5e5145a5c5
|
|
| MD5 |
33f2446e5cb672841eaf509a5311a5b4
|
|
| BLAKE2b-256 |
b2c7a14636ad9eab955d008d0bd7d5f1a7fe608927bb9d4875aef92b3adf96fb
|
Provenance
The following attestation bundles were made for contextpilot_ai-0.3.1-py3-none-any.whl:
Publisher:
release.yml on msousa202/ContextPilot
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
contextpilot_ai-0.3.1-py3-none-any.whl -
Subject digest:
948cef31a6e45e1ab44444964b664d3823cba1370362d66ddbe18f5e5145a5c5 - Sigstore transparency entry: 2256399345
- Sigstore integration time:
-
Permalink:
msousa202/ContextPilot@63ff844e47001b9023a3b40468787639c6e7d561 -
Branch / Tag:
refs/tags/v0.3.1 - Owner: https://github.com/msousa202
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@63ff844e47001b9023a3b40468787639c6e7d561 -
Trigger Event:
push
-
Statement type: