Skip to main content

ContextPilot

PyPI CI License: MIT Python 3.10+

Cut your LLM API costs 60-80% with one line of code.

ContextPilot is a Python middleware library that compresses LLM context before each API call. It wraps OpenAI and Anthropic SDKs, runs a compression pipeline, and falls back to the original payload if quality drops. Works as a Python library, a local proxy, an MCP server, or a CLI migration tool.

Website: contextpilot.org | PyPI: contextpilot-ai


How it works

Every API call goes through four steps:

  1. Analyze each message block for staleness, redundancy, relevance, and density
  2. Compress by summarizing history, deduplicating system prompts, pruning irrelevant RAG chunks, and stripping structural noise
  3. Quality gate checks the predicted score. If it drops below the threshold (default 72/100), the original payload goes out instead
  4. Forward the optimized (or original) payload to the provider; the response comes back unchanged

No prompt content ever leaves your machine. Telemetry is numerical metadata only.


Before you install it against real code

Compression runs in your own process, in memory. Your prompts and responses go to the LLM provider you're already using and nowhere else, not to us, not to any third party. The local event log is metadata only (token counts, latency, quality scores, no prompt text) and it stays on your machine unless you deliberately opt in to a hosted dashboard, which doesn't exist yet, so there's currently nothing to opt into even if you tried. See Privacy below, SECURITY.md for the full data-handling policy, and docs/limitations.md for an honest list of tradeoffs and what this tool doesn't do.


Benchmarks

Measured on realistic production conversation patterns. Each scenario uses actual repetition patterns developers encounter: accumulated context, repeated RAG chunks, repeated error traces, multi-agent handoffs.

Scenario Tokens Reduction Quality Latency
AI coding assistant, 25 turns, growing project context 5,810 → 1,118 80.8% 82.8/100 10ms
RAG chatbot, 18 turns, 5 retrieved chunks per query 4,980 → 1,034 79.2% 83.4/100 9ms
Multi-agent code review, 4 agents x 6 rounds 19,619 → 4,049 79.4% 83.9/100 22ms
Production debugging, 20 turns, repeated tracebacks 3,814 → 928 75.7% 82.4/100 9ms
LangChain tool agent, 15 turns, 3 tool outputs/turn 5,368 → 1,278 76.2% 83.7/100 8ms
Document Q&A, 16 turns, full spec prepended each query 4,561 → 1,110 75.7% 83.9/100 8ms

The quality gate skips compression whenever quality drops below threshold. In all 6 scenarios above, quality held at 82-84/100, well above the default 72.

Cost at scale (most impactful scenario: multi-agent on Claude Opus):

Volume Without ContextPilot With ContextPilot Monthly saving
100 calls/day $29/day $6/day $701/mo
1,000 calls/day $294/day $61/day $7,006/mo
10,000 calls/day $2,943/day $607/day $70,065/mo

Run python benchmarks/benchmark_readme.py to reproduce locally. These benchmarks top out around 20K tokens per conversation; see docs/limitations.md for where the performance budget is and isn't independently verified yet.


Integration surfaces

Surface Entry point Best for
Python library contextpilot.wrap(client) Backend apps, RAG pipelines, agents
Proxy (service) contextpilot service install Claude Code, GPT Codex, Aider, always on
Proxy (manual) contextpilot proxy --port 8432 Temporary sessions or per-project use
MCP server claude mcp add contextpilot -- contextpilot mcp Claude Desktop, Claude Code
CLI migration contextpilot migrate ./src/ Existing codebases with 50+ LLM calls

If you're using Claude Code, Codex CLI, or another agent that already does its own session-level context compaction, ContextPilot is complementary, not a replacement: it trims the payload of each individual API call, while the coding tool manages the overall conversation.


Quick Start

Python library

pip install contextpilot-ai

OpenAI:

import contextpilot
from openai import OpenAI

client = contextpilot.wrap(OpenAI())

response = client.chat.completions.create(
    model="gpt-4o",
    messages=messages  # compressed transparently
)

Anthropic:

import contextpilot
from anthropic import Anthropic

client = contextpilot.wrap(Anthropic())

response = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=1024,
    messages=messages
)

That's the full integration. No other code changes required.


Proxy for Claude Code, GPT Codex, and Aider

The proxy intercepts every request from your AI coding tool and compresses it before it reaches the provider.

Recommended: install as a background service

One command. Runs automatically on every login, no terminal to keep open.

pipx install "contextpilot-ai[proxy]"
contextpilot service install

That's it. ContextPilot will:

  • Start silently on login (Windows Task Scheduler / macOS launchd / Linux systemd)
  • Set ANTHROPIC_BASE_URL permanently in your environment
  • Restart automatically if it ever crashes
  • Compress every Claude Code, GPT Codex, and Aider request with zero ongoing effort

Restart VS Code (or open a new terminal) once to pick up the environment variable.

contextpilot service status     # confirm it's running
contextpilot service uninstall  # remove if you ever want to stop

Manual: start per session

Useful for temporary use or when you only want compression for a specific project:

# Terminal 1 (keep this open)
contextpilot proxy --port 8432

# Terminal 2 (set the env var, then use your tool normally)
export ANTHROPIC_BASE_URL=http://localhost:8432      # Linux / macOS
$env:ANTHROPIC_BASE_URL = "http://localhost:8432"    # Windows PowerShell

# OpenAI SDK / GPT Codex / Aider
export OPENAI_BASE_URL=http://localhost:8432/v1

python -m contextpilot proxy --port 8432 works as a fallback if contextpilot is not in your PATH.


MCP Server for Claude Desktop and Claude Code

Register once:

claude mcp add contextpilot -- contextpilot mcp

Restart Claude Code (or reload the VS Code window). ContextPilot appears as a connected MCP server. Claude will call optimize_context when processing large contexts, include contextpilot.wrap() in any LLM code it generates for you, and report savings on request via the contextpilot://savings resource.

To verify: ask Claude Code "What MCP tools do you have available?" and you should see optimize_context and optimize_llm_code.


CLI Migration for existing codebases

# Preview what would change
contextpilot migrate ./src/ --dry-run

# Rewrite files in place
contextpilot migrate ./src/ --apply

Uses AST parsing (not regex) to find every OpenAI() and Anthropic() instantiation and wrap it with contextpilot.wrap(). Designed for codebases with 50+ LLM calls where manual refactoring isn't practical.


Savings Report

contextpilot report

Reads the local event log (~/.contextpilot/events.jsonl) and shows token savings, compression ratio, quality scores, and estimated cost saved. No dashboard required.

  ContextPilot · Savings Report
  ────────────────────────────────────────
  Calls logged   :  142

  Token reduction
  ████████████░░░░░░░░░░░░░░░░  37.3% saved
  284,391 → 178,203  (saved 106,188 tokens)

  Quality avg    :  91.4 / 100  ✓
  Fallback rate  :  8/142  (5.6%)
  Est. cost saved:  ~$0.5309

  Log: /home/user/.contextpilot/events.jsonl

Agent Memory Middleware

Compress inter-agent context handoffs in LangChain, CrewAI, and AutoGen pipelines that otherwise multiply tokens 5-30x:

from contextpilot.middleware import AgentMemory

memory = AgentMemory(
    compression_level="aggressive",
    preserve_keys=["final_answer", "tool_outputs"],
)

compressed = memory.compress_handoff(agent_a.run(task))
result = agent_b.run(task, context=compressed)

Configuration

Drop a contextpilot.yaml in your project root:

compression:
  level: balanced          # conservative | balanced | aggressive
  quality_threshold: 72    # fallback to original if score drops below this
  history_window: 6        # keep last N turns verbatim
  rag_relevance_min: 0.15  # drop RAG chunks below this relevance score

shadow_testing:
  enabled: false
  sample_rate: 0.05        # fraction of calls sent both compressed and uncompressed

telemetry:
  enabled: true
  # The two fields below are reserved for the future hosted dashboard.
  # That service doesn't exist yet, so setting api_key today has no effect,
  # nothing gets sent anywhere. Local logging to ~/.contextpilot/events.jsonl
  # always works and needs neither of these.
  endpoint: https://api.contextpilot.org/v1/telemetry
  api_key: ${CONTEXTPILOT_API_KEY}

Environment variable overrides: CONTEXTPILOT_COMPRESSION_LEVEL, CONTEXTPILOT_QUALITY_THRESHOLD, CONTEXTPILOT_API_KEY.


Privacy

Telemetry sends numerical metadata only: token counts, latency, quality scores, model IDs, timestamps. No prompt content, no response content, no PII ever leaves your environment. This is an architectural guarantee, not a policy: compression runs in-process, and the telemetry schema has no field for content, so there's nothing to accidentally send even if that changed.

Local logging (~/.contextpilot/events.jsonl) is on by default and never leaves your machine. Remote sync to a hosted dashboard is opt-in only, requires an explicit API key, and today that endpoint isn't live yet, so enabling it is a no-op rather than a silent data leak.

See SECURITY.md for the full data handling policy, proxy trust model, and vulnerability reporting process, and docs/limitations.md for what this tool doesn't do yet.


Installation

Library (inside a project)

pip install contextpilot-ai                    # core library
pip install "contextpilot-ai[proxy]"           # + proxy server (starlette, uvicorn)
pip install "contextpilot-ai[openai]"          # + openai SDK
pip install "contextpilot-ai[anthropic]"       # + anthropic SDK
pip install "contextpilot-ai[mcp]"             # + MCP server
pip install "contextpilot-ai[all]"             # everything

CLI / proxy (recommended: pipx)

pipx installs CLI tools in isolated environments and wires them into your PATH automatically, no virtualenv activation needed in new terminals:

pipx install "contextpilot-ai[proxy,mcp]"

Without pipx:

pip install "contextpilot-ai[proxy,mcp]"

If contextpilot is not recognized after install, use the module form:

python -m contextpilot service install
python -m contextpilot proxy --port 8432
python -m contextpilot mcp

Contributing

See CONTRIBUTING.md.


License

MIT, see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

contextpilot_ai-0.4.0.tar.gz (103.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

contextpilot_ai-0.4.0-py3-none-any.whl (57.2 kB view details)

Uploaded Python 3

File details

Details for the file contextpilot_ai-0.4.0.tar.gz.

File metadata

  • Download URL: contextpilot_ai-0.4.0.tar.gz
  • Upload date:
  • Size: 103.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for contextpilot_ai-0.4.0.tar.gz
Algorithm Hash digest
SHA256 7b523968124ef4db7ff527c4d5306df538fa70bfe26252fd32f88447584b4d3c
MD5 8992a3d5f703a6bebb4eff579f7d5bef
BLAKE2b-256 e7ff81d5e6dc9335b9adb6c94b110b05b188b69ba7918a560f6643bf88ae8531

See more details on using hashes here.

Provenance

The following attestation bundles were made for contextpilot_ai-0.4.0.tar.gz:

Publisher: release.yml on msousa202/ContextPilot

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file contextpilot_ai-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: contextpilot_ai-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 57.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for contextpilot_ai-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a76021ddb8abd7b7ca358be2145c64a8ebb73dfb6a7b8288ea717d2c31a488ab
MD5 a1aa8e72404e486aa09d5f525e65175d
BLAKE2b-256 da85ba44f7e93db0c3aad4c9cddf9cab97d77f09f8b0e822d442d1cbebe6cf85

See more details on using hashes here.

Provenance

The following attestation bundles were made for contextpilot_ai-0.4.0-py3-none-any.whl:

Publisher: release.yml on msousa202/ContextPilot

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.4.1

2 files

This release

0.4.0 This release

2 files

0.3.1

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page