Skip to main content

ContextPilot

PyPI CI License: MIT Python 3.10+ Tests

Cut your LLM API costs 60-80% with one line of code.

ContextPilot is a Python middleware library that compresses LLM context before each API call. It wraps OpenAI and Anthropic SDKs, runs a compression pipeline, and falls back to the original payload if quality drops. Works as a Python library, a local proxy, an MCP server, or a CLI migration tool.

Website: contextpilot.org | PyPI: contextpilot-ai


How it works

Every API call goes through four steps:

  1. Analyze each message block for staleness, redundancy, relevance, and density
  2. Compress by summarizing history, deduplicating system prompts, pruning irrelevant RAG chunks, and stripping structural noise
  3. Quality gate checks the predicted score. If it drops below the threshold (default 72/100), the original payload goes out instead
  4. Forward the optimized (or original) payload to the provider; the response comes back unchanged

No prompt content ever leaves your machine. Telemetry is numerical metadata only.


Benchmarks

Measured on realistic production conversation patterns. Each scenario uses actual repetition patterns developers encounter: accumulated context, repeated RAG chunks, repeated error traces, multi-agent handoffs.

Scenario Tokens Reduction Quality Latency
AI coding assistant, 25 turns, growing project context 5,810 → 1,118 80.8% 82.8/100 10ms
RAG chatbot, 18 turns, 5 retrieved chunks per query 4,980 → 1,034 79.2% 83.4/100 9ms
Multi-agent code review, 4 agents x 6 rounds 19,619 → 4,049 79.4% 83.9/100 22ms
Production debugging, 20 turns, repeated tracebacks 3,814 → 928 75.7% 82.4/100 9ms
LangChain tool agent, 15 turns, 3 tool outputs/turn 5,368 → 1,278 76.2% 83.7/100 8ms
Document Q&A, 16 turns, full spec prepended each query 4,561 → 1,110 75.7% 83.9/100 8ms

The quality gate skips compression whenever quality drops below threshold. In all 6 scenarios above, quality held at 82-84/100, well above the default 72.

Cost at scale (most impactful scenario: multi-agent on Claude Opus):

Volume Without ContextPilot With ContextPilot Monthly saving
100 calls/day $29/day $6/day $701/mo
1,000 calls/day $294/day $61/day $7,006/mo
10,000 calls/day $2,943/day $607/day $70,065/mo

Run python benchmarks/benchmark_readme.py to reproduce locally.


Integration surfaces

Surface Entry point Best for
Python library contextpilot.wrap(client) Backend apps, RAG pipelines, agents
Proxy (service) contextpilot service install Claude Code, GPT Codex, Aider — always on
Proxy (manual) contextpilot proxy --port 8432 Temporary sessions or per-project use
MCP server claude mcp add contextpilot -- contextpilot mcp Claude Desktop, Claude Code
CLI migration contextpilot migrate ./src/ Existing codebases with 50+ LLM calls

Quick Start

Python library

pip install contextpilot-ai

OpenAI:

import contextpilot
from openai import OpenAI

client = contextpilot.wrap(OpenAI())

response = client.chat.completions.create(
    model="gpt-4o",
    messages=messages  # compressed transparently
)

Anthropic:

import contextpilot
from anthropic import Anthropic

client = contextpilot.wrap(Anthropic())

response = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=1024,
    messages=messages
)

That's the full integration. No other code changes required.


Proxy for Claude Code, GPT Codex, and Aider

The proxy intercepts every request from your AI coding tool and compresses it before it reaches the provider.

Recommended: install as a background service

One command. Runs automatically on every login, no terminal to keep open.

pipx install "contextpilot-ai[proxy]"
contextpilot service install

That's it. ContextPilot will:

  • Start silently on login (Windows Task Scheduler / macOS launchd / Linux systemd)
  • Set ANTHROPIC_BASE_URL permanently in your environment
  • Restart automatically if it ever crashes
  • Compress every Claude Code, GPT Codex, and Aider request with zero ongoing effort

Restart VS Code (or open a new terminal) once to pick up the environment variable.

contextpilot service status     # confirm it's running
contextpilot service uninstall  # remove if you ever want to stop

Manual: start per session

Useful for temporary use or when you only want compression for a specific project:

# Terminal 1 — keep this open
contextpilot proxy --port 8432

# Terminal 2 — set the env var, then use your tool normally
export ANTHROPIC_BASE_URL=http://localhost:8432      # Linux / macOS
$env:ANTHROPIC_BASE_URL = "http://localhost:8432"    # Windows PowerShell

# OpenAI SDK / GPT Codex / Aider
export OPENAI_BASE_URL=http://localhost:8432/v1

python -m contextpilot proxy --port 8432 works as a fallback if contextpilot is not in your PATH.


MCP Server for Claude Desktop and Claude Code

Register once:

claude mcp add contextpilot -- contextpilot mcp

Restart Claude Code (or reload the VS Code window). ContextPilot appears as a connected MCP server. Claude will call optimize_context when processing large contexts, include contextpilot.wrap() in any LLM code it generates for you, and report savings on request via the contextpilot://savings resource.

To verify: ask Claude Code "What MCP tools do you have available?" and you should see optimize_context and optimize_llm_code.


CLI Migration for existing codebases

# Preview what would change
contextpilot migrate ./src/ --dry-run

# Rewrite files in place
contextpilot migrate ./src/ --apply

Uses AST parsing (not regex) to find every OpenAI() and Anthropic() instantiation and wrap it with contextpilot.wrap(). Designed for codebases with 50+ LLM calls where manual refactoring isn't practical.


Savings Report

contextpilot report

Reads the local event log (~/.contextpilot/events.jsonl) and shows token savings, compression ratio, quality scores, and estimated cost saved. No dashboard required.

  ContextPilot — Savings Report
  ────────────────────────────────────────
  Calls logged   :  142

  Token reduction
  ████████████░░░░░░░░░░░░░░░░  37.3% saved
  284,391 → 178,203  (saved 106,188 tokens)

  Quality avg    :  91.4 / 100  ✓
  Fallback rate  :  8/142  (5.6%)
  Est. cost saved:  ~$0.5309

  Log: /home/user/.contextpilot/events.jsonl

Agent Memory Middleware

Compress inter-agent context handoffs in LangChain, CrewAI, and AutoGen pipelines that otherwise multiply tokens 5-30x:

from contextpilot.middleware import AgentMemory

memory = AgentMemory(
    compression_level="aggressive",
    preserve_keys=["final_answer", "tool_outputs"],
)

compressed = memory.compress_handoff(agent_a.run(task))
result = agent_b.run(task, context=compressed)

Configuration

Drop a contextpilot.yaml in your project root:

compression:
  level: balanced          # conservative | balanced | aggressive
  quality_threshold: 72    # fallback to original if score drops below this
  history_window: 6        # keep last N turns verbatim
  rag_relevance_min: 0.15  # drop RAG chunks below this relevance score

shadow_testing:
  enabled: false
  sample_rate: 0.05        # fraction of calls sent both compressed and uncompressed

telemetry:
  enabled: true
  endpoint: https://api.contextpilot.org/v1/telemetry
  api_key: ${CONTEXTPILOT_API_KEY}

Environment variable overrides: CONTEXTPILOT_COMPRESSION_LEVEL, CONTEXTPILOT_QUALITY_THRESHOLD, CONTEXTPILOT_API_KEY.


Privacy

Telemetry sends numerical metadata only: token counts, latency, quality scores, model IDs, timestamps. No prompt content, no response content, no PII ever leaves your environment. This is an architectural guarantee, not a policy.

See SECURITY.md for the full data handling policy, proxy trust model, and vulnerability reporting process.


Installation

Library (inside a project)

pip install contextpilot-ai                    # core library
pip install "contextpilot-ai[proxy]"           # + proxy server (starlette, uvicorn)
pip install "contextpilot-ai[openai]"          # + openai SDK
pip install "contextpilot-ai[anthropic]"       # + anthropic SDK
pip install "contextpilot-ai[mcp]"             # + MCP server
pip install "contextpilot-ai[all]"             # everything

CLI / proxy (recommended: pipx)

pipx installs CLI tools in isolated environments and wires them into your PATH automatically, no virtualenv activation needed in new terminals:

pipx install "contextpilot-ai[proxy,mcp]"

Without pipx:

pip install "contextpilot-ai[proxy,mcp]"

If contextpilot is not recognized after install, use the module form:

python -m contextpilot service install
python -m contextpilot proxy --port 8432
python -m contextpilot mcp

Contributing

See CONTRIBUTING.md.


License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

contextpilot_ai-0.3.0.tar.gz (84.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

contextpilot_ai-0.3.0-py3-none-any.whl (49.0 kB view details)

Uploaded Python 3

File details

Details for the file contextpilot_ai-0.3.0.tar.gz.

File metadata

  • Download URL: contextpilot_ai-0.3.0.tar.gz
  • Upload date:
  • Size: 84.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for contextpilot_ai-0.3.0.tar.gz
Algorithm Hash digest
SHA256 9df1e4a6149054128f2b9d9794133613897e37ab106974cf662f9cf069ef64d3
MD5 606012eb29a22cdb34ea271686d18630
BLAKE2b-256 2c3d156ec0179c33e270362b3c94655a2058ef04e8e5022879286571af7a11ca

See more details on using hashes here.

Provenance

The following attestation bundles were made for contextpilot_ai-0.3.0.tar.gz:

Publisher: release.yml on msousa202/ContextPilot

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file contextpilot_ai-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: contextpilot_ai-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 49.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for contextpilot_ai-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 26dc59e84c5e65d7cb87031ab22296b456a9153736c49df6460c59a78337dad6
MD5 49b36594efbaebcfb118b299a09e44dc
BLAKE2b-256 cec1ef5ada63643971fc7cae8fcbd269a7243a9f41da5d556864dbf898f1265f

See more details on using hashes here.

Provenance

The following attestation bundles were made for contextpilot_ai-0.3.0-py3-none-any.whl:

Publisher: release.yml on msousa202/ContextPilot

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.4.1

2 files

0.4.0

2 files

0.3.1

2 files

This release

0.3.0 This release

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page