Context Guard
Dual-layer Context Health Guardrail and Adaptive State-Ledger Compression Engine for LLM Agent Harnesses.
1. Executive Summary
Autonomous LLM agents and multi-turn coding assistants frequently degrade over extended sessions due to four critical context failure modes:
- Context Poisoning & Error Propagation: The agent hallucinates or makes a mistake, the user issues a correction (e.g., "No, don't use Flask, use FastAPI"), but subsequent turns keep referencing the poisoned context and repeating the rejected pattern.
- Distraction & Echo Loops: Consecutive turns suffer from repetitive phrasing, circular logic, or verbatim code generation loops that consume context budget without advancing the task.
- Directive Clash & Contradiction: System directives or pinned user constraints are gradually forgotten or violated as context windows fill up.
- Noise Ballooning & Context Rot: High cumulative context token payloads are carried across low-information single-word queries (e.g., "ok", "continue", "next"), diluting model attention.
Context Guard solves this through a dual-layer architecture:
- Layer 1 (Deterministic Heuristic Evaluator): A zero-latency, rule-based health evaluator that scores context degradation between 0 and 100 (
🟢 Healthy,🟡 Degraded,🔴 Critical). - Layer 2 (Adaptive State-Ledger Compressor): Strips ANSI noise, deep repetitive stack traces, and pleasantries, consolidating older conversation turns into a structured Active State Ledger while keeping recent raw turns intact.
Context Guard can be deployed as:
- A FastAPI OpenAI-compatible reverse proxy (
/v1/chat/completions) with live streaming support. - A Model Context Protocol (MCP) server for Cursor, Claude Code, and AGY.
- A Python SDK library embedded directly into LangGraph, AutoGen, or custom agent loops.
2. Architecture Flow
Client Request (OpenAI Payload)
│
▼
┌───────────────────────────────────────┐
│ Context Guard Interceptor │
└───────────────────┬───────────────────┘
│
▼
┌───────────────────────────────────────┐
│ Layer 1: Deterministic Evaluator │
│ - Repetition / Echo Loops (+35) │
│ - Error Propagation / Poisoning (+45)│
│ - Directive Clashes (+40) │
│ - Noise Ballooning (+25) │
└───────────────────┬───────────────────┘
│
┌─────────────────┴─────────────────┐
▼ ▼
[Score < 25 (GREEN)] [Score >= 25 (YELLOW/RED)]
│ │
│ Passes through ▼
│ unmodified ┌───────────────────────────┐
│ │ Layer 2: Adaptive │
│ │ State-Ledger Compressor │
│ │ - Noise & Trace Stripping │
│ │ - State Ledger Markdown │
│ │ - Preserves N Raw Turns │
│ │ - Intervention Directive │
│ └─────────────┬─────────────┘
│ │
└─────────────────┬─────────────────┘
▼
Forwarded Upstream Request
(e.g., OpenAI, Anthropic, vLLM, Ollama)
│
▼
Telemetry Headers Injected on Response:
- X-Context-Health-Status: 🟢 Healthy | 🟡 Degraded | 🔴 Critical
- X-Context-Penalty-Score: 0 - 100
- X-Context-Tokens-Saved: N tokens
3. Quickstart Guide
Installation
# Clone the repository
git clone https://github.com/your-org/context-guard.git
cd context-guard
# Install dependencies using uv
uv sync --all-extras
Option A: Running as an MCP Server
Context Guard natively implements the Model Context Protocol (MCP), exposing context inspection and ledger generation tools directly to IDEs like Cursor, Claude Code, and Gemini CLI.
Tools Provided:
inspect_context_health: Evaluates conversation messages for failure modes and returns a standardizedHealthReport.compress_context_buffer: Automatically compresses bloated older context into a State Ledger.generate_state_ledger: Extracts goals, hard constraints, active variables, and pending tasks.context_audit: Prompt template instructing an LLM to self-audit its context health.
Cursor Configuration (~/.cursor/mcp.json or project .cursor/mcp.json):
{
"mcpServers": {
"context-guard": {
"command": "uv",
"args": ["run", "--directory", "/absolute/path/to/context-guard", "context-guard-mcp"]
}
}
}
Claude Desktop Configuration (claude_desktop_config.json):
{
"mcpServers": {
"context-guard": {
"command": "uv",
"args": ["run", "--directory", "/path/to/context-guard", "context-guard-mcp"]
}
}
}
Option B: Running as a Drop-in Reverse Proxy
Run Context Guard in front of Groq, OpenRouter, OpenAI, vLLM, Ollama, or any OpenAI-compatible API. By default, it routes to Groq (https://api.groq.com/openai/v1), allowing you to use high-throughput Groq API keys (gsk_...) and models like llama-3.1-8b-instant or llama-3.3-70b-versatile.
Point your agent harness or client's base_url to http://localhost:8080/v1.
Running via CLI:
# Configure Groq (or override with any OpenAI-compatible upstream)
export OPENAI_BASE_URL="https://api.groq.com/openai/v1"
export GROQ_API_KEY="gsk_..."
# Launch the proxy (listening on 0.0.0.0:8080)
uv run context-guard
Running via Docker / Docker Compose:
# Launch container
docker compose up -d
# Verify health check
curl http://localhost:8080/health
# {"status":"ok","service":"context-guard"}
Client Configuration Example:
from openai import OpenAI
# Point client to Context Guard reverse proxy
client = OpenAI(
base_url="http://localhost:8080/v1",
api_key="gsk_...", # forwarded intact to Groq
)
response = client.chat.completions.create(
model="llama-3.1-8b-instant",
messages=[
{"role": "user", "content": "Build an API using FastAPI."},
{"role": "assistant", "content": "Here is the implementation..."},
{"role": "user", "content": "No, do not use Pydantic v1, use v2."},
],
)
Option C: Python Library SDK Integration
You can integrate Context Guard directly into custom agent loops or LangGraph orchestrators:
import asyncio
from context_guard.evaluators import DeterministicEvaluator
from context_guard.compressors import ContextCompressor
async def main():
messages = [
{"role": "user", "content": "Goal: Build high-throughput event pipeline with Redis."},
{"role": "assistant", "content": "Setting up consumer group... " * 15},
{"role": "user", "content": "Never use blocking keys."},
{"role": "assistant", "content": "Updated code without blocking keys... " * 15},
{"role": "user", "content": "What is our current memory consumption?"},
]
# 1. Evaluate context health
evaluator = DeterministicEvaluator()
report = evaluator.evaluate(messages, pinned_constraints=["Never use blocking keys"])
print(f"Status: {report.status.value} (Score: {report.penalty_score}/100)")
print(f"Action: {report.recommended_action}")
# 2. Compress context if degraded
if report.penalty_score >= 25:
compressor = ContextCompressor(preserve_recent_turns=3)
compressed = await compressor.compress(messages)
print(f"Reduced tokens by: {compressed.token_reduction_pct}%")
print("\nGenerated State Ledger:")
print(compressed.state_ledger_markdown)
# 3. Format final payload for model inference
next_payload = compressor.format_for_inference(compressed)
asyncio.run(main())
4. Telemetry Header Reference
Every response forwarded through the reverse proxy includes real-time telemetry headers:
| Header | Example Value | Description |
|---|---|---|
X-Context-Health-Status |
🟢 Healthy🟡 Degraded🔴 Critical |
Categorical health status of the conversation history prior to upstream forwarding. |
X-Context-Penalty-Score |
0 to 100 |
Cumulative penalty score calculated across all detected failure modes. |
X-Context-Tokens-Saved |
384 |
Estimated token count saved by noise stripping and State Ledger compression ($0$ if uncompressed). |
5. Development & Testing
# Run full test suite (evaluators, compressors, proxy, mcp)
uv run pytest
# Check formatting and style
uv run ruff check .
uv run ruff format --check .
# Auto-fix formatting
uv run ruff format .
6. License
Context Guard is open-source software licensed under the Apache-2.0 License.
Metadata
Release files for context-guards 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| context_guards-0.1.0.tar.gz | 117.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| context_guards-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 149.2 kB
Release files / context_guards-0.1.0.tar.gz
| Download URL | context_guards-0.1.0.tar.gz |
|---|---|
| Size | 117.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
01dfbb9ec07aab6aa754d7f271bdbdb0d45d1998b413ee5f4dd315f87d9e5bde
|
|
BLAKE2b-256 checksum How to use checksums |
716eaa8962da298257e327bd551758f6cab3ca7cc6009047b15b77670dc12d60
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / context_guards-0.1.0-py3-none-any.whl
| Download URL | context_guards-0.1.0-py3-none-any.whl |
|---|---|
| Size | 31.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
16f3165dc3561623a9558dcec073a0a965926c26f7f9b756228fdf929d82a1ad
|
|
BLAKE2b-256 checksum How to use checksums |
e1d6f50017dec98734f355b757d3b2d80146c73ac3a6adbe7c7f32cdd15fc42e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log