Skip to main content

Context Guard

Python 3.10+ FastAPI Model Context Protocol License: Apache-2.0

Dual-layer Context Health Guardrail and Adaptive State-Ledger Compression Engine for LLM Agent Harnesses.


1. Executive Summary

Autonomous LLM agents and multi-turn coding assistants frequently degrade over extended sessions due to four critical context failure modes:

  1. Context Poisoning & Error Propagation: The agent hallucinates or makes a mistake, the user issues a correction (e.g., "No, don't use Flask, use FastAPI"), but subsequent turns keep referencing the poisoned context and repeating the rejected pattern.
  2. Distraction & Echo Loops: Consecutive turns suffer from repetitive phrasing, circular logic, or verbatim code generation loops that consume context budget without advancing the task.
  3. Directive Clash & Contradiction: System directives or pinned user constraints are gradually forgotten or violated as context windows fill up.
  4. Noise Ballooning & Context Rot: High cumulative context token payloads are carried across low-information single-word queries (e.g., "ok", "continue", "next"), diluting model attention.

Context Guard solves this through a dual-layer architecture:

  • Layer 1 (Deterministic Heuristic Evaluator): A zero-latency, rule-based health evaluator that scores context degradation between 0 and 100 (🟢 Healthy, 🟡 Degraded, 🔴 Critical).
  • Layer 2 (Adaptive State-Ledger Compressor): Strips ANSI noise, deep repetitive stack traces, and pleasantries, consolidating older conversation turns into a structured Active State Ledger while keeping recent raw turns intact.

Context Guard can be deployed as:

  • A FastAPI OpenAI-compatible reverse proxy (/v1/chat/completions) with live streaming support.
  • A Model Context Protocol (MCP) server for Cursor, Claude Code, and AGY.
  • A Python SDK library embedded directly into LangGraph, AutoGen, or custom agent loops.

2. Architecture Flow

                      Client Request (OpenAI Payload)
                                    │
                                    ▼
                ┌───────────────────────────────────────┐
                │       Context Guard Interceptor       │
                └───────────────────┬───────────────────┘
                                    │
                                    ▼
                ┌───────────────────────────────────────┐
                │   Layer 1: Deterministic Evaluator    │
                │  - Repetition / Echo Loops (+35)      │
                │  - Error Propagation / Poisoning (+45)│
                │  - Directive Clashes (+40)            │
                │  - Noise Ballooning (+25)             │
                └───────────────────┬───────────────────┘
                                    │
                  ┌─────────────────┴─────────────────┐
                  ▼                                   ▼
        [Score < 25 (GREEN)]                [Score >= 25 (YELLOW/RED)]
                  │                                   │
                  │ Passes through                    ▼
                  │ unmodified          ┌───────────────────────────┐
                  │                     │ Layer 2: Adaptive         │
                  │                     │ State-Ledger Compressor   │
                  │                     │ - Noise & Trace Stripping │
                  │                     │ - State Ledger Markdown   │
                  │                     │ - Preserves N Raw Turns   │
                  │                     │ - Intervention Directive  │
                  │                     └─────────────┬─────────────┘
                  │                                   │
                  └─────────────────┬─────────────────┘
                                    ▼
                     Forwarded Upstream Request
               (e.g., OpenAI, Anthropic, vLLM, Ollama)
                                    │
                                    ▼
               Telemetry Headers Injected on Response:
               - X-Context-Health-Status: 🟢 Healthy | 🟡 Degraded | 🔴 Critical
               - X-Context-Penalty-Score: 0 - 100
               - X-Context-Tokens-Saved: N tokens

3. Quickstart Guide

Installation

# Clone the repository
git clone https://github.com/your-org/context-guard.git
cd context-guard

# Install dependencies using uv
uv sync --all-extras

Option A: Running as an MCP Server

Context Guard natively implements the Model Context Protocol (MCP), exposing context inspection and ledger generation tools directly to IDEs like Cursor, Claude Code, and Gemini CLI.

Tools Provided:

  • inspect_context_health: Evaluates conversation messages for failure modes and returns a standardized HealthReport.
  • compress_context_buffer: Automatically compresses bloated older context into a State Ledger.
  • generate_state_ledger: Extracts goals, hard constraints, active variables, and pending tasks.
  • context_audit: Prompt template instructing an LLM to self-audit its context health.

Cursor Configuration (~/.cursor/mcp.json or project .cursor/mcp.json):

{
  "mcpServers": {
    "context-guard": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/context-guard", "context-guard-mcp"]
    }
  }
}

Claude Desktop Configuration (claude_desktop_config.json):

{
  "mcpServers": {
    "context-guard": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/context-guard", "context-guard-mcp"]
    }
  }
}

Option B: Running as a Drop-in Reverse Proxy

Run Context Guard in front of Groq, OpenRouter, OpenAI, vLLM, Ollama, or any OpenAI-compatible API. By default, it routes to Groq (https://api.groq.com/openai/v1), allowing you to use high-throughput Groq API keys (gsk_...) and models like llama-3.1-8b-instant or llama-3.3-70b-versatile.

Point your agent harness or client's base_url to http://localhost:8080/v1.

Running via CLI:

# Configure Groq (or override with any OpenAI-compatible upstream)
export OPENAI_BASE_URL="https://api.groq.com/openai/v1"
export GROQ_API_KEY="gsk_..."

# Launch the proxy (listening on 0.0.0.0:8080)
uv run context-guard

Running via Docker / Docker Compose:

# Launch container
docker compose up -d

# Verify health check
curl http://localhost:8080/health
# {"status":"ok","service":"context-guard"}

Client Configuration Example:

from openai import OpenAI

# Point client to Context Guard reverse proxy
client = OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="gsk_...",  # forwarded intact to Groq
)

response = client.chat.completions.create(
    model="llama-3.1-8b-instant",
    messages=[
        {"role": "user", "content": "Build an API using FastAPI."},
        {"role": "assistant", "content": "Here is the implementation..."},
        {"role": "user", "content": "No, do not use Pydantic v1, use v2."},
    ],
)

Option C: Python Library SDK Integration

You can integrate Context Guard directly into custom agent loops or LangGraph orchestrators:

import asyncio
from context_guard.evaluators import DeterministicEvaluator
from context_guard.compressors import ContextCompressor


async def main():
    messages = [
        {"role": "user", "content": "Goal: Build high-throughput event pipeline with Redis."},
        {"role": "assistant", "content": "Setting up consumer group... " * 15},
        {"role": "user", "content": "Never use blocking keys."},
        {"role": "assistant", "content": "Updated code without blocking keys... " * 15},
        {"role": "user", "content": "What is our current memory consumption?"},
    ]

    # 1. Evaluate context health
    evaluator = DeterministicEvaluator()
    report = evaluator.evaluate(messages, pinned_constraints=["Never use blocking keys"])
    print(f"Status: {report.status.value} (Score: {report.penalty_score}/100)")
    print(f"Action: {report.recommended_action}")

    # 2. Compress context if degraded
    if report.penalty_score >= 25:
        compressor = ContextCompressor(preserve_recent_turns=3)
        compressed = await compressor.compress(messages)
        print(f"Reduced tokens by: {compressed.token_reduction_pct}%")
        print("\nGenerated State Ledger:")
        print(compressed.state_ledger_markdown)

        # 3. Format final payload for model inference
        next_payload = compressor.format_for_inference(compressed)


asyncio.run(main())

4. Telemetry Header Reference

Every response forwarded through the reverse proxy includes real-time telemetry headers:

Header Example Value Description
X-Context-Health-Status 🟢 Healthy
🟡 Degraded
🔴 Critical
Categorical health status of the conversation history prior to upstream forwarding.
X-Context-Penalty-Score 0 to 100 Cumulative penalty score calculated across all detected failure modes.
X-Context-Tokens-Saved 384 Estimated token count saved by noise stripping and State Ledger compression ($0$ if uncompressed).

5. Development & Testing

# Run full test suite (evaluators, compressors, proxy, mcp)
uv run pytest

# Check formatting and style
uv run ruff check .
uv run ruff format --check .

# Auto-fix formatting
uv run ruff format .

6. License

Context Guard is open-source software licensed under the Apache-2.0 License.

Metadata

Release files for context-guards 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for context-guards 0.1.0
File Size Uploaded
context_guards-0.1.0.tar.gz 117.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for context-guards 0.1.0
File Interpreter ABI Platform
context_guards-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 149.2 kB

Release files / context_guards-0.1.0.tar.gz

Download URL context_guards-0.1.0.tar.gz
Size 117.3 kB
Tags Source
SHA-256 checksum
How to use checksums
01dfbb9ec07aab6aa754d7f271bdbdb0d45d1998b413ee5f4dd315f87d9e5bde
BLAKE2b-256 checksum
How to use checksums
716eaa8962da298257e327bd551758f6cab3ca7cc6009047b15b77670dc12d60
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / context_guards-0.1.0-py3-none-any.whl

Download URL context_guards-0.1.0-py3-none-any.whl
Size 31.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
16f3165dc3561623a9558dcec073a0a965926c26f7f9b756228fdf929d82a1ad
BLAKE2b-256 checksum
How to use checksums
e1d6f50017dec98734f355b757d3b2d80146c73ac3a6adbe7c7f32cdd15fc42e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page