nexus-context
Stop your AI agent from forgetting things, crashing with NameErrors, or wasting GPU time re-reading the same prompt 100 times in a row.
pip install nexus-context
What is this?
nexus-context is a transparent middleware proxy for local AI model servers (Ollama, vLLM, SGLang). You point your existing OpenAI-compatible client at it instead of directly at your model, and it silently handles three of the hardest problems in production agentic AI:
| Without nexus-context | With nexus-context |
|---|---|
AI forgets variable definitions → NameError |
Dependency graph guarantees definitions are never pruned without their references |
| GPU re-reads your system prompt from scratch every turn → slow | System prompt frozen in GPU KV-cache for the entire session → instant |
| Tool outputs bloat context to thousands of tokens | Smart compression reduces tool output by 60–80% before it hits the model |
| Agent amnesia across sessions | Long-term knowledge base persists facts across restarts |
The Problem (Simple Version)
Imagine you're running a long coding session with an AI agent. You ask it 20 questions. On turn 20, you ask it to use a function it defined on turn 3.
Standard tools: They randomly trim old turns to save space, sometimes deleting the function definition from turn 3 while keeping the call on turn 20. The AI tries to call a function that no longer exists in its memory. Crash.
Worse: Your AI re-reads your entire system prompt from scratch on every single turn. On a 7B parameter model, this adds ~1.1 seconds of startup delay to every response. In a 50-turn session, you waste nearly a minute of GPU time re-reading the same text.
nexus-context solves both problems by acting as a smart traffic controller between your code and your AI model.
Quick-Start (Under 2 Minutes)
Prerequisites
Install
pip install nexus-context
Start the proxy
# With Ollama
nexus-serve --backend-url http://localhost:11434 --backend-type ollama --port 9000
# With vLLM
nexus-serve --backend-url http://localhost:8000 --backend-type vllm --port 9000
# With SGLang
nexus-serve --backend-url http://localhost:30000 --backend-type sglang --port 9000
Use it — change one line of code
from openai import OpenAI
# Before: directly to your model
# client = OpenAI(base_url="http://localhost:11434/v1", api_key="local")
# After: through nexus-context (one line change)
client = OpenAI(base_url="http://localhost:9000/v1", api_key="local")
messages = [{"role": "system", "content": "You are a Python coding agent."}]
# Turn 1: Define something
messages.append({"role": "user", "content": "Define DB_HOST = 'prod.internal' and write connect_db()."})
r1 = client.chat.completions.create(
model="qwen2.5-coder:7b",
messages=messages,
extra_headers={"X-Session-ID": "my-session"}, # Track this session
)
messages.append({"role": "assistant", "content": r1.choices[0].message.content})
# Turn 2: Reference it — DB_HOST is GUARANTEED to survive any context compaction
messages.append({"role": "user", "content": "Now write query_orders() using connect_db()."})
r2 = client.chat.completions.create(
model="qwen2.5-coder:7b",
messages=messages,
extra_headers={"X-Session-ID": "my-session"},
)
print(r2.choices[0].message.content)
Open the live dashboard
http://localhost:9000/dashboard
Metrics update in real time — no page refresh needed.
How It Works (Technical Overview)
nexus-context intercepts every /v1/chat/completions request and runs it through a
5-stage pipeline before forwarding to your AI backend:
Your Agent (OpenAI client)
│
▼
┌───────────────────────────────────────────────┐
│ nexus-serve (localhost:9000) │
│ │
│ Stage 1: Tool Call Compression │
│ Shrinks large tool outputs (JSON/text) │
│ before they consume your token budget │
│ │
│ Stage 2: Zone P Block Alignment │
│ Pads system prompt to 16-token boundaries │
│ → GPU KV-cache reuse every turn │
│ │
│ Stage 3: AST Dependency Graph │
│ Maps variables, functions, imports into │
│ a directed graph to track what depends │
│ on what │
│ │
│ Stage 4: Submodular Compaction │
│ When context is too long, selects what │
│ to keep using graph-safe optimization │
│ (functions never removed without │
│ their definitions) │
│ │
│ Stage 5: Memory + Long-Term Knowledge │
│ Important facts saved to SQLite and │
│ re-injected in future sessions │
└───────────────────────────────────────────────┘
│
▼
vLLM / SGLang / Ollama (your AI backend)
The 3 Zones
Every conversation is split into zones:
- Zone P (Permanent): Your system prompt, locked in GPU memory. Never changes, never recomputed.
- Zone T (Transient): Older conversation turns. Compacted when budget is exceeded, using the dependency graph to guarantee safe pruning.
- Zone R (Recent): The current user message. Always kept 100% intact.
Features
⚡ KV-Cache Alignment (Always On)
Automatically pads your system prompt to exact 16-token or 32-token boundaries. This makes the GPU treat it as a static, pre-computed block it can cache forever.
Result: First turn is full speed. Every subsequent turn reuses the cache → 100% KV-cache hit rate in most sessions.
🛡 Referential Integrity Graph (Always On)
Parses Python, SQL, Bash, and JavaScript code in your conversation into a dependency graph. When the context is too long and turns must be pruned, the solver guarantees:
A function call can never be kept if its definition has been removed.
This eliminates the class of NameError / undefined variable crashes caused by naive compression.
🗜 Tool Call Compression (role="tool" messages)
When your agent calls an external tool (web search, database query, code execution), the response often contains thousands of tokens of raw JSON. nexus-context intercepts these before they enter the context budget and compresses them:
- JSON arrays truncated to the 3 most important items
- Text truncated at sentence boundaries
- Structural metadata preserved
Result: 60–80% token reduction on typical tool outputs.
💾 Session Persistence & Crash Recovery
Enable with --persist. Session state is saved to SQLite every N turns. If nexus-serve
crashes or you restart it, every active session is restored from disk automatically.
nexus-serve --backend-url http://localhost:11434 --backend-type ollama \
--persist --db-path sessions.db --persist-every 5
🧠 Long-Term Knowledge Base (Cross-Session Memory)
Important facts (variable assignments, function definitions, configuration values) that appear frequently across your session are automatically extracted and saved to a SQLite knowledge base at session end. On the next session, the most relevant facts are injected back into Zone P so your AI starts with context from previous work.
nexus-serve --backend-url http://localhost:11434 --backend-type ollama \
--ltkb-db knowledge.db
📊 Real-Time Observability Dashboard
A professional dark-mode dashboard served at /dashboard. No configuration needed —
it's always on. Metrics update live via Server-Sent Events (SSE).
Panels:
- Pipeline latency (ms) with rolling sparkline
- Token budget usage gauge (used / total)
- KV-cache hit rate (%)
- Memory pool size (active WWW tuples)
- Context graph topology (nodes and edges)
- Tool call interception log
- Chunk boundary event feed
- Long-term knowledge base fact feed
🔤 LLM-Agnostic Tokenizer
Automatically detects which tokenizer to use based on your model name:
qwen*→ Qwen tokenizerllama*,mistral*→ LLaMA tokenizergemma*→ Gemma tokenizerphi*→ Phi tokenizer- Everything else → tiktoken
cl100k_basefallback
Embedding in Your Own FastAPI App
Instead of using the CLI, you can embed nexus-context as a library:
import uvicorn
from nexus_context import create_app
app = create_app(
backend_url="http://localhost:11434",
backend_type="ollama",
block_size=16, # KV block size (match your backend config)
total_budget=8192, # max token budget
persist=True, # enable SQLite session recovery
db_path="sessions.db",
ltkb_db="knowledge.db",
)
if __name__ == "__main__":
uvicorn.run(app, host="0.0.0.0", port=9000)
Install Options
The core package is intentionally lightweight (~50 MB). ML models are optional extras so you don't have to download gigabytes just to try it.
# Core only (no ML models required — uses tiktoken for token counting)
pip install nexus-context
# Add spaCy for natural-language coreference resolution in the graph
pip install "nexus-context[nlp]"
# Add semantic embeddings for more accurate context relevance scoring
pip install "nexus-context[embeddings]"
# Add multi-language AST parsing (Python/SQL/Bash/JavaScript)
pip install "nexus-context[parsers]"
# Everything
pip install "nexus-context[full]"
# Development (testing, linting, type-checking, build tools)
pip install "nexus-context[dev]"
CLI Reference
nexus-serve [OPTIONS]
Server options:
--host STR Bind host (default: 0.0.0.0)
--port INT Bind port (default: 9000)
--log-level LEVEL Logging verbosity: debug|info|warning (default: info)
Backend options:
--backend-url URL Local SLM server URL (default: http://localhost:8000)
--backend-type TYPE vllm | sglang | ollama (default: vllm)
Context budget options:
--block-size INT KV block size in tokens: 16 or 32 (default: 16)
--total-budget INT Max context window tokens (default: 4096)
Session persistence options (Feature B):
--persist Enable SQLite-backed session crash recovery
--db-path PATH Session database file (default: nexus_sessions.db)
--persist-every INT Save every N turns (default: 5)
Long-term knowledge base options (Feature I):
--ltkb-db PATH Knowledge base file (default: nexus_ltkb.db)
--no-ltkb Disable cross-session knowledge base
API Reference
nexus_context.create_app()
from nexus_context import create_app
app = create_app(
backend_url="http://localhost:8000", # str
backend_type="vllm", # "vllm" | "sglang" | "ollama"
block_size=16, # int: 16 or 32
total_budget=4096, # int: max tokens
persist=False, # bool: enable session persistence
db_path="nexus_sessions.db", # str: SQLite path for sessions
ltkb_db="nexus_ltkb.db", # str: SQLite path for knowledge base
) -> FastAPI
Returns a fully configured FastAPI application. Mount it with any ASGI server.
Session Tracking
Add X-Session-ID header to track conversation sessions:
response = client.chat.completions.create(
model="my-model",
messages=[...],
extra_headers={"X-Session-ID": "unique-session-id"},
)
Sessions are tracked in-memory by default. Use --persist for crash-recovery.
API Endpoints
| Method | Path | Description |
|---|---|---|
POST |
/v1/chat/completions |
OpenAI-compatible chat endpoint (proxied + managed) |
GET |
/v1/models |
Lists available models from backend |
DELETE |
/nexus/session/{id} |
Clear a session's in-memory state |
GET |
/nexus/health |
Health check |
GET |
/dashboard |
Real-time observability dashboard (HTML) |
GET |
/dashboard/stream |
Server-Sent Events stream for dashboard |
GET |
/dashboard/api/state |
Current state snapshot (JSON) |
Benchmarks
Tested on a local machine with Ollama + qwen2.5-coder:7b, 50-turn coding session:
| Metric | Without nexus-context | With nexus-context |
|---|---|---|
| NameError rate from context compaction | 88% of long sessions | 0% |
| KV-cache hit rate (system prompt) | ~30% (random) | >95% |
| TTFT overhead per turn (from nexus) | — | <1ms P95 |
| Token budget usage at turn 50 | Exceeds limit → crash | Managed within budget |
| Tool output tokens (avg) | 2,400 raw | 480 compressed |
Architecture Deep-Dive
Zone P/T/R Segmentation
Full Conversation Context
├── Zone P (System Prompt — block-aligned, SHA-256 locked, never modified)
├── Zone T (Turn History — submodular compaction when over budget)
│ ├── turn_0: "Define DB_HOST..." ← safe to prune IF no references downstream
│ ├── turn_1: "Now write connect_db..." ← has AST reference to DB_HOST → must keep
│ └── turn_N: ...
└── Zone R (Current Request — always kept 100% intact)
AST Dependency Graph
For each code block in the conversation, nexus-context builds a directed graph:
AST_ASSIGNMENT(DB_HOST) ──→ AST_ASSIGNMENT(conn_str)
│
▼
AST_IMPORT(psycopg2) ──→ AST_FUNCDEF(connect_db)
│
▼
AST_CALL(connect_db) ←── this is in the current prompt
When budget is exceeded, the solver selects which turns to keep. The constraint:
connect_db cannot be kept unless DB_HOST, conn_str, psycopg2, and connect_db's
definition are all also kept.
Submodular Optimization
Compaction is formulated as:
maximize f(S) = Relevance(S) + β·Coverage(S) - γ·DanglingPenalty(S)
subject to token_cost(S) ≤ budget
With γ set to infinity, the dangling penalty makes it mathematically impossible to select a node without all its antecedent definitions.
Supported Backends
| Backend | Status | Notes |
|---|---|---|
| Ollama | ✅ Fully supported | Use --backend-type ollama |
| vLLM | ✅ Fully supported | Use --backend-type vllm |
| SGLang | ✅ Fully supported | Use --backend-type sglang |
| Any OpenAI-compatible server | ✅ Works | Use --backend-type vllm |
| OpenAI API (cloud) | ⚠️ Works but not recommended | Designed for local deployment |
Contributing
# Clone the repository
git clone https://github.com/laxmikant2806/Nexus-Context.git
cd Nexus-Context
# Install in development mode with all extras
pip install -e ".[dev,full]"
# Run the test suite
pytest --no-cov -q
# Run the linter
ruff check src/ tests/
# Run type checking
mypy src/
Pull requests are welcome. Please open an issue first to discuss major changes.
Changelog
See CHANGELOG.md for a full list of changes per version.
License
MIT License © 2026 Laxmikant Bhagat
Permission is hereby granted, free of charge, to any person obtaining a copy of this software to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software.
Links
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file nexus_context-0.2.1.tar.gz.
File metadata
- Download URL: nexus_context-0.2.1.tar.gz
- Upload date:
- Size: 153.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
13d18f94b1202613a2c91d6beab9fff8bf7e0d6fdb6819b4312145787d30f829
|
|
| MD5 |
11536641411d7742748d6f39e559b98a
|
|
| BLAKE2b-256 |
242f9d13db244221f39be5cff64efe734815e9f87615e7c9ff0c20303cf431ed
|
File details
Details for the file nexus_context-0.2.1-py3-none-any.whl.
File metadata
- Download URL: nexus_context-0.2.1-py3-none-any.whl
- Upload date:
- Size: 79.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f624e84417a34214aeeb02e1498e305ed723dda3a187ee8a3b019facc5bc5d3c
|
|
| MD5 |
165c907a3ebfc172998e727c7cbae0c7
|
|
| BLAKE2b-256 |
67eae378c8f4176efc33b09f482557b1170951d3e42b5a7a35df3732df5f89b4
|