ContextGraph Studio MVP: repository indexing, retrieval planning, and context-pack delivery.
Project description
ContextGraph Studio
ContextGraph Studio is local context-retrieval infrastructure for AI coding agents.
Repository: https://github.com/eugenechoi98/contextgraph
It indexes a repository, builds searchable code/document chunks, and returns a structured ContextPack through CLI, FastAPI, or MCP. It is not a general RAG chatbot and it does not include a chat UI.
Default mode does not require an LLM API key, does not download an embedding model, and does not require trust_remote_code.
What It Is
ContextGraph Studio helps coding agents find the files, symbols, traces, and graph neighbors that matter for a development task.
Main entry points:
- CLI:
cgstudio index,cgstudio retrieve,cgstudio eval - API: FastAPI routes for scan, retrieve, trace, and health
- MCP: stdio tools for agent clients
Why It Exists
Large coding tasks often fail because the agent misses the right files before it starts editing. ContextGraph Studio focuses on that earlier step: building a reproducible local index and returning a compact context pack that an agent can use.
Demo
ContextGraph Studio can index its own repository and retrieve a structured ContextPack.
Try these queries after cgstudio index .:
how does the graph traversal work
BM25 candidate lane source code
how are embeddings stored and retrieved
The output highlights required files, supporting files, graph paths, and a trace_id for follow-up inspection.
Verified local eval results:
- BM25-only MRR:
0.7538 - BM25 + Graph MRR:
0.8179 - Recall@5 remained
0.8077 - Early SWE-bench Lite localization smoke:
3 / 3critical files hit
This is an early localization smoke, not the full SWE-bench Lite benchmark.
Core Capabilities
- Repository scanner with intake exclusions for eval fixtures, reports, temp files, caches, and generated outputs
- SQLite storage for files, entities, chunks, traces, eval runs, embeddings, and relations
- FTS5 BM25 retrieval
- Source-code, schema, and config candidate lanes
- Python AST parser, including module-level static assignment chunks
- TypeScript / TSX / JavaScript / JSX Tree-sitter parser
- SQL parser and JSON / YAML / TOML config parser
- Optional graph retrieval, weighted RRF fusion, and token budgeting
- Optional local vector search
- CLI, FastAPI, and MCP stdio integration
- Golden-dataset eval and lightweight SWE-bench Lite localization smoke
Architecture Overview
repository scan
-> source classification
-> parser layer
-> structure-aware chunks
-> SQLite / FTS5 / optional embeddings / graph relations
-> BM25 / optional vector / optional graph recall
-> weighted RRF
-> token budget
-> ContextPack
files, entities, and chunks are the content source of truth. chunks_fts, embeddings, relations, traces, and eval_runs are derived or supporting layers.
Quick Start
Development install from a source checkout:
Windows PowerShell:
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e .[dev]
.\.venv\Scripts\cgstudio.exe init-db
.\.venv\Scripts\cgstudio.exe index .
.\.venv\Scripts\cgstudio.exe retrieve "verify token auth flow"
.\.venv\Scripts\cgstudio.exe retrieve "verify token auth flow" --repo-id <repo-id>
cgstudio index . prints a repo_id. If you have indexed only one repository, retrieve can use the latest successful scan automatically. If you have indexed multiple repositories, pass --repo-id <repo-id> explicitly.
Linux / macOS:
python -m venv .venv
./.venv/bin/python -m pip install -e '.[dev]'
./.venv/bin/cgstudio init-db
./.venv/bin/cgstudio index .
./.venv/bin/cgstudio retrieve "verify token auth flow"
./.venv/bin/cgstudio retrieve "verify token auth flow" --repo-id <repo-id>
cgstudio index . prints a repo_id. If you have indexed only one repository, retrieve can use the latest successful scan automatically. If you have indexed multiple repositories, pass --repo-id <repo-id> explicitly.
PyPI publication is pending. Until then, release artifact validation uses a local wheel:
python -m pip install <local-wheel>
cgstudio --help
Verified local uvx smoke:
uvx --from <local-wheel> cgstudio --help
After PyPI publication, the planned command shape is:
uvx --from contextgraph-studio cgstudio --help
CLI Usage
Common commands:
.\.venv\Scripts\cgstudio.exe init-db
.\.venv\Scripts\cgstudio.exe index .
.\.venv\Scripts\cgstudio.exe retrieve "update login route token verification" --task-hint api_endpoint --max-tokens 4000
.\.venv\Scripts\cgstudio.exe get-trace <trace-id>
.\.venv\Scripts\cgstudio.exe serve --host 127.0.0.1 --port 8000
Debug commands are also available for BM25, vector, and graph search.
MCP Integration
ContextGraph Studio exposes a stdio MCP server:
.\.venv\Scripts\cgstudio.exe mcp
Supported MCP tools:
retrieve_contextscan_repoget_trace
Example client config:
Replace the command value with the absolute path to your local cgstudio executable.
MCP stdio protocol integration is tested with the official Python MCP SDK. Claude Code, Cursor, and other MCP-capable clients can use the same tools, but client-specific setup depends on your local installation.
Optional Local Vector Search
Default mode is BM25 + Graph and performs no model download:
VECTOR_INDEX_ENABLED=false
HYBRID_VECTOR_ENABLED=false
Recommended local code-specialized vector baseline:
- model:
nomic-ai/CodeRankEmbed - revision:
3c4b60807d71f79b43f3c4363786d9493691f8b1 - cache location: outside this repository
- requires explicit
EMBEDDING_TRUST_REMOTE_CODE=true
Lightweight fallback:
sentence-transformers/all-MiniLM-L6-v2
Deferred high-resource target:
nomic-ai/nomic-embed-code
Eval
Run the local golden dataset:
.\.venv\Scripts\cgstudio.exe eval --repo-id <repo-id> --dataset eval\fixtures\contextgraph_golden.json
Eval records route diagnostics so reports can distinguish requested, executed, and participating retrieval routes.
SWE-bench Lite Localization
ContextGraph Studio includes a lightweight localization smoke, not the full SWE-bench harness.
It can inspect SWE-bench Lite rows and run up to three isolated localization cases:
.\.venv\Scripts\cgstudio.exe swebench-localize `
--manifest eval\manifests\generated\swebench_lite_manifest.json `
--cache-dir <swebench-cache-dir> `
--dry-run
Boundaries:
- no patch apply
- no Docker
- no target repository tests
- no dependency install in target repositories
- no full 300-case benchmark
- GitHub checkout requires explicit
--allow-network - SWE-bench cache and per-instance SQLite DBs must stay outside tracked source paths
Latest three-instance smoke recovered critical hits for Astropy, Django, and Matplotlib after adding Python module-level assignment chunks. See eval/analysis/ for the detailed reports.
Early localization evidence:
- official SWE-bench Lite smoke sample:
3instances - critical hit rate:
3 / 3 - Astropy rank:
1 - Django rank:
7 - Matplotlib rank:
2
This is an early localization smoke, not a full SWE-bench Lite benchmark. It does not execute patches, run target tests, or use a Docker harness.
Current-machine performance evidence only:
- Astropy unchanged snapshot reuse improved from about
20.9 minto about4.4 min - relative improvement: about
79% - this is not a general SLA
Security And Data Boundaries
Default behavior:
- does not execute target repository code
- does not install target repository dependencies
- does not run target repository tests
- does not apply patches
- does not run Docker
- does not download embeddings
- does not require API keys
Sensitive config and assignment handling is key-name based. Obvious sensitive names such as password, secret, token, api_key, private_key, and credential keep the key or symbol searchable but do not write raw values into chunks, traces, or ContextPacks.
Known Limitations
- Parser coverage is intentionally minimal and static.
- TypeScript / JavaScript route detection only supports clear static Express-like patterns.
- SQL parsing covers small
CREATEstatement patterns. - Config parsing uses simple path extraction for arrays such as
servers[].host. - Graph relation types such as
uses_table,configures, andcross_lang_callsare not implemented yet. - Optional vector quality depends on explicit local model configuration.
Roadmap
Deferred items:
nomic-ai/nomic-embed-code7B local smokeuses_tableconfigurescross_lang_calls- complete 300-case SWE-bench Lite benchmark
- Docker harness
- patch apply
- target repository tests
- frontend Studio
- PyPI release
Development
Run tests:
.\.venv\Scripts\python.exe -m pytest tests
GitHub Actions CI runs a lightweight no-model validation:
- checkout
- setup Python
pip install -e .[dev]pytest tests
Before publishing changes, check that generated DBs, reports, caches, .env, third-party checkouts, and model files are not tracked.
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file contextgraph_studio-0.1.0.tar.gz.
File metadata
- Download URL: contextgraph_studio-0.1.0.tar.gz
- Upload date:
- Size: 124.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
23c18135de8cf374117e9912b8c5c2d8aeb807986e3712a3925e8af36eb15d45
|
|
| MD5 |
b66e6c765ea2bd3c23e15368e68069d9
|
|
| BLAKE2b-256 |
9095712a08e21b5771d08f90b0fcae96d63e7f20131045395cb20341261496e4
|
File details
Details for the file contextgraph_studio-0.1.0-py3-none-any.whl.
File metadata
- Download URL: contextgraph_studio-0.1.0-py3-none-any.whl
- Upload date:
- Size: 105.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
26f984a3b04e1c6b2a642b38e8096d3985ac27e013f3f3262e06385f633c9861
|
|
| MD5 |
293d8307b2db371bd7e1c6be0225cfe0
|
|
| BLAKE2b-256 |
7e04e4730ba501b07e316d2138bd8c2b9579dc31761183d5f20dcfbae33f0136
|