Skip to main content

The cgh CLI landing screen: banner, command list, and examples

PyPI version Python versions npx CI License: MIT and CC BY-NC-SA

Local code graph, shared memory and guardrails for AI coding assistants.

Parses your repo into a graph of files, functions, classes, Terraform resources, and Markdown documentation -- then exposes it as an MCP server so Claude Code, Cursor, Codex, Gemini, and IBM Bob can do symbol-level lookups instead of reading entire files. On top of the graph: a knowledge and session memory every connected agent shares, and a confidentiality layer (findings, egress gate, per-agent guard hooks) that decides what an agent may read and what may reach a cloud model.

Result: 40-60% fewer context tokens on typical navigation tasks, learnings that survive context clears, and nothing leaving the machine without a gate.

pip install cgh && cgh init && cgh serve

Why cgh: the measured gains

cgh's job is to keep an agent's working context small and its round-trips few. It answers code questions from the graph, returning exact file:line, instead of the agent reading whole files or grepping. That shows up as fewer context tokens and fewer turns, at equal correctness.

The figures below come from a two-arm benchmark: the same tasks run twice against the same repo, once with cgh available and once with Read/Grep only, scored on each session's token usage, turn count, and answer correctness. Cost is compared only across tasks both arms got right, so a cheap wrong answer never reads as a saving.

On code-navigation tasks: about 40 to 60% fewer context tokens and 20 to 40% fewer turns, correctness unchanged.

The gap is widest on multi-file questions, where a graph beats text search. For "what breaks if I change _backend?", the agent has to follow call edges across a module:

turns context tokens
Read / Grep 13 2607
cgh 6 886

That question is one command:

cgh callers: the call graph of resolve_import rendered as a tree with exact file and line

Delegating writes. The cgh-codegen plugin does the same for writes: predictable, pattern-following code (tests, stubs, config, boilerplate) is handed to a cheap or local model that mirrors an existing file, and the reference never enters the primary model's context. cgh picks the file to mirror from the graph, so selecting it costs zero model tokens.

task reference size primary-model tokens (without -> with)
generate an auth-stripping test suite 108KB test file ~27,980 -> ~100 (99.6%)
generate a models.pyi type stub 41KB module ~14,000 -> ~100 (99.3%)

Per task, what the agent does instead of reading files:

Task Without cgh With cgh
Find where process_data is defined Read 3-5 files (~2,000 tokens) symbol_lookup (< 50 tokens)
Find all callers of save_record Read every candidate file find_callers (< 50 tokens)
Understand blast radius of utils.py Read imports manually subgraph (< 100 tokens)
Find docs about reconciliation Read all .md files search_docs (< 50 tokens)
Build context for a task 5-10 file reads (~5,000 tokens) context_for_task (< 200 tokens)

What this is not. The billed cost, once the model's prompt cache is counted, is roughly a wash on the read side: the cache dominates the invoice, so fewer turns do not cut it much. cgh's gain there is a smaller working context and fewer turns, not a smaller bill. On a trivial one-file edit cgh adds nothing. Run-to-run variance is real (around 20%), so read these as ratios over a task set rather than a single guaranteed number.


Install

pip install cgh                 # or: pipx install cgh / uv tool install cgh
pip install "cgh[full]"         # plugins, extra language parsers, precise Python calls

No Python? Run the standalone binary through npm, or download it from the latest release:

npx @altikva/cgh serve          # fetches the binary for your OS, verifies it, runs it
npx @altikva/cgh --egress serve # the egress build, with the model-calling plugins

The binary uses the SQLite backend; uvx cgh bundles DuckDB and every plugin. One-line installers for macOS, Linux, WSL, Git Bash and Windows PowerShell, corporate mirror settings, optional extras and the cgh: command not found fix are in docs/INSTALL.md. Python 3.11 through 3.14.

Quick start

# 1. Initialize (interactive wizard)
cgh init

# 2. Build the graph
cgh index

# 3. Check what was indexed
cgh stats

# 4. Start the MCP server for your AI tool
cgh serve --watch --reindex

cgh status tells you what the graph holds and whether it still matches the working tree:

cgh status: backend, owner, scan freshness, import coverage and file count

How it works

AI Assistant (Claude / Cursor / Codex / Gemini / IBM Bob)
    |  symbol_lookup("process_data")
    |  search_docs("reconciliation")
    |  context_for_task("fix auth bug")
    v
MCP server (codegraph)          <-- stdio, no network
    |  SQL graph query + BM25 FTS
    v
DuckDB graph DB (.codegraph/graph.duckdb)   <-- embedded, file-based
SQLite FTS5 (.codegraph/fts.db)       <-- BM25 full-text search
    |  indexed from
    v
Your source files (.py / .ts / .java / .go / .rs / .vue / .tf / .md)
    ^
File watcher (watchdog)         <-- live incremental updates on save

Instead of reading services.py (800 tokens) to find where verify_token is defined, your AI calls symbol_lookup("verify_token") and gets back the file, the line range, the kind and the docstring, then reads only those lines.

Languages

Six languages get a full symbol graph, functions and classes and the calls between them, with their imports resolved into File -> File edges:

imports resolved
Python .py .pyw yes
TypeScript, JavaScript .ts .tsx .js .jsx .mjs .cjs yes
Java .java yes
Go .go yes
Rust .rs yes
Vue SFC .vue yes

C# and Ruby parse too, behind pip install "cgh[langs]".

Alongside them, Markdown becomes a heading tree with its links and code references, Terraform yields resources, variables and outputs, and JSON, TOML, YAML and SQL expose their structure as sections. Every one of them is searchable and reachable from the graph.


Documentation

Guide What it covers
Install one-line installers, extras, corporate mirrors, PATH
CLI reference every verb and flag
Configuration config.toml, environment variables, .cghignore
MCP tools the tools your agent calls, by category
Integrations Claude Code, Cursor, Codex, Gemini, IBM Bob
Federation one parent repo querying its sub-repos read-only
Session memory knowledge and plans that survive a context clear
Security findings, secure mode, the guard, the MCP auth key
Plugins installing them, disabling them, writing one
Parsers the parser interface and how to add a language
Graph schema the nodes and edges the index holds
Embedding (SDK) using cgh as a library

Limitations

  • CALLS resolution is name-based by default. A call is linked to a same-file function of that name, falling back to all repo functions with that name only when there is no same-file match, so cross-file call edges are best-effort. For Python you can opt into precise cross-file resolution with pip install cgh[lsp] and precise_calls = true (jedi-backed); other languages stay name-based.
  • Terraform HCL uses regex, not a full grammar. Complex meta-arguments may be missed.
  • Imports resolve to files in your repo, never to dependencies. Python, JS/TS, Vue, Java, Go and Rust each map an import onto the file it names, following that language's own layout rules. Anything outside the repo stays unresolved on purpose: the standard library, a Go module you do not own, an external crate or npm package gets no node and no edge, because inventing one would be a lie about your code. Cross-repo edges are not inferred either, each federated scope is canonical for its own files.
  • Markdown code refs are heuristic. PascalCase and snake_case patterns are matched, so a ref can be a false positive.
  • Large repos take minutes to index. Incremental updates stay fast (well under a second per changed file), and a pull or merge reindexes only the changed files via the git hooks.

License

Dual-licensed under MIT and CC BY-NC-SA 4.0: both licenses apply together and you must comply with both. In practice that means non-commercial use, share-alike derivatives, attribution, and no warranty. Copyright (c) 2026 ALTIKVA. See LICENSE or the canonical notice at https://www.altikva.com/licenses/LICENSE-1.0.

Plugin exception: a plugin that talks to cgh only through the documented plugin interfaces (the cgh entry-point group and the public plugin API) is not treated as a derivative work and may be licensed under any terms its author chooses, including commercial ones. Using cgh itself stays under the dual license whatever plugins are installed. Full wording in LICENSE.

Metadata

Release files for cgh 0.14.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cgh 0.14.3
File Size Uploaded
cgh-0.14.3.tar.gz 394.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cgh 0.14.3
File Interpreter ABI Platform
cgh-0.14.3-py3-none-any.whl Python 3 none any Details

Total release size: 867.2 kB

Release files / cgh-0.14.3.tar.gz

Download URL cgh-0.14.3.tar.gz
Size 394.5 kB
Tags Source
SHA-256 checksum
How to use checksums
7b6442809e6f133525164f421a644b381cb30d4db6c104c7077b90690e259e34
BLAKE2b-256 checksum
How to use checksums
0ad7966ce1dc843b470d33b74a286af235600ba71d78cc1bc0f6dba5e180c988
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / cgh-0.14.3-py3-none-any.whl

Download URL cgh-0.14.3-py3-none-any.whl
Size 472.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
31749d3b78345eafa764c5cbb0eae130228171b12145fcf8f4a48f6696afa877
BLAKE2b-256 checksum
How to use checksums
5b1e325aa719214c79262b07fe82e680749bd9ebb638a2d517b5a1783efa1978
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page