Skip to main content

megabrain

megabrain

Ask a codebase a question. Get the exact code back.

PyPI MIT No LLM in the retrieval path MCP ready Studio web UI


Point megabrain at a repo and ask "how does auth work" in plain English. It finds all the related code — in ~200 ms, using no LLM, just math on embeddings — and an LLM narrates a walkthrough with the real code spliced in from disk. Nothing is invented: every line shown is copied verbatim.

Use it from the terminal, as an MCP server inside Claude Code, Codex, Antigravity, Cursor, Windsurf or Gemini CLI (megabrain install wires up whichever you have), as a Python library, or as a full local web app.

🖥️ megabrain studio — the whole engine, in your browser

One command turns megabrain into a local studio — nothing canned, every pixel driven by the live engine:

megabrain serve ~/repo        #  → open http://localhost:2134
  • Search — every related file ranked in ~200 ms; click one for a chunk heatmap where signal glows and noise dims, code syntax-highlighted.
  • Prune — the money shot: what the engine read vs what it ignored, side by side.
  • Ask — watch a broad question fan out into parallel sub-agents, their tool calls and prose streaming into per-agent cards, then a synthesis with the real code spliced in as it types.
  • Providers, live — Claude SDK · OpenRouter · Ollama, auto-detected. Switch the narrator without leaving the page, pick the model, and start ollama serve in one click to go fully local.
  • Add a repo → it scans first — you SEE exactly what will index and what's skipped and why (.gitignore · vendored · generated · too-big), edit the .megabrainignore, then a live progress bar indexes it file by file.
  • Embeddings you can see — which model each index used, and re-index with another (cloud pplx or a local, code-tuned jina) behind the same bar — the query embedding switches to match, so search keeps working.

Keyboard-driven, dark/light, zero build step, no CDN. Want the JSON API without the UI? megabrain serve-api ~/repo mounts the exact same endpoints, no studio.

Quickstart — the easy path, no API keys

Everything runs on your machine: ask narrates on your Claude Code subscription, embeddings run locally on Ollama. No cloud keys.

pip install 'megabrain[claude]'                      # engine + Claude Code narration

ollama pull unclemusclez/jina-embeddings-v2-base-code # local, code-tuned embeddings, one time
export MEGABRAIN_EMBED_BASE_URL=http://localhost:11434/v1
export MEGABRAIN_EMBED_MODEL=unclemusclez/jina-embeddings-v2-base-code

megabrain index ~/your/repo                           # once — incremental after
megabrain ask   ~/your/repo "how does auth work end to end"

ask uses your logged-in claude CLI (free on your plan); embeddings never leave your machine. No OpenRouter, no Anthropic key.

Which model? On Claude Code, ask narrates with Haiku by default (fast + cheap on your plan). Bump it with a Claude alias — export MEGABRAIN_ASK_MODEL=sonnet (or opus). ⚠️ On the claude provider this must be a Claude model (haiku/sonnet/ opus/a claude-* id), not an OpenRouter slug like google/….

Inside your AI coding assistant

megabrain speaks MCP, and MCP is portable — the same stdio server runs in every assistant. One command wires up whichever ones you have installed:

megabrain install            # detects + registers; --list to preview, --remove to undo
Registered megabrain in 3 platform(s):
  ✓ Claude Code  registered   ~/.claude.json
  ✓ Codex        registered   ~/.codex/config.toml
  ✓ Antigravity  registered   ~/.gemini/antigravity/mcp_config.json
  · Cursor       skipped (not installed)

Supported: Claude Code · Codex · Antigravity · Cursor · Windsurf · Gemini CLI (--platform <name> for just one). It only ever writes the megabrain key — your other MCP servers are left alone — and it pins the entry to the interpreter megabrain is installed in, so re-running it repairs a config that drifted to an old checkout. Prefer to do it by hand? claude mcp add megabrain -- python3 -m megabrain.mcp_server, or copy the equivalent entry into your assistant's MCP config.

Then use megabrain_ask / megabrain_query instead of grep + Read chains — one call replaces minutes of file-crawling. Five tools, deliberately lean — megabrain exposes only what it alone can do (your agent already has Read/Grep for single files): megabrain_ask (narrated walkthrough, real code spliced), megabrain_query (no LLM, ~200 ms — a flat, relevance-ranked list of exactly the chunks worth reading, with the code, noise dropped), megabrain_index, plus megabrain_forge (teach it a new file type) and megabrain_flows (the opt-in workflow cache).

Commands

megabrain install                                # register the MCP server with your assistants
megabrain index  ~/repo                          # build / update the index
megabrain scan   ~/repo                          # census: what WOULD index + what's skipped & why
megabrain ask    ~/repo "how does X work"        # narrated walkthrough + real code
megabrain query  ~/repo "retry logic"            # raw code map, no LLM (~200 ms)
megabrain query  ~/repo "retry logic" --prune    # flat signal-only chunks, no LLM (drops the noise)
megabrain get    ~/repo src/x.py --symbol Foo    # one file or symbol
megabrain forge  ~/repo                          # teach it your repo's file types (below)
megabrain serve  ~/repo                          # studio web UI at / + the JSON API
megabrain serve-api ~/repo                       # the JSON API only, no UI

Scope to a sub-folder (~/repo/src/auth), search several repos at once (~/a,~/b), and the index auto-refreshes when files change on disk.

megabrain serve ~/repo serves megabrain studio (the web UI, above) at /, and megabrain serve-api ~/repo exposes the same JSON API with no UI mounted. And megabrain scan is the studio's add-repo census on the CLI — what would index and everything skipped with a reason (.gitignore · vendored · generated · too-big): --write applies the proposed .megabrainignore, and megabrain index --scan indexes with those smart filters on (a plain index stays byte-identical).

Rather use the cloud?

No Claude Code or Ollama? One key runs everything through OpenRouter — embeddings and narration — with sensible defaults:

export OPENROUTER_API_KEY=...
megabrain ask ~/repo "how does X work"

megabrain auto-picks the narrator: Claude when its SDK is installed, otherwise OpenRouter. Embeddings always go through OpenRouter or a local endpoint (Anthropic has no embeddings API).

Pin the provider and models with env vars (any OpenRouter slug):

export MEGABRAIN_CHAT_PROVIDER=openrouter                          # pin openrouter (skip claude auto-pick)
export MEGABRAIN_ASK_MODEL=google/gemini-3.1-flash-lite-preview    # the `ask` narration model
export MEGABRAIN_EMBED_MODEL=perplexity/pplx-embed-v1-0.6b         # the embedding model

The full provider matrix — native APIs, hybrid, fully-local GPU, per-provider defaults — is in docs/ARCHITECTURE.md.

100% open-source stack (measured, no closed-weight anything)

Every default above uses a proprietary model somewhere (pplx embeddings, Gemini/Claude narration). If you want zero closed weights — private code, an air-gapped box, or just principle — this combo is measured, not a guess, and holds up:

# 1. embeddings — Apache 2.0, code-tuned, runs on your machine, $0
ollama serve
ollama pull unclemusclez/jina-embeddings-v2-base-code    # 322 MB, one time
export MEGABRAIN_EMBED_BASE_URL=http://localhost:11434/v1
export MEGABRAIN_EMBED_MODEL=unclemusclez/jina-embeddings-v2-base-code

# 2. narration — Apache 2.0 (Qwen), via OpenRouter (or self-host on the same Ollama)
export MEGABRAIN_CHAT_PROVIDER=openrouter
export MEGABRAIN_ASK_MODEL=qwen/qwen3-coder

megabrain index ~/your/repo --force
megabrain ask   ~/your/repo "how does X work"

Retrieval recall (R@1 on a 22-question golden set, sdk-server — does the right file land #1):

stack R@1 weights cost
pplx + closed narrator (the cloud default above) 0.591 closed ~$0.01/ask
jina-code (local) + qwen3-coder (this section) 0.455 all open $0 embed + ~$0.01/ask on OpenRouter, or $0 fully self-hosted

Does ask actually still work? Ran the same two real questions against sdk-server with this exact stack:

  • "where is barge-in handled when the user interrupts mid-speech" → correctly narrated from turn_controller.py, citing 4 files total (event_bus.py, bot_handler.py, webhooks.py too) — broader than the closed-default run.
  • "how does an inbound websocket client get authenticated" → correctly narrated from transports/client/handler.py, the same file the closed stack found.

Both answers were grounded (every code block spliced verbatim, nothing invented) and landed on the right file — the open stack is a real, usable alternative, not a token gesture. The one real cost: qwen/qwen3-coder narrates in ~20-25 s per ask vs ~6 s for Gemini Flash — output-bound, not retrieval-bound, so it's the same trade-off as the cloud cheap-vs-fast pick. qwen3-coder also runs on the same local Ollama for a fully air-gapped setup (no OpenRouter call at all) — just slower without a GPU. Full comparison + a weaker general-purpose local embedder (e5-large, 0.364 R@1) in docs/GUIDE.md §2b.

Compared to claude-context (measured, not vibes)

claude-context (Zilliz) is the closest open-source peer: an MCP server that also does AST-chunked, no-LLM semantic retrieval over a repo. We actually ran it — same repo, same questions, both at their best.

Setup: pinecall/sdk-server (173 source files) · 22 natural-language questions with hand-labelled ground-truth files (barge-in, VAD, turn control, billing…) · R@1 = the right file ranked #1, R@5 = a right file in the top 5 unique files.

megabrain claude-context
R@1 0.864 0.818
R@5 1.000 0.909
Search latency ~22 ms warm · ~370 ms cold ~1400 ms
Vector store SQLite file (zero infra) Milvus + etcd + MinIO (3 containers)
Chunks for the repo 575 1400
LLM in the retrieval path no no
Narrated answer (ask) yes — real code spliced in no (returns chunks; your agent synthesizes)

Both were given their own default embedder (megabrain: pplx-embed-v1-0.6b; claude-context: text-embedding-3-small). To check the gap wasn't just the embedder, we re-ran claude-context on megabrain's exact embedder — it scored 0.727 R@1, i.e. lower. So the difference comes from the retrieval design (tiered CORE/RELATED, import-graph expansion, 4000-char AST merge), not from which embedding model was picked. Fine-grained chunking (2.4× more chunks) also means its top-1 is a fragment, where megabrain's is a whole file with its symbol index.

Caveats, honestly: one repo, 22 questions — this is an indicative result, not a benchmark suite. More importantly, the golden set is ours, on a corpus megabrain has been tuned against, so treat the absolute numbers as home-field. The reproducible parts are the qualitative ones: claude-context needs a Milvus stack, returns chunks rather than a grounded walkthrough, mixes README.md/PROTOCOL.md into code answers (it doesn't separate docs from code), and its get_indexing_status reported ✅ fully indexed while the index was still growing in the background (200 → 1400 chunks), so an agent that trusts it will silently search a partial index. Run it yourself before believing either of us.

How it works

stage what happens
index code is split over its syntax tree (whole functions / classes, never arbitrary line windows), embedded once, stored in SQLite. Incremental by hash.
query no LLM — your question is embedded and matched by vector similarity. Returns every related file in ~200 ms; nothing is dropped.
ask one LLM call narrates the answer and cites code as [[k]]; the engine replaces each citation with the verbatim block from disk. The model can only point at code, never rewrite it — so nothing is hallucinated. Broad questions fan out into parallel sub-agents, then a parent synthesizes.
forge for a file type the engine doesn't index yet (.toml, .astro, a private DSL), an LLM writes a chunking strategy — accepted only after it partitions every matching file exactly. One-time, at your command, off the query path.
flows (opt-in) turn it on and every ask caches its cross-file walkthrough; the next related question retrieves the whole workflow at once. Off by default — plain query/ask are unchanged.

Languages: Python · JS/TS · Markdown built in; Ruby · Go · Rust · PHP with pip install 'megabrain[languages]'; anything else via megabrain forge (below).

forge — megabrain writes its own chunkers

Repos carry more than code: .toml, .yaml, .astro, .proto, private DSLs… Anything outside the registry is invisible to retrieval. megabrain forge fixes that per repo:

megabrain forge ~/repo --list        # census: which text file types aren't indexed (free)
megabrain forge ~/repo               # LLM-write a chunking strategy per type, validate, install
megabrain forge ~/repo --dry-run     # show the generated code without installing

For each uncovered extension, an LLM (same provider stack as ask) writes a ChunkStrategy from the contract source + real sample files, and it is only accepted after chunking every matching file in the repo with a clean exact-line partition (validate_partition — failures feed a repair loop, and nothing unvetted ever installs). The vetted module lands in .megabrain/strategies/<ext>.py, sha-recorded in a user-level trust store (~/.megabrain/trust.json), and from then on every index — including the 60 s auto-refresh — loads it automatically. Hand-written strategies work the same way: drop the file in .megabrain/strategies/ and approve it with megabrain trust ~/repo.

Real run on pallets/click: forge detected .toml (11 files) and .yaml (8 workflows), generated both strategies on the first attempt (~28 s total), and "which workflow runs the test suite?" went from missing entirely to ranking .github/workflows/tests.yaml #1.

--specialize — measure a hand-written chunker (no LLM)

For a file type the engine ALREADY reads but chunks poorly (a giant lookup table blobs; a class of many tiny methods merges), you can hand-write a better strategy and have the engine measure it before it installs:

megabrain forge ~/repo --specialize          # census: covered files the built-in chunks poorly
# write a ChunkStrategy into .megabrain/strategies/<ext>.py, then gate it:
python -c "from megabrain.forge.specialize import gate_strategy; \
           print(gate_strategy('~/repo', open('strat.py').read(), '.py'))"

gate_strategy indexes the built-in vs your candidate for real, scores span-IoU

  • hit@1 on neutral probes over every file the candidate changes, and installs (trust-gated) only if it beats a literature-tuned baseline — never on a whisper of improvement.

We tried letting an LLM write these and removed it. Across four repos the generated chunkers lost to a five-line deterministic recipe. And the deeper, measured finding: on a real query set (the sdk-server golden) tighter chunks LOWER retrieval ranking — the 4000-char merge concentrates a file's evidence and that is what wins R@1 (4000 → 0.86, 2000 → 0.82, blob-split → 0.77). Tighter chunks help navigation (fewer lines to read) but not retrieval. The built-in default is a genuine optimum; leave it alone unless you measure a win. Specialization is for the rare pathological file, gated hard.

flows — self-caching workflow retrieval (opt-in, off by default)

Every ask synthesizes a cross-file workflow ("VAD detects speech → TurnController.on_vad_start → cancel TTS") that the engine used to discard. Turn the flow cache on and it keeps them: the next related question — even worded completely differently — retrieves the whole workflow at once.

megabrain ask ~/repo "how does X work"       # unchanged: flows are OFF by default
megabrain flows ~/repo --enable              # opt in for this repo; asks now cache their flows
megabrain index ~/repo --warm-flows 12       # or pre-fill: discover the repo's 12 top workflows now
megabrain flows ~/repo                        # list what's cached · --clear to reset
  • Off by default — plain query/ask behave byte-for-byte as before, at zero cost. It's a mode a team turns on so its megabrain accumulates the repo's workflows from use (great for onboarding).
  • Rules intact: the LLM + the one embed happen at ask time (write path); the read path is pure cosine. Flows only add their source files to the bundle when missing (never displace real files → completeness only rises), and the narrator gets the cached flow as non-citable context. Any flow whose cited files change sha is pruned on the next index — a stale walkthrough can't outlive its code, and ask splices real code regardless.

Validated on sdk-server: --warm-flows 5 discovered and cached the system's main workflows; a paraphrase ("how does the bot stop talking when the user cuts in") retrieved the barge-in flow cached from a differently-worded question.

See it live

bernardocastro.dev/megabrain — search 7 popular open-source repos and watch the engine rank the files and pick the exact code chunks, live. Or run it locally: python examples/webui/server.py.

Learn more

  • docs/GUIDE.md — step-by-step: providers, indexing, the 2000-vs-4000 budget choice, custom chunkers, and the flow cache
  • docs/ARCHITECTURE.md — the full design, the locked rules, and the measurements behind them
  • examples/ — programmatic API · a custom .sql chunker · the web demo
  • CONTRIBUTING.md — the best first PR is a new language

MIT · github.com/bernatch22/megabrain

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

megabrain-0.9.1.tar.gz (215.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

megabrain-0.9.1-py3-none-any.whl (185.0 kB view details)

Uploaded Python 3

File details

Details for the file megabrain-0.9.1.tar.gz.

File metadata

  • Download URL: megabrain-0.9.1.tar.gz
  • Upload date:
  • Size: 215.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.0

File hashes

Hashes for megabrain-0.9.1.tar.gz
Algorithm Hash digest
SHA256 8b89ac1b3dfa8e33ff1ad6221f53dc5f7c2ee7df18cc0e4c827346eea64b831c
MD5 bc499c23197d33306c7e295b14396bd7
BLAKE2b-256 51f9007634621fba94c9f5e62ba5e62535afbb0951c5338f51e91e252b055215

See more details on using hashes here.

File details

Details for the file megabrain-0.9.1-py3-none-any.whl.

File metadata

  • Download URL: megabrain-0.9.1-py3-none-any.whl
  • Upload date:
  • Size: 185.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.0

File hashes

Hashes for megabrain-0.9.1-py3-none-any.whl
Algorithm Hash digest
SHA256 5c3a7c33eeba161c8f3d5e153fd8a7f600837092fa65a0d4c386b5f789a4b2c7
MD5 1f9a3434fae783070f675c388a715d5e
BLAKE2b-256 dc455342babc424ec256c239b59405c18398f5dc7c23bcb3f4f9a7f00d9bd969

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page