Skip to main content

megabrain

megabrain

Ask a codebase a question. Get the exact code back.

PyPI MIT No LLM in the retrieval path MCP ready


Point megabrain at a repo and ask "how does auth work" in plain English. It finds all the related code — in ~200 ms, using no LLM, just math on embeddings — and an LLM narrates a walkthrough with the real code spliced in from disk. Nothing is invented: every line shown is copied verbatim.

Use it from the terminal, as an MCP server inside Claude Code, or as a Python library.

Quickstart — the easy path, no API keys

Everything runs on your machine: ask narrates on your Claude Code subscription, embeddings run locally on Ollama. No cloud keys.

pip install 'megabrain[claude]'                      # engine + Claude Code narration

ollama pull nomic-embed-text                          # local embeddings, one time
export MEGABRAIN_EMBED_BASE_URL=http://localhost:11434/v1
export MEGABRAIN_EMBED_MODEL=nomic-embed-text

megabrain index ~/your/repo                           # once — incremental after
megabrain ask   ~/your/repo "how does auth work end to end"

ask uses your logged-in claude CLI (free on your plan); embeddings never leave your machine. No OpenRouter, no Anthropic key.

Inside Claude Code

Register it as an MCP server and research any indexed repo without leaving Claude Code:

claude mcp add megabrain -- python3 -m megabrain.mcp_server

Then use megabrain_ask / megabrain_query instead of grep + Read chains — one call replaces minutes of file-crawling. Tools: megabrain_ask (narrated walkthrough), megabrain_query (raw code map, no LLM), megabrain_get, megabrain_chunks, megabrain_index.

Commands

megabrain index  ~/repo                          # build / update the index
megabrain ask    ~/repo "how does X work"        # narrated walkthrough + real code
megabrain query  ~/repo "retry logic"            # raw code map, no LLM (~200 ms)
megabrain get    ~/repo src/x.py --symbol Foo    # one file or symbol
megabrain forge  ~/repo                          # teach it your repo's file types (below)
megabrain serve-api ~/repo                       # long-running HTTP API (warm state)

Scope to a sub-folder (~/repo/src/auth), search several repos at once (~/a,~/b), and the index auto-refreshes when files change on disk.

Rather use the cloud?

No Claude Code or Ollama? One key runs everything through OpenRouter — embeddings and narration — with sensible defaults:

export OPENROUTER_API_KEY=...
megabrain ask ~/repo "how does X work"

megabrain auto-picks the narrator: Claude when its SDK is installed, otherwise OpenRouter. Embeddings always go through OpenRouter or a local endpoint (Anthropic has no embeddings API). The full provider matrix — native APIs, hybrid, fully-local GPU — is in docs/ARCHITECTURE.md.

How it works

stage what happens
index code is split over its syntax tree (whole functions / classes, never arbitrary line windows), embedded once, stored in SQLite. Incremental by hash.
query no LLM — your question is embedded and matched by vector similarity. Returns every related file in ~200 ms; nothing is dropped.
ask one LLM call narrates the answer and cites code as [[k]]; the engine replaces each citation with the verbatim block from disk. The model can only point at code, never rewrite it — so nothing is hallucinated. Broad questions fan out into parallel sub-agents, then a parent synthesizes.

Languages: Python · JS/TS · Markdown built in; Ruby · Go · Rust · PHP with pip install 'megabrain[languages]'.

forge — megabrain writes its own chunkers

Repos carry more than code: .toml, .yaml, .astro, .proto, private DSLs… Anything outside the registry is invisible to retrieval. megabrain forge fixes that per repo:

megabrain forge ~/repo --list        # census: which text file types aren't indexed (free)
megabrain forge ~/repo               # LLM-write a chunking strategy per type, validate, install
megabrain forge ~/repo --dry-run     # show the generated code without installing

For each uncovered extension, an LLM (same provider stack as ask) writes a ChunkStrategy from the contract source + real sample files, and it is only accepted after chunking every matching file in the repo with a clean exact-line partition (validate_partition — failures feed a repair loop, and nothing unvetted ever installs). The vetted module lands in .megabrain/strategies/<ext>.py, sha-recorded in a user-level trust store (~/.megabrain/trust.json), and from then on every index — including the 60 s auto-refresh — loads it automatically. Hand-written strategies work the same way: drop the file in .megabrain/strategies/ and approve it with megabrain trust ~/repo.

Real run on pallets/click: forge detected .toml (11 files) and .yaml (8 workflows), generated both strategies on the first attempt (~28 s total), and "which workflow runs the test suite?" went from missing entirely to ranking .github/workflows/tests.yaml #1.

See it live

bernardocastro.dev/megabrain — search 7 popular open-source repos and watch the engine rank the files and pick the exact code chunks, live. Or run it locally: python examples/webui/server.py.

Learn more

  • docs/ARCHITECTURE.md — the full design, the locked rules, and the measurements behind them
  • examples/ — programmatic API · a custom .sql chunker · the web demo
  • CONTRIBUTING.md — the best first PR is a new language

MIT · github.com/bernatch22/megabrain

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

megabrain-0.6.0.tar.gz (114.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

megabrain-0.6.0-py3-none-any.whl (96.6 kB view details)

Uploaded Python 3

File details

Details for the file megabrain-0.6.0.tar.gz.

File metadata

  • Download URL: megabrain-0.6.0.tar.gz
  • Upload date:
  • Size: 114.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.0

File hashes

Hashes for megabrain-0.6.0.tar.gz
Algorithm Hash digest
SHA256 7a1fee868a61c63fb35eb974897d8657b1e901c8ec3d91ae1e4d0492a21715ed
MD5 d2d5b20040d6a774f8fb2553d336410e
BLAKE2b-256 bb550165007d0b2ceee536bb38804d320784e199f996bbbfa6a87222d655a855

See more details on using hashes here.

File details

Details for the file megabrain-0.6.0-py3-none-any.whl.

File metadata

  • Download URL: megabrain-0.6.0-py3-none-any.whl
  • Upload date:
  • Size: 96.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.0

File hashes

Hashes for megabrain-0.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e27e432a80c842cd7d7e391b1f8c6a1ab8b52582f1cc34c9d37372d11955553f
MD5 cf1613ef6d989a1f35cc25ece1331b72
BLAKE2b-256 5d7a76fb7117e21040d81a4db162b46633565487605092b5b21297b676212256

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page