Skip to main content

A local context layer for AI tools: mirror your repositories, index them into a knowledge graph, and serve it over MCP so agents answer from real source instead of guessing.

Project description

contextlake, all your real context in one local lake. Pebble the otter surfacing from a misty lake cradling a glowing pebble of context.

contextlake

All your real context, in one local lake.

A local context layer for your AI tools: mirror your repositories, index them
into a knowledge graph, and serve it over MCP, so agents answer from real source instead of guessing.

CI PyPI Python 3.10+ for the knowledge layer, 3.9+ for the mirror core Offline-first License: MIT


Why contextlake

Your AI assistant is only as good as what it can actually see. Point it at one file and it's sharp; ask it about the system, which service calls this API, who depends on that package, where a symbol is really defined across dozens of repos, and it starts guessing.

contextlake gives your tools the real source to read. It mirrors your repositories to your machine, indexes them into a queryable knowledge graph, and serves that graph to your editor over MCP. Everything runs locally and offline, no code leaves your machine, and it carries no credentials of its own.

How it works

contextlake is three layers you adopt one at a time. The mirror is useful on its own, and each layer above it is optional.

contextlake architecture. On the left, your repos: a GitLab group, plus optional Figma, Jira, and other MCP connectors. In the centre, contextlake indexes and mirrors them into a graph and embeddings, a wiki, and connectors. On the right, it serves the result over MCP to your AI tools: Claude Code, Windsurf, Kiro, Cursor, and Postman.

  1. Mirror: clone every repo you can reach in a GitLab group, GitHub org, Bitbucket workspace, or Gitea/Codeberg/Forgejo owner into a faithful copy of its namespace tree, each on its most active branch, kept fresh with one command.
  2. Knowledge layer (optional): parse the mirror into a code + dependency graph across 14 languages plus Terraform infrastructure, SQL schema, and package manifests (npm / PyPI / NuGet / Maven), add semantic search, a council-verified wiki (each page reviewed and scored before publishing, low-confidence pages dropped), and connectors to Atlassian / Figma / GitLab / Slack.
  3. Serve: expose it all over MCP and an offline interactive graph visualizer, so agents can answer "where is X defined?" or "who calls Y?" instead of grepping.

Each layer has its own guide: the mirror in Usage & config, the knowledge layer and serving in Knowledge layer, and the whole flow start to finish in QUICKSTART.

Install

pip install "contextlake[kb]"       # the full tool: mirror + graph, search, wiki, MCP server
pip install contextlake             # mirror-only core (no pip dependencies at all)

Everything in the quickstart below needs the [kb] extra (Python 3.10+); the plain install is just the mirroring CLI and runs on Python 3.9+.

Prefer an isolated, zero-setup install? uv fetches the right Python and an isolated environment for you:

uv tool install "contextlake[kb]"            # install the CLI on your PATH
uvx --from "contextlake[kb]" contextlake --help   # …or run it once, without installing
# pipx install "contextlake[kb]"             # pipx works too
Install extras (the mirror needs none, add these for the knowledge layer)
Extra Adds When you need it
[kb] The knowledge layer: parse → graph → wiki → MCP server Anything beyond mirroring
[kb-full] [kb] + the built-in CPU embedder + sqlite-vec ANN One-step local semantic search, no Ollama or API key
[kb-vec] The sqlite-vec ANN backend Faster vector search than the pure-Python fallback
[kb-local] The built-in CPU embedder (model2vec, ~30 MB) Semantic search with no Ollama or API key
[kb-fastembed] A higher-quality ONNX embedder (~90 MB) Better semantic ranking
[llm-local] A built-in CPU model for the wiki (llama-cpp) wiki --llm builtin with no Ollama or API key

[llm-local] is the one extra a plain pip install cannot finish on its own: llama-cpp-python publishes no wheels to PyPI (llama.cpp is built per hardware backend, so upstream ships one index per accelerator), so pip compiles C++ unless you point it at one. Let contextlake do it:

contextlake doctor --fix llm-local     # add --dry-run to see the exact command first

This applies to pip installs only: the standalone binary has the index preconfigured and installs the runtime on its first run, and the full Docker image ships it baked in.

Docker (turnkey / air-gapped: models baked in)

The published image bundles the knowledge layer plus the built-in CPU models (embedder + a small wiki LLM), so it runs with no Ollama, no API key, and no model download at runtime. The PyPI wheel stays the primary install; reach for the image on locked-down or offline machines. Runs as a non-root user.

docker run -v "$PWD:/work" ghcr.io/sayak-sarkar/contextlake doctor
docker run -v "$PWD:/work" ghcr.io/sayak-sarkar/contextlake kb index

The -v mount is what makes the run worth doing: everything contextlake persists, the knowledge store included, is written under it as .contextlake/, so it is still there on the host after the container exits. Drop the -v and the run is ephemeral.

The container runs as uid 1000, and a bind mount keeps the host's ownership, so if your host account is not uid 1000 the write fails with a permission error. Pass your own ids to fix it:

docker run -u "$(id -u):$(id -g)" -v "$PWD:/work" ghcr.io/sayak-sarkar/contextlake kb index

It fails rather than falling back on purpose. Before 5.1.0 the store was written inside the container instead, so the run appeared to succeed and the index was gone the moment the container exited.

A :slim tag is also published, no llama-cpp-python, no baked wiki-LLM GGUF, much smaller pull. Semantic search still works (the embedder is pure Python); point the wiki tier at Ollama/OpenAI/Anthropic/cli instead of the built-in LLM.

docker run -v "$PWD:/work" ghcr.io/sayak-sarkar/contextlake:slim doctor
From source (for contributors)
git clone https://github.com/sayak-sarkar/contextlake && cd contextlake
pip install -e ".[kb]"
Update & uninstall

Upgrade in place (whichever installer you used):

pipx upgrade contextlake                       # pipx
pip install --upgrade "contextlake[kb-full]"   # pip
uv tool upgrade contextlake                     # uv
docker pull ghcr.io/sayak-sarkar/contextlake   # image

Your store and config carry forward. Confirm with contextlake --version, then run contextlake doctor.

doctor is load-bearing here, not a formality. An upgrade that changes how code is parsed leaves every existing shard describing the old parse, and a plain kb index will not notice: it skips repos whose HEAD commit has not moved, and upgrading contextlake does not move anyone's HEAD. doctor compares the parser version recorded in each shard against the running one and names the repos that need rebuilding. When it asks for one, force it:

contextlake kb index --force

Upgrading to 5.0.0 specifically requires this, for every indexed repo: that release changed the parser and every shard written before it is stale.

Uninstall the tool, then optionally remove what it created (it never writes inside your repos, so your source is never touched):

pipx uninstall contextlake        # or: pip uninstall contextlake
rm -rf ~/.contextlake             # store + kb.toml + graph/dashboard exports (optional)
rm -f  ~/.contextlake.ini         # mirror config (optional)
# mirrored repos live in your work_dir (default ~/work), delete only if unwanted

Prerequisites: git, and, only for fleet mirroring, the platform's token env var (GITLAB_TOKEN with read_api + read_repository, or GITHUB_TOKEN / BITBUCKET_TOKEN / GITEA_TOKEN); on GitLab an authenticated glab works instead. The knowledge layer needs neither. Once installed, contextlake, python -m contextlake, and python3 run-contextlake.py are equivalent.

Quickstart: one repo, no setup

You don't need GitLab or any config to try contextlake on a repo you already have. No install? Run it once with uvx: prefix any command below with uvx --from "contextlake[kb]" (e.g. uvx --from "contextlake[kb]" contextlake kb index --source .).

contextlake kb index                     # parse the current repo into a local knowledge graph
contextlake kb graph --overview --open   # open the interactive graph in your browser
contextlake kb serve                     # …or serve it to your AI IDE over MCP

Wire it into your editor in one line, no config file needed (it uses the local ~/.contextlake/kb store you just built):

claude mcp add contextlake-kb -- contextlake kb serve      # Claude Code
# zero-install variant: claude mcp add contextlake-kb -- uvx --from "contextlake[kb]" contextlake kb serve

The contextlake graph visualizer showing a repository's symbols as a navigable node graph, with a type-glyph legend, search, and a corner minimap

contextlake kb graph, a whole codebase as one offline, navigable graph.

Everything lands in a local store (~/.contextlake/kb), nothing leaves your machine. Index any path with --source PATH, or every git repo under a directory with --workspace DIR.

Want the full path, mirror a GitLab fleet → graph → wired editor in a few minutes? QUICKSTART.md walks the whole flow.

Fleet mode: mirror a whole org

Where contextlake goes beyond single-repo tools is mirroring and cross-referencing a whole fleet: a GitLab group, a GitHub org, a Bitbucket workspace, or a Gitea/Codeberg/Forgejo owner. Copy the example config and set your platform, group and workspace:

cp .contextlake.ini.example ~/.contextlake.ini
[contextlake]
work_dir = ~/work
gitlab_group = your-gitlab-group
# or any other platform:
# platform = github
# group = your-org
contextlake mirror status      # see where you stand (read-only)
contextlake mirror sync        # fetch → clone → update → branches → verify → audit

Auth is one env var: the platform's token (GITLAB_TOKEN / GITHUB_TOKEN / BITBUCKET_TOKEN / GITEA_TOKEN), carried in headers and the child environment, never in URLs or argv, so .contextlake.ini holds only non-secret settings and is gitignored by default. (On GitLab, an authenticated glab works too; public orgs on other platforms need no token at all.) It runs across hundreds of repos concurrently, with an adaptive worker pool, retries with backoff, and never stomps on the feature branch you're in the middle of.

Behind a slow / TLS-inspecting corporate proxy (e.g. Zscaler) where glab's API calls time out? Set GITLAB_TOKEN (a read_api token) and contextlake enumerates projects via its own HTTP client, which tolerates the slow DNS where glab's short dial timeout fails.

Commands at a glance

Run any command as contextlake <command>; each has scoped help via contextlake <command> --help. Each verb lives under the noun it belongs to, mirror for mirroring git repositories, kb for the knowledge layer, except init, bootstrap, version, completion, and doctor, which span both tiers or neither. Per-command docs live with their layer: the mirror commands in usage.md; the knowledge-layer commands (kb index, kb embed, kb connect, kb wiki, kb query, kb owners, kb impact, kb graph, …) in knowledge-layer.md, and kb serve/kb steer in serve.md.

Command What it does
init Guided setup: write your mirror + knowledge-layer config (--skip-interactive for non-interactive)
mirror status Show the workspace sync state vs GitLab (read-only)
mirror sync The full pipeline: fetch → clone → update → branches → verify → audit
mirror fetch · mirror clone · mirror update The sync steps, individually
mirror branches Switch each repo to its most active branch
mirror verify · mirror audit Check the mirror vs GitLab; report repo health, age & drift (JSON + CSV)
bootstrap Turnkey: sync + index + connect + embed + enrich + wiki + steer (--no-enrich to skip)
kb index Build the code/dependency graph (--workspace, incremental, --watch)
kb source Manage connectors: add/list/remove/test/enable/disable knowledge sources; edits kb.toml for you, comments preserved
kb connect Link repos to Atlassian / Figma / GitLab items (--watch to keep refreshing)
kb embed Build semantic-search vectors (zero-config built-in CPU model, Ollama, or an API; incremental, --watch)
kb enrich Query connected sources with codebase-derived terms and store the results in a searchable @enrich partition that feeds the wiki
kb ingest Aggregate external docs into the graph + semantic store (built-in files/web/api/graphql/mcp sources, or plugins)
kb wiki [<repo>…] LLM-synthesized, council-verified wiki pages (all repos, or just the named ones); --llm builtin|ollama|openai|anthropic|cli enables the LLM tier inline
kb query Search the index (--kind, --repo, --as-of <commit>)
kb owners (alias kb who-knows) Likely owners / SMEs for a repo (or --path), ranked from git history
kb impact (alias kb blast-radius) Change-impact / blast radius: what depends on a symbol (--hops, --repo to disambiguate)
kb graph Visualize the graph, offline interactive HTML / DOT / Mermaid / JSON
kb dashboard Local knowledge-system dashboard UI (--serve; --sample for the bundled demo fleet; --site DIR for a static offline export)
kb serve Expose the graph over MCP (--transport stdio/http/sse)
kb steer Write editor steering, AGENTS.md, .mcp.json, .vscode/mcp.json, .windsurfrules, skills
kb lint · doctor · kb eval Graph health · environment check · retrieval-quality scoring

Global options apply to any command: --dry-run (preview without changing anything), -v/-q (verbosity), --log-file PATH, --config PATH, --version. Output is colorized on a TTY and plain when piped; set NO_COLOR to force-disable.

For runs nobody watches, the systemd timer in examples/, cron, CI, there is a second set: --log-format json (one JSON object per line, every line stamped with a run id), --metrics-file PATH (Prometheus textfile-collector output), --redact (the --log-file copy is already scrubbed of workspace paths, group and repo names), and --access-log. See Reading the console output.

Knowledge layer

Beyond mirroring, the optional contextlake.kb layer turns your repos into a knowledge graph and serves it to AI tools over MCP. It can link repos directly to the Atlassian / Figma / GitLab / Slack items and code symbols that reference them, add semantic search, write a curated wiki, visualize the graph (offline interactive HTML, fleet overview, a symbol's neighbourhood, or a single repo), and generate per-tool steering files + a skills library. Most of it needs no model; the rest works with a local Ollama or any OpenAI-compatible endpoint.

One command sets it all up (configs are read from their default locations):

contextlake bootstrap

Full guide: docs/knowledge-layer.md.

The dashboard

contextlake kb dashboard --serve opens a local, offline-first window into everything the knowledge layer builds: a fleet overview, per-repo anatomy, the cross-repo architecture graph, change-impact (blast radius), health, search, and a Chat tab to ask questions about the fleet in plain language (free graph router always on, LLM-synthesized prose opt-in via --llm-chat). Try it with zero setup via contextlake kb dashboard --serve --sample.

The contextlake dashboard fleet overview: stat cards, a knowledge-confidence bar, and repos grouped by namespace, with a Cards/List/Table layout switcher.

The dashboard: a guided tour, step by step, with screenshots.

Documentation

License

MIT, see LICENSE. Pebble the otter is the project mascot; deep context, clear answers.

Project details


Release history Release notifications | RSS feed

This version

5.1.1

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

contextlake-5.1.1.tar.gz (1.8 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

contextlake-5.1.1-py3-none-any.whl (1.8 MB view details)

Uploaded Python 3

File details

Details for the file contextlake-5.1.1.tar.gz.

File metadata

  • Download URL: contextlake-5.1.1.tar.gz
  • Upload date:
  • Size: 1.8 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for contextlake-5.1.1.tar.gz
Algorithm Hash digest
SHA256 12330f8c0f7b7f41a7908fab05d310ad06c94452da6c5fd6699dbf9796f147aa
MD5 ee5a81d3232ead993b24699d93b081ed
BLAKE2b-256 99dca064b76a17ed1a6d30bf6dbcbc9dbb882b59ad22462c621ff02860c15611

See more details on using hashes here.

Provenance

The following attestation bundles were made for contextlake-5.1.1.tar.gz:

Publisher: release.yml on sayak-sarkar/contextlake

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file contextlake-5.1.1-py3-none-any.whl.

File metadata

  • Download URL: contextlake-5.1.1-py3-none-any.whl
  • Upload date:
  • Size: 1.8 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for contextlake-5.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 df1a90a1beb786e20f47f49567f3c2bdda0cc647c39ed934ae9ac4b36f3f88e0
MD5 767f952b62865d919cdce61d2b2a3886
BLAKE2b-256 268bd1170688c9671da7c8cabef98021ee1f0298490f0cffc40637631c64ff47

See more details on using hashes here.

Provenance

The following attestation bundles were made for contextlake-5.1.1-py3-none-any.whl:

Publisher: release.yml on sayak-sarkar/contextlake

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page