contextlake
All your real context, in one local lake.
A local context layer for your AI tools: mirror your repositories, index them
into a knowledge graph, and serve it over MCP, so agents answer from real source instead of guessing.
Why contextlake
Your AI assistant is only as good as what it can actually see. Point it at one file and it's sharp; ask it about the system, which service calls this API, who depends on that package, where a symbol is really defined across dozens of repos, and it starts guessing.
contextlake gives your tools the real source to read. It mirrors your repositories to your machine, indexes them into a queryable knowledge graph, and serves that graph to your editor over MCP. Everything runs locally and offline, no code leaves your machine, and it carries no credentials of its own.
How it works
contextlake is three layers you adopt one at a time. The mirror is useful on its own, and each layer above it is optional.
- Mirror: clone every repo you can reach in a GitLab group, GitHub org, Bitbucket workspace, or Gitea/Codeberg/Forgejo owner into a faithful copy of its namespace tree, each on its most active branch, kept fresh with one command.
- Knowledge layer (optional): parse the mirror into a code + dependency graph across 14 languages plus Terraform infrastructure, SQL schema, and package manifests (npm / PyPI / NuGet / Maven), add semantic search, a council-verified wiki (each page reviewed and scored before publishing, low-confidence pages dropped), and connectors to Atlassian / Figma / GitLab / Slack.
- Serve: expose it all over MCP and an offline interactive graph visualizer, so
agents can answer "where is
Xdefined?" or "who callsY?" instead of grepping.
Each layer has its own guide: the mirror in Usage & config, the knowledge layer and serving in Knowledge layer, and the whole flow start to finish in QUICKSTART.
Install
pip install "contextlake[kb]" # the full tool: mirror + graph, search, wiki, MCP server
pip install contextlake # mirror-only core (one dependency: argcomplete)
Everything in the quickstart below needs the [kb] extra (Python 3.10+); the plain
install is just the mirroring CLI. Both need Python 3.10 or newer: one floor for the whole
tool, since the split floor the mirror core used to allow only ever surprised people.
Prefer an isolated, zero-setup install? uv fetches the right
Python and an isolated environment for you:
uv tool install "contextlake[kb]" # install the CLI on your PATH
uvx --from "contextlake[kb]" contextlake --help # …or run it once, without installing
# pipx install "contextlake[kb]" # pipx works too
Docker, the standalone binaries, the full extras table, upgrading, and uninstalling all live on one page: Install and upgrade. If an install misbehaves, see Troubleshooting.
Prerequisites: git, and, only for fleet mirroring, the platform's token env var
(GITLAB_TOKEN with read_api + read_repository, or GITHUB_TOKEN /
BITBUCKET_TOKEN / GITEA_TOKEN); on GitLab an authenticated
glab works instead. The knowledge layer needs
neither. Once installed, contextlake, python -m contextlake, and
python3 run-contextlake.py are equivalent.
Quickstart: one repo, no setup
You don't need GitLab or any config to try contextlake on a repo you already have.
No install? Run it once with uvx: prefix any command
below with uvx --from "contextlake[kb]" (e.g. uvx --from "contextlake[kb]" contextlake kb index --source .).
contextlake kb index # parse the current repo into a local knowledge graph
contextlake kb graph --overview --open # open the interactive graph in your browser
contextlake kb serve # …or serve it to your AI IDE over MCP
Wire it into your editor in one line, no config file needed (it uses the local
~/.contextlake/kb store you just built):
claude mcp add contextlake-kb -- contextlake kb serve # Claude Code
# zero-install variant: claude mcp add contextlake-kb -- uvx --from "contextlake[kb]" contextlake kb serve
contextlake kb graph, a whole codebase as one offline, navigable graph.
Everything lands in a local store (~/.contextlake/kb), nothing leaves your machine. Index
any path with --source PATH, or every git repo under a directory with --workspace DIR.
Want the full path, mirror a GitLab fleet → graph → wired editor in a few minutes? QUICKSTART.md walks the whole flow.
Fleet mode: mirror a whole org
Where contextlake goes beyond single-repo tools is mirroring and cross-referencing a whole fleet: a GitLab group, a GitHub org, a Bitbucket workspace, or a Gitea/Codeberg/Forgejo owner. Copy the example config and set your platform, group and workspace:
cp .contextlake.ini.example ~/.contextlake.ini
[contextlake]
work_dir = ~/work
gitlab_group = your-gitlab-group
# or any other platform:
# platform = github
# group = your-org
contextlake mirror status # see where you stand (read-only)
contextlake mirror sync # fetch → clone → update → branches → verify → audit
Auth is one env var: the platform's token (GITLAB_TOKEN / GITHUB_TOKEN /
BITBUCKET_TOKEN / GITEA_TOKEN), carried in headers and the child environment, never in
URLs or argv, so .contextlake.ini holds only non-secret settings and is gitignored by
default. (On GitLab, an authenticated glab works too; public orgs on other platforms need
no token at all.) It runs across hundreds of repos concurrently, with an adaptive worker
pool, retries with backoff, and never stomps on the feature branch you're in the middle
of.
Behind a slow / TLS-inspecting corporate proxy (e.g. Zscaler) where
glab's API calls time out? SetGITLAB_TOKEN(aread_apitoken) and contextlake enumerates projects via its own HTTP client, which tolerates the slow DNS whereglab's short dial timeout fails.
Commands at a glance
Run any command as contextlake <command>; each has scoped help via
contextlake <command> --help. Each verb lives under the noun it belongs to, mirror for
mirroring git repositories, kb for the knowledge layer, except init, bootstrap,
version, completion, and doctor, which span both tiers or neither. Per-command docs live
with their layer: the mirror commands in usage.md;
the knowledge-layer build commands (kb index, kb embed, kb connect, kb wiki, …) in
knowledge-layer.md; the query commands
(kb query, kb impact, kb owners) in ask-the-graph.md;
and kb serve/kb steer in serve.md.
| Command | What it does |
|---|---|
init |
Guided setup: write your mirror + knowledge-layer config (--skip-interactive for non-interactive) |
mirror status |
Show the workspace sync state vs GitLab (read-only) |
mirror sync |
The full pipeline: fetch → clone → update → branches → verify → audit |
mirror fetch · mirror clone · mirror update |
The sync steps, individually |
mirror branches |
Switch each repo to its most active branch |
mirror verify · mirror audit |
Check the mirror vs GitLab; report repo health, age & drift (JSON + CSV) |
bootstrap |
Turnkey: sync + index + connect + embed + enrich + wiki + steer (--no-enrich to skip) |
kb index |
Build the code/dependency graph (--workspace, incremental, --watch; a directory holding git repos is refused with the right command, --bundle to index it as one repo anyway) |
kb source |
Manage connectors: add/list/remove/test/enable/disable knowledge sources; edits kb.toml for you, comments preserved |
kb connect |
Link repos to Atlassian / Figma / GitLab items (--watch to keep refreshing) |
kb embed |
Build semantic-search vectors (zero-config built-in CPU model, Ollama, or an API; incremental, --watch) |
kb enrich |
Query connected sources with codebase-derived terms and store the results in a searchable @enrich partition that feeds the wiki |
kb ingest |
Aggregate external docs into the graph + semantic store (built-in files/web/api/graphql/mcp sources, or plugins) |
kb wiki [<repo>…] |
LLM-synthesized, council-verified wiki pages (all repos, or just the named ones); --llm builtin|ollama|openai|anthropic|cli enables the LLM tier inline (builtin needs doctor --fix llm-local first on a pip install; ollama needs no compiler) |
kb query |
Search the index (--kind, --repo, --as-of <commit>) |
kb owners (alias kb who-knows) |
Likely owners / SMEs for a repo (or --path), ranked from git history |
kb impact (alias kb blast-radius) |
Change-impact / blast radius: what depends on a symbol (--hops, --repo to disambiguate) |
kb graph |
Visualize the graph, offline interactive HTML / DOT / Mermaid / JSON |
kb dashboard |
Local knowledge-system dashboard UI (--serve; --sample for the bundled demo fleet; --site DIR for a static offline export) |
kb serve |
Expose the graph over MCP (--transport stdio/http/sse; --tool-concurrency N bounds concurrent tool calls, default 2, raising it makes the server slower) |
kb steer |
Write editor steering, AGENTS.md, .mcp.json, .vscode/mcp.json, .windsurfrules, skills |
kb lint · doctor · kb eval |
Graph health · environment check · retrieval-quality scoring |
Global options apply to any command: --dry-run (preview without changing anything),
-v/-q (verbosity), --log-file PATH, --config PATH, --version. Output is colorized on
a TTY and plain when piped; set NO_COLOR to force-disable.
For runs nobody watches, the systemd timer in examples/, cron, CI, there is a
second set: --log-format json (one JSON object per line, every line stamped with a run id),
--metrics-file PATH (Prometheus textfile-collector output), --redact (the --log-file copy
is already scrubbed of workspace paths, group and repo names), and --access-log. See
Reading the console output.
Knowledge layer
Beyond mirroring, the optional contextlake.kb layer turns your repos into a knowledge
graph and serves it to AI tools over MCP. It can link repos directly to the Atlassian /
Figma / GitLab / Slack items and code symbols that reference them, add semantic search,
write a curated wiki, visualize the graph
(offline interactive HTML, fleet overview, a symbol's neighbourhood, or a single repo), and
generate per-tool steering files + a skills library. Most of it needs no model; the rest
works with a local Ollama or any OpenAI-compatible endpoint.
One command sets it all up (configs are read from their default locations):
contextlake bootstrap
Full guide: docs/knowledge-layer.md.
The dashboard
contextlake kb dashboard --serve opens a local, offline-first window into everything the
knowledge layer builds: a fleet overview, per-repo anatomy, the cross-repo architecture
graph, change-impact (blast radius), health, search, and a Chat tab to ask questions
about the fleet in plain language (free graph router always on, LLM-synthesized prose
opt-in via --llm-chat). Try it with zero setup via contextlake kb dashboard --serve --sample.
The dashboard: a guided tour, step by step, with screenshots.
Documentation
- QUICKSTART.md, install → bootstrap → wire your editor, in minutes
- docs/install.md, every install channel, upgrading, and uninstalling
- docs/troubleshooting.md, it broke, now what
- docs/dashboard.md, the dashboard, a guided tour with screenshots
- docs/usage.md, the mirror commands and branch safety
- docs/knowledge-layer.md, the graph, connectors, search, wiki
- docs/ask-the-graph.md,
kb query,kb impact,kb owners - docs/serve.md, serve the graph over MCP + wire your editor
- docs/keep-fresh.md, bootstrap, scheduling, and re-indexing on commit
- docs/explained.md, what changes for you, and why it is built this way
- docs/benchmarks.md, where the token/cost/correctness impact comes from, and how to measure it yourself
- docs/internals.md, architecture and internals
- docs/releasing.md, maintainer runbook: versioning, tagging, publishing
- CHANGELOG.md · ROADMAP.md · CONTRIBUTING.md · BRANDING.md
License
MIT, see LICENSE. Pebble the otter is the project mascot; deep context, clear answers.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file contextlake-6.7.0.tar.gz.
File metadata
- Download URL: contextlake-6.7.0.tar.gz
- Upload date:
- Size: 1.9 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
414992f1f843a455ed601b204730597adc2084fe7860408f633d760ea39f7e3f
|
|
| MD5 |
02b1225947bed54e8c91881821579b9c
|
|
| BLAKE2b-256 |
70785e834c8ddc816dbc920f2e362505a0fb469fe153e3ba32457b22e62feffc
|
Provenance
The following attestation bundles were made for contextlake-6.7.0.tar.gz:
Publisher:
release.yml on sayak-sarkar/contextlake
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
contextlake-6.7.0.tar.gz -
Subject digest:
414992f1f843a455ed601b204730597adc2084fe7860408f633d760ea39f7e3f - Sigstore transparency entry: 2412495358
- Sigstore integration time:
-
Permalink:
sayak-sarkar/contextlake@de74382c55a3df7387acbd25ba2a6859f6b68827 -
Branch / Tag:
refs/tags/v6.7.0 - Owner: https://github.com/sayak-sarkar
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@de74382c55a3df7387acbd25ba2a6859f6b68827 -
Trigger Event:
push
-
Statement type:
File details
Details for the file contextlake-6.7.0-py3-none-any.whl.
File metadata
- Download URL: contextlake-6.7.0-py3-none-any.whl
- Upload date:
- Size: 1.8 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
44ff9323ae70b87f6dd9a0472b5d011ac491b9fe172826ff14e953120fd34d81
|
|
| MD5 |
dbe103877e9f57106bb77867f856f95d
|
|
| BLAKE2b-256 |
b770a3fb9feb912312b2f6738c02f10b1ec7b679e444e30871e90d58b9ff37da
|
Provenance
The following attestation bundles were made for contextlake-6.7.0-py3-none-any.whl:
Publisher:
release.yml on sayak-sarkar/contextlake
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
contextlake-6.7.0-py3-none-any.whl -
Subject digest:
44ff9323ae70b87f6dd9a0472b5d011ac491b9fe172826ff14e953120fd34d81 - Sigstore transparency entry: 2412495396
- Sigstore integration time:
-
Permalink:
sayak-sarkar/contextlake@de74382c55a3df7387acbd25ba2a6859f6b68827 -
Branch / Tag:
refs/tags/v6.7.0 - Owner: https://github.com/sayak-sarkar
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@de74382c55a3df7387acbd25ba2a6859f6b68827 -
Trigger Event:
push
-
Statement type: