Skip to main content

Lockkeeper

The skill router and prompt-injection firewall for AI coding agents

Give Claude Code, Codex, Cursor and other AI agents the few skills, MCP servers and tools that fit each task, instead of all of them.
Smaller context window, better tool choices, and no unvetted skill instructions reaching your agent.

PyPI tests Python 3.11+ zero dependencies macOS, Linux, Windows License: FSL-1.1-ALv2

Quickstart · Ways to use it · Benchmark · Firewall · FAQ · Docs


What is Lockkeeper?

AI coding agents get better with skills (SKILL.md files), MCP servers, plugins and tools. But every one you install adds to what the agent has to read and choose from. With hundreds installed, your context window fills up before work starts, and the agent often picks the wrong skill or none at all.

Lockkeeper is a local skill router. It indexes everything installed across all your agents, and for each task it hands the agent a small, complementary set, up to 10 capabilities by default (you choose the size), with the exact file to read for each. Before anything reaches your agent, its built-in firewall can check skills and live tool calls for prompt injection.

$ lockkeeper route --runtime claude "migrate the auth module to the new token API"

[primary] skill: api-migration
[context] mcp: context7
[integration] tool: mcp__context7__query_docs
[verification] agent: code-reviewer
[support] skill: python-patterns
context savings: loaded 6 of 7,540 eligible capabilities (7,534 kept out of context)

Lockkeeper routing a payment-webhook audit task to two primary skills in the terminal

Why developers use Lockkeeper

  • 🎯 Better skill choices. On a public benchmark of real agent tasks, Lockkeeper ranks a correct skill first 65% of the time among 26,000 real skills (up from 35%) and 55% among 79,000, no model required. See the benchmark.
  • 📉 A context window that stays small. With 58,018 capabilities in the library, a routed task still carries a median of about 8,700 tokens of skills instead of about 77.5 million. Adding skills to the library doesn't grow your prompt.
  • 🛡 Safer skills and plugins. Scan any skill, plugin or MCP config for hidden instructions and data exfiltration before your agent reads it, and block hostile tool calls live.
  • 🔌 Works where you already work. Automatic routing in Claude Code, an MCP server for Codex, Cursor, Windsurf, Cline and other clients, and a CLI for everything else.
  • 🔒 Local, private and dependency-free. Pure Python standard library. No GPU, API key or cloud service needed. Telemetry is off unless you say yes.

Quickstart

1. Install from PyPI (Python 3.11+, macOS, Linux and Windows):

pipx install lockkeeper     # or: pip install lockkeeper  ·  uv tool install lockkeeper
lockkeeper init             # finds every AI agent on this machine and connects it

Prefer not to use a terminal? Paste the prompt in PROMPT.md into the AI agent you already use; it installs and configures Lockkeeper for you. Working from source? git clone https://github.com/Hannay001/lockkeeper.git && cd lockkeeper && ./install.sh

2. Index what you have installed:

lockkeeper rebuild    # indexes every skill, agent, command, MCP server and plugin it finds
lockkeeper doctor     # shows each agent found and how many skills it has

3. Route a task:

lockkeeper route "write unit tests for a python data pipeline"

Then pick how your agent should use it, below.

Four ways to use Lockkeeper

1. Route every prompt automatically (Claude Code)

lockkeeper hooks install claude

Every prompt you send now reaches Claude Code with a short note naming the installed skills that fit it and the exact files to read. Slash commands and short replies like "thanks" pass through untouched, and the hook never blocks a prompt. Undo with lockkeeper hooks remove claude.

2. As an MCP server (Codex, Cursor, Windsurf, Cline and any MCP client)

lockkeeper mcp gives your agent three tools, route, search and audit, and keeps the index loaded between calls so answers are fast.

claude mcp add lockkeeper -- lockkeeper mcp          # Claude Code
# Codex: ~/.codex/config.toml
[mcp_servers.lockkeeper]
command = "lockkeeper"
args = ["mcp", "--runtime", "codex"]
{ "mcpServers": { "lockkeeper": { "command": "lockkeeper", "args": ["mcp"] } } }

The JSON form works for Cursor (~/.cursor/mcp.json), Windsurf, Cline and most other clients.

3. From the command line and scripts

lockkeeper route --runtime codex "add rate limiting to a REST endpoint"
lockkeeper search "pdf tables"
lockkeeper route --json --stdin < task.txt      # whole prompts, machine-readable output

4. As a firewall for skills and plugins

lockkeeper audit ~/Downloads/some-skill --recursive --strict   # exit 2 = hostile
lockkeeper hooks install claude --firewall                      # block hostile tool calls live

Supported agents

Agent Skills and tools indexed How the agent gets its routes
Claude Code ✓ Automatically on every prompt (hooks install claude), or MCP
OpenAI Codex CLI ✓ MCP (lockkeeper mcp) or CLI
Cursor, Windsurf, Cline ✓ MCP
GitHub Copilot, Gemini CLI, OpenCode ✓ MCP
Jcode, Hermes ✓ MCP or CLI

Lockkeeper reads the formats you already use: SKILL.md Agent Skills, agents and commands in Markdown, plugin manifests, and MCP server configs. The installer also detects agent tools it doesn't know by name.

Proven on a public benchmark

Routing claims should be measurable. Lockkeeper is tested against SkillRouter Eval Core, the public benchmark from the SkillRouter paper (arXiv:2603.22455): 75 real agent tasks with known correct skills, hidden among real SKILL.md files from public repositories, including 780 deliberately misleading look-alikes.

Before this release Lockkeeper today
Correct skill ranked first, 26,000 skills 34.7% 65.3%
Correct skill ranked first, 79,141 skills 25.3% 54.7%
Needed skills included in the routed set (79k) 20.6% 52.1%
Time to route a ~180-word task, 26k skills 6.5 s 0.7 s

On the full pool, Lockkeeper's standard-library ranker scores between the paper's general-purpose embedding models (Qwen3-Embedding-0.6B at 53.3%, Gemini embedding at 56.0%) and roughly double its BM25 keyword baseline (28.0%), without loading a model. Methods, per-change results and caveats: docs/BENCHMARK.md.

Reproduce it yourself (downloads the ~400 MB dataset once):

python3 scripts/bench_routing.py prepare --home /tmp/lk-bench --size 26000
python3 scripts/bench_routing.py run --home /tmp/lk-bench

Your prompt stays flat as your library grows. On the 79,141-skill benchmark pool (about 157M tokens of skill text), six everyday tasks each routed to 10 capabilities: a median of about 16,000 tokens even if the agent reads every one, over 99.98% kept out of context. The 26,000-skill pool gave about the same (17,600). Reproduce with python3 scripts/bench_context_savings.py.

How it works

flowchart LR
    T["Your task or prompt"] --> S["Rank every installed capability<br/>names, descriptions, body keywords,<br/>rare words weighted higher"]
    S --> P["Apply your policy<br/>deny lists, required roles"]
    P --> B["Build a bundle<br/>roles first, then close matches,<br/>up to your size (10 by default)"]
    B --> F["Keep only what this<br/>agent can actually run"]
    F --> R["Routed set with<br/>exact files to read"]
  • One index for every agent on your machine, deduplicated, and refreshed automatically when you install, remove or update a skill.
  • Reads what each skill is about, not just its one-line description: the most distinctive words of every skill's body are indexed at rebuild.
  • Bundles, not long lists. A route fills complementary roles (primary method, context, integration, verification, support), then tops up with close matches only, never past your bundle size (10 by default, 3 to 20).
  • Optional upgrades, never required: an embedding sidecar for semantic re-ranking, and a decision-model stage (for example Laya or a cross-encoder) that starts in shadow mode so you can measure it before trusting it.

Prompt-injection firewall for skills and MCP

Lockkeeper audit flagging a skill as hostile for an instruction override and a data-exfiltration pipeline

Skills and plugins are instructions your agent follows. Lockkeeper's scanner finds text that tries to override the agent, commands that send secrets or files to the network, credential-store access, code that decodes and runs hidden payloads, destructive commands, and invisible Unicode, across Markdown, configs and scripts.

  • CI-ready verdicts: clean, suspect, hostile with exit codes 0, 1, 2.
  • Live protection: a Claude Code hook blocks hostile tool calls before they run.
  • Evidence: signed receipts prove what was scanned and that results weren't altered.
  • Optional: dependency CVE checks against osv.dev, and a second-pass LLM review.

Full details: docs/FIREWALL.md.

How Lockkeeper compares

Routes each task Skills, MCP, plugins and tools, across agents Uses what a skill's body says Injection firewall Needs a model or GPU
Loading every skill into context ✗ – ✓, at a huge token cost ✗ no
Built-in skill lists (name and description only) agent guesses one agent ✗ ✗ no
Learned skill routers (e.g. SkillRouter, 1.2B parameters) ✓ skills only ✓ ✗ yes
MCP server managers ✗ MCP only – ✗ no
Skill security scanners ✗ ✗ ✓ ✓ some
Lockkeeper ✓ ✓ ✓ ✓ no

FAQ

How do I stop too many skills from filling my Claude Code context window?

Keep your everyday skills where Claude Code loads them, and put the large collection in a skills library that Lockkeeper indexes but your agent doesn't load on its own. Then run lockkeeper hooks install claude: each prompt arrives with the few library skills that fit it, and the agent reads only those SKILL.md files instead of carrying every description in its context.

Does Lockkeeper work with MCP servers?

Both ways. It indexes the MCP servers and tools your agents have configured and routes to them, and it is itself an MCP server (lockkeeper mcp) that Codex, Cursor, Windsurf, Cline and other clients can call.

How do I check a skill from GitHub for prompt injection before installing it?

Run lockkeeper audit path/to/skill --recursive --strict. A hostile verdict (exit code 2) means don't install it. See docs/FIREWALL.md.

Does Lockkeeper send my prompts or code anywhere?

No. Routing, indexing and auditing run locally. The only network features are opt-in: the osv.dev dependency check, the LLM scan, remote decision providers, and the optional embedding sidecar, which downloads its model once.

Is there telemetry?

Only if you say yes. lockkeeper init and lockkeeper hooks install ask once, in your terminal (never in scripts or CI), and lockkeeper telemetry on|off changes your answer at any time. It shares anonymous daily counts (which commands ran and how fast), never prompts, skill names or file paths, and DO_NOT_TRACK=1 always turns it off. See docs/TELEMETRY.md.

Is Lockkeeper free to use?

Yes, for you and your company, including at work and on commercial projects: use it, change it and share it. What the Functional Source License (FSL-1.1-ALv2) doesn't allow is offering Lockkeeper, or a product built from it, to others as a commercial product or service that competes with it. Each release becomes Apache 2.0 two years after it ships, and versions up to 1.2.0 remain under the MIT license.

Do I need a GPU, an API key or an embedding model?

No. The core uses only the Python standard library. Embeddings and decision models are optional add-ons.

How many skills can Lockkeeper handle?

It's tested with up to 79,141 skills. At typical sizes (hundreds to a few thousand) routing and re-indexing take well under a second to a few seconds.

Will it choose worse skills than my agent would on its own?

Measure it: scripts/bench_routing.py runs the public benchmark, and scripts/eval_decision.py evaluates labeled tasks from your own history.

Documentation

Guide What's in it
Configuration All commands, projects and policy packs, freshness, long prompts, hooks, MCP, optional models
Firewall What the scanner detects, verdicts, receipts, live hooks
Benchmark Methods, full results, comparison with published routers, caveats
Telemetry Exactly what opt-in telemetry collects, and how to turn it off
Architecture How the index, router and firewall fit together
Roadmap What's shipped and what's next

Contributing

Issues, ideas and pull requests are welcome. To run the tests:

HOME="$(mktemp -d)" python3 -m unittest discover -s tests -p "test_*.py" -t .

By submitting a pull request, you agree to license your contribution under the project's license.

Found a security issue or a way past the firewall? Please report it privately per SECURITY.md.


Built and maintained by Himanshu (@Hannay001) · Functional Source License (FSL-1.1-ALv2)

If Lockkeeper saves you context or catches something nasty, a ⭐ helps other developers find it.

Release files for lockkeeper 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lockkeeper 1.3.0
File Size Uploaded
lockkeeper-1.3.0.tar.gz 214.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lockkeeper 1.3.0
File Interpreter ABI Platform
lockkeeper-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 368.5 kB

Release files / lockkeeper-1.3.0.tar.gz

Download URL lockkeeper-1.3.0.tar.gz
Size 214.6 kB
Tags Source
SHA-256 checksum
How to use checksums
89f204da4d855f28e928027beb923b204a874d12aff1e179d616237560598ee2
BLAKE2b-256 checksum
How to use checksums
4d34af18c8cc36dcbc409a9198603fdc6c1f9ce6078377592c6d54b214bb538f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / lockkeeper-1.3.0-py3-none-any.whl

Download URL lockkeeper-1.3.0-py3-none-any.whl
Size 153.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
420b18fbcdb3d052e870a7ece6b68e42e629473240a017df69bf3c54a54def55
BLAKE2b-256 checksum
How to use checksums
75d295c85369a720baf5d00ad2dd8ec17fb2e2aa03b250618a84668d22591f79
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

1.4.0

2 release files

This release

1.3.0 This release

2 release files

1.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page