kbx — Local Knowledge Base with Hybrid Search
Give your AI agents persistent memory. Index your markdown notes, meeting transcripts, and documentation into a hybrid search engine. Search with keywords or natural language. Everything runs locally — your data never leaves your machine.
kbx combines SQLite FTS5 full-text search with LanceDB vector search using Qwen3 embeddings — all on-device, with Apple Silicon acceleration via MLX.
You can read more about kbx's progress in the CHANGELOG.
Quick Start
# Install
pip install kbx # core CLI + FTS5 search
pip install "kbx[search]" # + vector search (Qwen3 embeddings)
pip install "kbx[search,mlx]" # + Apple Silicon acceleration
# Set up a knowledge base
kbx init # create kbx.toml in the current directory
# Index your markdown files
kbx index run # index everything under memory/
kbx index run --no-embed # text-only index (fast, no model needed)
# Search
kbx search "quarterly planning" # hybrid search (FTS5 + vector)
kbx search "quarterly planning" --fast # keyword-only (~instant, no model needed)
kbx search "MFA rollout" --json # structured output for scripts
# Browse
kbx view "memory/notes/decisions.md" # read a document
kbx view "#a1b2c3" # by content-hash prefix
kbx list --type notes --from 2026-01-01
Using with AI Agents
kbx is built for agentic workflows. The --json output format, structured error responses, and built-in agent playbook make it a natural fit for AI assistants.
# Orient: get a compressed overview of all entities (~2K tokens)
kbx context
# Search with structured output
kbx search "authentication" --fast --json --limit 5
# Look up a person
kbx person find "Wren" --json
# Timeline of everything mentioning a project
kbx person timeline "Helix Refactor" --from 2026-01-01 --json
# Take notes that persist across sessions
kbx memory add "Decision: use Postgres" --tags decision,infra --pin
kbx memory add "Promoted to Staff" --entity "Soren"
# Pin important docs to the context window
kbx pin "memory/notes/priorities.md"
When you run kbx --help, it prints an agent playbook alongside the standard CLI help — a complete reference for AI agents to self-orient and use the knowledge base effectively.
MCP Server
kbx exposes an MCP server for tighter integration with Claude Desktop, Claude Code, Cursor, and other MCP-compatible tools.
Tools exposed:
kb_search— Hybrid or FTS-only search with date/tag filterskb_person_find— Entity lookup by name, alias, or partial matchkb_person_timeline— Chronological document list for an entitykb_view— Retrieve a document by path, glob, or#hashkb_context— Compressed entity index for session orientationkb_memory_add— Create notes or record facts about entitieskb_pin/kb_unpin— Pin documents to the context windowkb_usage— Index status and usage instructions
Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"kbx": {
"command": "/Users/YOU/.local/bin/kbx",
"args": ["mcp"]
}
}
}
Note: Claude Desktop does not inherit your shell PATH. Use the full path to
kbx— find it withwhich kbx(typically~/.local/bin/kbxwhen installed viauv tool install).
Claude Code (.claude/settings.local.json):
{
"mcpServers": {
"kbx": {
"command": "kbx",
"args": ["mcp"],
"type": "stdio"
}
}
}
See MCP plugin docs for full tool parameter reference.
Python API
Use kbx as a library in your own applications:
from kb import KnowledgeBase
with KnowledgeBase(thread_safe=True) as kb:
# Search
results = kb.search("cloud migration")
# Entities
people = kb.list_entities(entity_type="person")
alice = kb.get_entity("Wren")
timeline = kb.get_entity_timeline("Wren")
# Context
ctx = kb.context()
# Index
kb.index()
The KnowledgeBase class manages the full lifecycle — DB connections, embedder, auto-reindexing of stale files. All methods return Pydantic models.
See architecture docs for the full API surface.
Architecture
Write-through principle: Markdown files are the source of truth. All data writes go to flat files first; the database is a derived index rebuilt from those files. The DB is disposable — delete it and re-index.
Markdown files (source of truth)
│
▼
┌─────────────────────────────────────────────────────┐
│ Source Adapters │
│ meetings.py — walk memory/meetings/YYYY/MM/DD/ │
│ memory.py — walk memory/people/, projects/, ... │
└────────────────────────┬────────────────────────────┘
│ ParsedDocument
▼
┌─────────────────────────────────────────────────────┐
│ Indexer │
│ chunk → embed → store → link entities │
└──────────┬──────────────────────────┬───────────────┘
│ │
▼ ▼
┌──────────────────┐ ┌─────────────────────────────┐
│ SQLite │ │ LanceDB │
│ docs, chunks, │ │ Qwen3-Embedding-0.6B │
│ FTS5, entities, │ │ 1024-dim vectors │
│ facts, mentions │ │ float32, instruction-aware │
└──────────────────┘ └─────────────────────────────┘
│ │
└────────────┬─────────────┘
▼
┌─────────────────────────────────────────────────────┐
│ Hybrid Search │
│ FTS5 (BM25) + Vector → RRF Fusion → Recency Weight │
└─────────────────────────────────────────────────────┘
Search
kbx supports two search modes:
| Mode | Flag | Speed | Method |
|---|---|---|---|
| Fast | --fast |
~instant | FTS5 keyword search only |
| Hybrid | (default) | ~2s | FTS5 + vector search + RRF fusion |
Hybrid search uses Reciprocal Rank Fusion (RRF) to combine keyword and semantic results, with a 90-day half-life recency weight. A strong-signal fast path skips vector search entirely when FTS5 produces a high-confidence match.
Score interpretation: 0.8+ strong | 0.5–0.8 worth reading | <0.5 noise
See search docs for the full pipeline, score normalisation, and fusion strategy.
Entity System
kbx automatically links people, projects, teams, and glossary terms to your documents:
kbx person find "Wren" --json # profile + linked documents
kbx person timeline "Wren" # chronological mentions
kbx person create "Soren" --role "SRE Lead" --team "Platform"
kbx project find "Helix Refactor" # project profile + linked docs
kbx entity stale --days 30 # entities not mentioned recently
Entities are seeded from memory/people/*.md and memory/projects/*.md files, then linked to documents via five-tier matching: YAML tags → title participants → title substrings → source IDs → content name matching.
See entity docs for the full linking pipeline.
Sync & Ingest
Pull meeting transcripts from external sources:
# Granola public-API sync (default — Bearer key)
kbx sync granola --since 2026-01-01
# Notion AI Meeting Notes sync
kbx sync notion --since 2026-01-01
# Granola zip export ingest
kbx ingest export.zip
# View and edit synced meeting notes
kbx granola view <calendar-uid>
kbx granola edit <calendar-uid> --append "Action: follow up with Wren"
Sync is incremental — only new or updated meetings are fetched. Attendees are automatically matched to existing entities.
Granola authentication
kbx sync granola calls Granola's public REST API at api.granola.ai/v1. Generate a
Personal API key in the Granola desktop app (Settings → Connectors → API keys —
requires a Business or Enterprise plan), then store it in the macOS Keychain:
security add-generic-password -a "$USER" -s granola_api_key -w "grn_..."
GRANOLA_API_KEY env var is checked first and overrides the Keychain entry. The
deprecated internal-API path (reads ~/Library/Application Support/Granola/supabase.json)
remains reachable via kbx sync granola --legacy for one release window and will be
removed once everyone is on the public API.
Configuration
kbx looks for configuration in this order:
$KBX_CONFIGenvironment variable./kbx.tomlin the current directory (walk up from CWD)~/.config/kbx/config.toml
Run kbx init to generate a starter config.
Optional Extras
| Extra | What it adds |
|---|---|
search |
LanceDB + sentence-transformers + NumPy for vector search |
mlx |
MLX backend for faster embeddings on Apple Silicon |
mcp |
MCP server for AI tool integration |
all |
Everything above plus test and dev dependencies |
Install with: pip install "kbx[search,mlx,mcp]"
Requires Python 3.10+.
Data Storage
Index stored in the data directory (configurable via kbx.toml or $KB_DATA_DIR):
kbx-data/
├── metadata.db # SQLite — documents, chunks, FTS5, entities, facts
└── vectors/ # LanceDB — Qwen3 embedding vectors (1024-dim)
The database is a derived index. Delete it and kbx index run to rebuild from your markdown files.
Development
git clone https://github.com/tenfourty/kbx.git
cd kbx
uv sync --all-extras
uv run pre-commit install
uv run pytest -x -q --cov # 1361 tests, 90%+ coverage
uv run mypy src/ # strict mode
Quick CI check locally:
make ci # mirror exact GitHub CI pipeline
make fix # auto-fix lint + format issues
See CONTRIBUTING.md for guidelines and testing docs for the test strategy.
Documentation
| Doc | What it covers |
|---|---|
| Architecture | System design, data flow, module dependencies, Python API |
| Search | FTS5 + vector + RRF fusion pipeline, score normalisation |
| Entities | Entity seeding, five-tier linking, disambiguation |
| Indexing | Walk → chunk → embed → store pipeline |
| Chunking | Markdown-aware chunking strategy |
| CLI Reference | All commands and options |
| Output Formatting | JSON, table, CSV, JSONL, jq, field selection |
| Context Layer | Compressed entity index for AI agents |
| Testing | Test strategy, fixtures, markers |
| MCP Plugin | MCP server tools and resources |
| MLX Plugin | Apple Silicon embedding acceleration |
| Granola Plugin | Meeting transcript sync (view, edit, push) |
| Notion Plugin | Notion AI Meeting Notes sync |
| Integration | Ingest, migrations, search quality |
License
Apache-2.0
Metadata
Release files for kbx 0.3.11
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| kbx-0.3.11.tar.gz | 614.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| kbx-0.3.11-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 806.4 kB
Release files / kbx-0.3.11.tar.gz
| Download URL | kbx-0.3.11.tar.gz |
|---|---|
| Size | 614.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
187b4727f2c54d67b10effeb0cffb0c1110ba8c4f468ad2bb6de648c1b2ac81d
|
|
BLAKE2b-256 checksum How to use checksums |
1ed7636c50fe13b775e15b03569a8832c57c36f618e1d27319d22ddb98782c64
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 25, 2026.
Transparency logRelease files / kbx-0.3.11-py3-none-any.whl
| Download URL | kbx-0.3.11-py3-none-any.whl |
|---|---|
| Size | 191.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e1f4c2ed31424e309ea35437663a253bd39a3775ccdf09ce8f72b02479a043dd
|
|
BLAKE2b-256 checksum How to use checksums |
6988b16f173b643db30e937ceff5a1899d51c70b3d0468adc0a6e83de75e5444
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 25, 2026.
Transparency log