Skip to main content

kvault

Persistent, structured memory for AI agents — plain Markdown, a CLI, zero services.

pip install knowledgevault

Your agent creates nodes (people, projects, notes), keeps every parent summary a rollup of what's below, and orients itself with one cheap command:

$ kvault tree
. « Knowledge Base » [3 children, 11 total] ~2026-06-07
  notes [1 children, 1 total] ~2026-04-11
    reading_list ~2026-04-11
  people [2 children, 5 total] ~2026-06-02
    contacts [2 children, 2 total] ~2026-06-02
      mike_torres ~2026-01-20
      sarah_chen ~2026-06-02
    friends [1 children, 1 total] ~2026-03-14
      alex_rivera ~2026-03-14
  projects [2 children, 2 total] ~2026-06-07
    launch_plan « Launch Plan — v2 » ~2026-06-07
    website_redesign ~2026-05-28

One outline line per node: title, size, and most-recent activity — about 15 tokens each, so a several-hundred-node KB orients an agent for a few thousand tokens. Anything pruned by --depth or --max-children is called out in place (…74 nodes below), so a partial view can never silently hide content.

Built for developers using AI coding tools who want their agent to remember things between sessions — contacts, projects, meeting notes, research — in a structured, navigable format. kvault needs no API keys, no hosted service, no database: any agent that can run shell commands can use it.

How it works

  • A node is a directory containing a single _summary.md — YAML frontmatter plus Markdown. Leaf nodes are entities (a person, a project); parent nodes summarize their descendants.
  • Parent summaries are the index. Every level is a comprehensive rollup of the subtree below it, written by the agent itself. Navigation is top-down reading, not blind grepping.
  • Writes propagate. kvault write returns the full ancestor chain so the agent rewrites those summaries in one follow-up call — the "2-call write workflow."
  • The KB instructs the agent. kvault init generates an AGENTS.md with the workflow, the rules (search before create, never fabricate, propagate everything), and a periodic maintenance playbook.

Quickstart (30 seconds)

pip install knowledgevault
kvault init ./my_kb --name "Your Name"

Then tell your agent:

"Use kvault CLI commands to manage my knowledge base at ./my_kb"

The agent reads the generated AGENTS.md and starts working.

Tool Setup
Project-instruction agents Keep AGENTS.md in the KB root so the agent reads the workflow automatically
Terminal agents Tell the agent: "Read AGENTS.md for the kvault workflow, then use shell commands to manage ./my_kb"
Custom-instruction agents Paste the generated AGENTS.md workflow into the workspace or system instructions

Agent skill included. skills/kvault/SKILL.md carries the full workflow in the portable SKILL.md agent-skills format, so the agent loads it on demand from any directory — no per-KB setup. Install it wherever your tool discovers skills:

# Claude Code
cp -r skills/kvault ~/.claude/skills/kvault

# OpenClaw (per workspace)
cp -r skills/kvault ~/.openclaw/workspace/skills/kvault

# Other agents: copy into your tool's skills directory, or paste the
# SKILL.md body into its custom instructions

Already have data? Point your agent at an export from a chat, email, or notes tool — see docs/importing-data.md.

The 2-call write workflow

# Call 1: write the node (stdin = frontmatter + markdown body)
kvault write people/contacts/sarah_chen --create --reasoning "Met at NeurIPS" --json --kb-root ./my_kb <<'EOF'
---
source: manual
aliases: [Sarah Chen, sarah@example.com]
---
# Sarah Chen
Research scientist at Acme AI...
EOF
# → {"success": true, "changed": true, "did": "created people/contacts/sarah_chen",
#    "notes": [{"code": "autofilled", "text": "name=Sarah Chen", ...}],
#    "ancestor_paths": ["people/contacts", "people", "."],
#    "ancestors": [{path, current_content}, ...], "journal_logged": true}

# Call 2: the agent rewrites the returned ancestors, including root
kvault update-summaries --json --kb-root ./my_kb <<'EOF'
[
  {"path": "people/contacts", "content": "# Contacts\n...updated..."},
  {"path": "people", "content": "# People\n...updated..."},
  {"path": ".", "content": "# Knowledge Base\n...updated..."}
]
EOF

# A small change to a long rollup can go as exact edits instead of the whole body
kvault update-summaries --json --kb-root ./my_kb <<'EOF'
[{"path": "people", "patches": [{"old_str": "12 contacts", "new_str": "13 contacts"}]}]
EOF

Each old_str must match the body exactly once and patches apply in order; a miss leaves that summary unwritten (the batch reports it in errors and continues). kvault write <path> --patches edits one node the same way. update-summaries rewrites only summaries that exist, and an item's meta merges onto the frontmatter.

In human mode the same write narrates its decisions under the receipt:

Created: people/contacts/sarah_chen
  autofilled  name=Sarah Chen
Journal: journal/2026-08/log.md
Ancestors to update: 3  (people/contacts, people, .)

Re-sending identical content is a detected no-op — the file is not rewritten, mtime and created/updated stay put, so the recency signal in the tree stays honest:

Unchanged: people/contacts/sarah_chen
  unchanged   body and metadata identical — file not rewritten, created 2026-08-10, updated 2026-08-10 preserved
Journal: journal/2026-08/log.md

Required frontmatter: source, aliases — kvault stamps created/updated automatically.

What kvault tells you

kvault reports what it decided, not what you asked for. A note is emitted only when kvault invented a value, deliberately changed nothing, half-failed, hid something, or fell back — silence means the operation went exactly as asked. Notes render as indented lines under the receipt (human mode) and as a notes array in --json and over MCP, each {code, text, level} plus detail, why and next when it has them — they ride in the JSON at every tier, and why/next print in human mode at --explain (compact search results keep code, text and next only). Batch commands collapse repeated notes by code ({code, count, examples}), so a 40-ancestor maintenance run emits one line, not forty. The vocabulary is closed — 11 codes:

Code Contract
autofilled kvault invented a value you did not supply
unchanged the operation ran and deliberately changed nothing
partial part succeeded, part did not; manual repair needed
created something came into existence as a side effect
removed something was destroyed, with a count
truncated you are not seeing everything that matched or exists
skipped kvault could not read something and continued without it
waited kvault blocked on, or broke, another process's lock
guessed an input was unusable and a fallback was chosen
propagate summaries are stale because of this operation: ancestor rollups, or other nodes that reference what moved
structure this write changed the tree's shape in a way worth a look: a near-duplicate name, a parent past the child ceiling, or a new root category

Tiers. -q/--quiet (receipt and warnings only) → normal → --explain (adds each note's why and the exact next command) → --trace (adds lock waits and mechanics). Flags work before or after the subcommand; KVAULT_VERBOSITY=quiet|normal|explain|trace sets the tier for hooks and cron jobs (flags win; a typo silently means normal). partial notes survive even --quiet — silencing a half-failure on request is a footgun.

--strict exits 3 when any warning-class note (partial, skipped, a broken lock) was emitted — for CI and unattended runs. check rejects it: its exit codes are already a contract, and its human output is frozen.

Durable ops log. Every successful mutating command (CLI and MCP) appends one row to .kvault/logs.db: op, path, did, notes, changed/partial flags, duration, session. kvault log tail shows what this KB's other agents and sessions did recently; KVAULT_SESSION groups the commands of one logical task, KVAULT_OPS_LOG=0 disables. A failed append can never fail a write — the CLI surfaces the miss as a skipped note.

Full 0.13.0 detail — every new JSON field, per-command changes, frozen surfaces — is in the CHANGELOG.

Structure that holds up under many agents

A KB rots in a specific way when several agents write to it: flat parents with dozens of children, infra/ beside infrastructure/, one new root category per sub-problem, and directories with no summary that no surface can see. kvault guards the write and audits the tree with the same rules, so the two never disagree.

At write time (kvault write --create, MCP kvault_write_node):

  • A create that would add a root category is refused unless you pass --new-root.
  • A create whose name has the same words as a sibling (release_note beside release_notes) is refused unless you pass --allow-similar. Near-duplicates (crm_prompts beside crm_prompts_legacy, infra beside infrastructure) and a parent pushed past 10 children are reported as a structure note, with the node to read next.
  • Missing intermediate parents get a stub summary (a created note; it says "Placeholder" until you rewrite it) instead of becoming an invisible directory.

In the audit (kvault check), one bounded group per prefix, all warn-only:

Prefix Meaning Fix
BRANCH: (hard) A parent — the root included — has more than 10 children kvault plan <path>
SUMMARY: A parent rollup is too short, misses children, has placeholder text, is too long, or accretes dated sections Rewrite as a rollup; too_long/stale_history means fold, never split — chronology belongs in journal/, detail in deep_context/
GHOST: A directory with no _summary.md — invisible to tree, search, and check Write a summary, or list it in .kvaultignore if it is tooling
SERIES: A parent whose children differ only by date or time words: a chronology written as nodes kvault plan folds them: the dated nodes become deep_context/ material of one current-state node you then write; new timeline entries go to journal/
SIBLINGS: Two sibling names under one parent share their words Merge, or nest one with kvault move
DUPLICATE: The same thing filed in two places anywhere in the KB: the same name at two depths, the same title, shared aliases, or near-identical text (buckets like a_m, tier × segment facets, homonyms, rollups and date series are exempt) Read both; same thing → fold into one node and park the other under its deep_context/ (one move to undo); different things → kvault mark --distinct-from
DANGLING: A summary links to, or lists as a child, a path with nothing there (moved, deleted, or never created) Point it at the node's current path (the finding names same-name nodes elsewhere) or drop it
LOOSE: A file outside the node convention: a legacy node file (Markdown with frontmatter, invisible to search), a supporting doc, or an artifact Adopt it as a node (plan emits kvault write <node> --create < file && git rm file), move it into <node>/deep_context/, or ignore it
JOURNAL: Files off journal/YYYY-MM/log.md, or a second history Fold into the canonical log, or list a deliberate second layout in .kvaultignore
STALE: A node's verify_by date has passed: it records facts that go stale (a pending change, an open review) Re-check them and rewrite what changed; still time-sensitive → kvault mark <path> --verify-by +14d; settled → --verify-by none
PENDING: / RETRACTED: Captured events never promoted; nodes citing retracted events Promote or resolve; rewrite and re-link

.kvaultignore at the KB root (one fnmatch pattern per line; a directory pattern covers its subtree) declares the tooling directories and files that are not nodes and are fine. kvault check --code DUPLICATE (repeatable) runs and reports one code; --max-findings 0 and --max-lines 0 return and print its full list. move and delete name the other summaries that still point at the old path (referrer_paths).

One path in. The guards run only on kvault write and the MCP write tools. A node written with a file tool or a shell redirect skips them and never reaches the ops log; on a real KB, 56 of the 59 nodes created in a month arrived that way. Agents and pipelines create nodes through kvault, and check is the backstop for the ones that did not.

Corrections stick. Nobody reviews a KB on a schedule; the owner corrects the agent in use, after the fact. A correction that only lives in a chat is re-proposed the next week, so it is recorded in the node's frontmatter, where every rule reads it: kvault mark <a> --distinct-from <b> (different things; the sibling and duplicate findings and the create guard stop for that pair), kvault mark <parent> --max-children N (this parent is meant to be this wide), kvault mark <parent> --series-ok (this chronology is intentional), kvault mark <node> --verify-by +14d (re-check this node's time-sensitive facts by then). Every plan question carries a default, and a batch may carry --new-root only when it does not increase the root count, so consolidation runs unattended and nothing waits on a person.

The engine: kvault plan turns findings into an ordered worklist with the exact commands — parents over the ceiling are clustered by leading word into new parents, each with a ready kvault move --batch payload; then ghosts, date series, duplicates, sibling collisions, dangling references, loose files, journal drift, stale facts, and summary rewrites. It never applies anything, and the judgment calls (are ml and machine_learning one initiative?) come back as questions. kvault move --batch --confirm runs a JSON list of {from, to} under one lock with one combined propagation list.

$ kvault plan --limit 2
Plan for .: 9 items (showing 2; --limit 0 for all)
1. cluster  projects → projects/atlas
     projects has 46 direct children (ceiling 10); 6 share the leading word 'atlas'
     kvault move --batch --confirm --kb-root /home/me/my_kb <<'EOF'
     [{"from": "projects/atlas_architecture", "to": "projects/atlas/atlas_architecture"}, …]
     EOF
     then: rewrite projects/atlas as a rollup of its 6 children
2. ghost    infra
     no _summary.md — invisible to tree, search, and check
     …
Questions (answer from evidence; defer only when the evidence is not there):
  - projects: are any of these groups one initiative? atlas, billing, ml, crm, support, …

The cadence — every session, nightly, weekly, monthly — lives in skills/kvault-maintenance/SKILL.md, written so a cron or systemd job can load it alone.

kvault check also catches stale propagation, and works as a pre-prompt hook:

{
  "hooks": {
    "UserPromptSubmit": [
      {"type": "command", "command": "kvault check --kb-root /absolute/path/to/my_kb"}
    ]
  }
}

CLI reference

Category Commands
Orient & discover kvault tree [path] [--depth N] [--max-children N] [--gist], kvault search "<query>" [--compact] [--parents gist]
Nodes kvault read <path>… [--parents gist], kvault write (stdin) [--new-root] [--allow-similar], kvault list, kvault delete, kvault move [--batch --dry-run]
Summaries kvault read-summary, kvault write-summary (stdin), kvault update-summaries (stdin JSON), kvault ancestors
Quality kvault validate, kvault check [--code X] [--max-findings N|0] [--max-children N], kvault plan [PATH] [--limit N], kvault mark <path> [--distinct-from X] [--max-children N] [--series-ok] [--verify-by DATE]
Journal & artifacts kvault journal, kvault artifact daily, kvault log tail, kvault log summary
Lifecycle kvault init, kvault status

Agent-facing commands accept --json for machine-readable output and --kb-root (auto-detected from cwd by default), before or after the subcommand. The output flags -q/--quiet, --explain, --trace, and --strict go on the group (kvault --strict write …) or after a note-reporting subcommand (write, write-summary, update-summaries, delete, move, mark, journal, search); see What kvault tells you.

MCP server (optional)

The CLI is the primary interface. For MCP-native clients, a stdio compatibility server ships with the [mcp] extra (Python 3.10+), bound to one KB root per process:

pip install "knowledgevault[mcp]"
kvault-mcp --kb-root /absolute/path/to/my_kb
{
  "mcpServers": {
    "kvault": {
      "command": "kvault-mcp",
      "args": ["--kb-root", "/absolute/path/to/my_kb"]
    }
  }
}

It exposes the same operations as the CLI (kvault_tree, kvault_search, kvault_read_node, kvault_write_node, kvault_capture and kvault_events, summary/journal/validation tools, kvault_log_tail for the ops log), plus a strict parent-summary workflow with stale-write detection. Results carry the same did/notes decision reporting as --json, placed before the bulk payload. The write, move and delete tools accept ancestors="content"|"paths": "paths" (the default) keeps ancestor_paths but omits the full ancestors[].current_content payload, which can exceed 45,000 characters on a mature KB; pass "content" to inline it. kvault_write_node and kvault_update_summaries take patches for exact edits. Unknown arguments are refused rather than ignored. Eight entity-era tools (kvault_read_entity, kvault_write_entity, ...) are registered only with --legacy-tools or KVAULT_MCP_LEGACY_TOOLS=1. Set KVAULT_ALLOWED_ROOTS to pin allowed roots on shared runtimes. Protocol details: ARCHITECTURE.md.

Every signal above is on the MCP surface too: kvault_check returns the check document (codes=[...] with max_findings=0 for one code's full list), kvault_plan the worklist, kvault_move_entities runs a batch, kvault_mark records a decision (verify_by included), kvault_write_node takes new_root and allow_similar, and kvault_prepare_summary_update returns child gists past the ceiling (children="content" for full bodies). kvault_validate_kb is integrity only.

Results stay small over MCP: the defaults fit clients that inline about 4 KB of tool output, and a result cut to fit says so in a truncated note. kvault_tree shows the deepest outline that fits max_chars (3,500 by default). kvault_search returns 8 compact hits by default (path, title, kind, date, one-line snippet; compact=false for scores and long snippets). A match under deep_context/ is folded into its node when that node matches about as well and is in the results (include_background=true lists them). parents="gist" on search and reads gives each ancestor's path, title, and first line, where parents="all" attaches every ancestor's full document to every hit (capped by total_max_chars). kvault_read_nodes reads up to 25 picked hits in one call under one budget that counts whole nodes (3,500 characters by default over MCP; raise it when the client can take more). kvault_status and kvault_search keep their results under max_chars the same way (0.17.1); kvault_check returns findings once, without the CLI document's per-category copies, and kvault_validate_kb carries each issue type's message and fix once.

It's just files

kvault produces Markdown with YAML frontmatter in a plain directory. No proprietary format, no database to export from. Your existing tools work out of the box:

Want to... Use
Semantic search Embed the .md files with any vector tool
Exact text search rg -n "phrase" ./my_kb
Visual browsing Open the KB directory in Obsidian or Logseq
Publish as a site Point Hugo, Jekyll, or Astro at the directory
CI validation Run kvault validate or kvault check in a GitHub Action
Bulk export find . -name _summary.md + yq over the frontmatter

Python API

from pathlib import Path
from kvault.core import operations as ops

kg_root = Path("my_kb")
outline = ops.build_outline(kg_root, depth=2)          # annotated tree as nested dict
node = ops.read_node(kg_root, "people/contacts/sarah_chen")
result = ops.write_node(kg_root, "people/contacts/new_person", "# Content", create=True)
matches = ops.search_nodes(kg_root, "sarah follow up")

Development

pip install -e ".[dev,mcp]"
pytest -q
ruff check .
black --check kvault/ tests/
mypy kvault/ --ignore-missing-imports

License

MIT

Metadata

Release files for knowledgevault 0.17.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for knowledgevault 0.17.1
File Size Uploaded
knowledgevault-0.17.1.tar.gz 279.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for knowledgevault 0.17.1
File Interpreter ABI Platform
knowledgevault-0.17.1-py3-none-any.whl Python 3 none any Details

Total release size: 479.0 kB

Release files / knowledgevault-0.17.1.tar.gz

Download URL knowledgevault-0.17.1.tar.gz
Size 279.5 kB
Tags Source
SHA-256 checksum
How to use checksums
cf02db506906d521f0b0a118fe6bac6f59dad357efcb29bf149f6933a3af5923
BLAKE2b-256 checksum
How to use checksums
c14936f15b360e7cb14f37c38baa1f7f7f90443118fd66e7dd19fa479765f978
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / knowledgevault-0.17.1-py3-none-any.whl

Download URL knowledgevault-0.17.1-py3-none-any.whl
Size 199.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0c3f9681e7546030b1187f45273aee02d3fa193fc17306c064db622da430a058
BLAKE2b-256 checksum
How to use checksums
34e076d62ae7902a497a8e15014cadfbbb28c6cd14b61422f54f5ffa738aedfc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.17.1 This release

2 release files

0.16.1

2 release files

0.16.0

2 release files

0.15.2

2 release files

0.15.1

2 release files

0.15.0

2 release files

0.13.0

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.3

2 release files

0.11.2

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.2

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page