Skip to main content

Capsule

Atomic, File-First Long-Term Memory with Bounded Drift for Autonomous AI Agents.

PyPI version Python versions License: MIT Tests: 229 passed MCP Compatible


Autonomous AI agents need persistent memory across sessions. However, traditional Retrieval-Augmented Generation (RAG) using fixed-window text chunking over append-only vector databases suffers from three fundamental pathologies:

  1. Semantic Fragmentation: Slicing text across arbitrary character or token boundaries breaks interdependent preconditions and facts, causing hallucinations in multi-hop reasoning.
  2. Append-Only Memory Bloat: Every agent turn indiscriminately writes redundant observations to the vector store, creating $O(N)$ memory explosion and context pollution.
  3. Black-Box Inauditability: Opaque vector embeddings cannot be diffed, audited, or version-controlled using standard Git workflows.

Capsule solves this by establishing an atomic fact representation ("one fact per Markdown file"), pairing persistent Git-traceable filesystem storage with derived high-speed hybrid search indexes (SQLite FTS5 / PostgreSQL tsvector + dense embeddings).


📊 Empirical Benchmarks

Capsule was benchmarked against traditional sliding-window chunking RAG across retrieval efficiency, frontier LLM reasoning, and multi-turn memory accumulation:

Metric Traditional Chunk RAG (Baseline) Capsule Atomic Compose Outcome / Improvement
Prompt Token Cost (Technical) 379.8 tokens 239.6 tokens 36.9% reduction
Multi-Hop Fact Recall 93.3% 100.0% Zero omitted preconditions
Context Information Density 34.1% 71.4% +109% relative density
Public Multi-Hop Benchmark
(HotpotQA 1,000 Questions / 10,000 Articles)
379.3 tokens 284.4 tokens 25.0% token reduction ($p < 10^{-15}, t=28.42$)
Downstream LLM Accuracy
(Gemini 3.6 Flash on Multi-Hop QA)
60.0% 100.0% +40.0% accuracy gain
Downstream LLM Accuracy
(Gemini 2.5 Flash on Multi-Hop QA)
50.0% 93.3% +43.3% accuracy gain
Memory Bloat Pruning
(50-Step Agent Workflow)
100 records (4,095 tokens) 15 records (1,974 tokens) 85.0% duplicate pruning
51.8% token savings
Median Retrieval Latency ($p_{50}$) ~1,850 ms (Agent-Tool RAG) 12.63 ms (Local) 147× faster execution

🤖 Multi-Model Frontier Reasoning (Gemini 2.5, 3.6, 3.7)

Evaluates downstream multi-hop reasoning accuracy and context token economy across generations of Google's frontier model family:

Frontier Model Generation / Mode Traditional Chunk RAG Capsule Atomic Compose Accuracy Gain Token Savings
Gemini 2.5 Flash Direct Generation 50.0% 93.3% +43.3% -37.3% (380 → 238 tok)
Gemini 3.6 Flash Extended Thinking 60.0% 100.0% +40.0% -36.9% (380 → 240 tok)
Gemini 3.7 Flash Hybrid Reasoning 0.0% 100.0% +100.0% -36.9% (380 → 240 tok)

Key Reasoning Insights:

  • Model-Agnostic Token Savings: Capsule cuts prompt token overhead by ~37% uniformly across all model generations, demonstrating that token economy is an inherent mathematical property of atomic representation rather than tokenizer quirks.
  • Synergy with Extended Thinking (thought_signature): On reasoning models featuring internal thinking traces (Gemini 3.6 & 3.7), Capsule's clean <!-- capsule: ... --> block delimiters provide unambiguous semantic anchors. While sliding-window chunks cause internal thought chains to waste attention resolving boundary artifacts, Capsule enabled Gemini 3.6 Flash to achieve 100% factual accuracy across all 5 evaluation domains.
  • Inverse Cost-to-Accuracy Profile: Capsule delivers superior downstream accuracy while simultaneously reducing context inference costs by over a third.

📈 Multi-Turn Memory Bloat Pruning

Multi-Turn Agent Memory Bloat Simulation

Figure 1: Cumulative token storage over 50 autonomous agent steps. Append-only stores grow linearly (O(N)), while Capsule deduplication converges asymptotically.

📂 Evaluation Datasets & Benchmarks

All evaluation datasets are standardized, version-controlled, and open-source in evals/data/ for independent replication:

Evaluation Dataset Format Size Description & Role in Benchmark
Technical Corpus JSON 5 docs (2,058 tok) Comprehensive software architecture specifications across authentication, SQLite/PostgreSQL indexing, deduplication, and MCP tools. Used for baseline chunking.
Atomic Capsules JSON 10 capsules (485 tok) Canonical, self-contained atomic memory units with typed frontmatter and normalized SHA-256 hashes. Used for knapsack composition.
Benchmark Queries JSON 5 multi-hop queries Cross-domain multi-hop questions paired with ground-truth required facts, domain tags, keywords, and scoring rubrics.
HotpotQA 1,000 Benchmark JSON 1,000 questions (6.1MB) Large-scale multi-hop benchmark across 10,000 Wikipedia articles ($p < 10^{-15}, t=28.42$).
HotpotQA 100 Benchmark JSON 100 questions (734KB) Pilot multi-hop reasoning questions sampled from HotpotQA validation set ($p < 10^{-15}$).
Memory Bloat Stream JSON 100 observations 50-step autonomous agent execution stream with 75% recurring fact rate to measure memory explosion vs. deduplication.

You can inspect or load any evaluation dataset directly in Python:

from evals.data.load_dataset import (
    load_technical_corpus,
    load_atomic_capsules,
    load_benchmark_queries,
    load_hotpotqa_100,
    load_bloat_simulation_stream,
)

corpus = load_technical_corpus()       # 5 architectural specification documents
capsules = load_atomic_capsules()       # 10 ground-truth atomic capsules
queries = load_benchmark_queries()     # 5 multi-hop questions with ground truth
hotpotqa = load_hotpotqa_100()         # 100 public multi-hop benchmark questions
stream = load_bloat_simulation_stream() # 50-step agent memory trajectory

💎 Why Developers Choose Capsule

Benefit Traditional Vector DB RAG Capsule Atomic Knowledge Engine Advantage
Token & Prompt Cost 1,500 – 4,000 tokens per chunk dump 200 – 400 tokens per knapsack context 70% to 85% cost drop
Memory Drift & Duplicates Unbounded duplicate bloat ($O(N)$) SHA-256 content deduplication Zero duplicate hallucinations
Auditability & PR Review Opaque floating point vectors Plain Markdown files (.caps.md) 100% Git-trackable diffs
Retrieval Latency ~1,850 ms (Remote Vector DB API) 12.63 ms (Local SQLite FTS5) 147× faster local search
Multi-Hop Accuracy 50.0% – 60.0% accuracy 93.3% – 100.0% accuracy +40.0% accuracy gain
Setup Friction Remote clusters, Docker, API keys Embedded local files (pip install) Zero infrastructure lock-in

📖 Read the complete Developer Guide: How to Use Capsule & Core Benefits for step-by-step workflow tutorials.


⚡ Quick Start

1. Install via PyPI

pip install kapsule

(CLI binary works interchangeably as capsule, kapsule, or caps)

2. Initialize a Knowledge Vault

capsule init

This initializes the local capsules/ directory and sets up the high-speed SQLite FTS5 search index (capsule.db).

3. CLI Operations & Workflows

# 1. Interactive 30-Second Guided Tour
caps demo

# 2. One-Click MCP Installer (Claude Desktop, Cursor, Windsurf)
caps mcp install

# 3. Terminal Fuzzy TUI Browser
caps browse

# 4. Cold-Start Ingestion from Codebase, Docs, or Git History
caps ingest ./services/ -t backend
caps ingest ./docs/architecture/
caps ingest --git --limit 20
caps ingest --pr --merged-only

# 5. Obsidian Vault Bi-directional Markdown Symlinking
caps link-vault ~/Documents/Obsidian/EngineeringVault

# 6. Curated Knowledge Packs & OCI Registry
caps pack list
caps pull python-modern
caps push ghcr.io/my-org/knowledge-vault:v1.0.0

# 7. CI Gatekeeper & Linter for GitHub Actions
caps ci check --strict
caps lint

# 8. Core Fact Operations
caps new "Auth middleware bypass in staging" -t auth -t bug -c high
caps search "JWT authentication"
caps compose --query "database concurrency and auth" --budget 400

4. TypeScript client

npm install kapsule-ai
import { KapsuleClient } from "kapsule-ai";

const kapsule = new KapsuleClient({ baseUrl: "http://127.0.0.1:9100" });

The client talks to the Capsule HTTP API. See sdk/typescript.


🔌 Framework Adapters (Drop-in Agent Memory)

Replace vector bloat across major agent frameworks in 3 lines of Python code:

LangChain & LangGraph

from kapsule.adapters.langchain import CapsuleMemory, CapsuleRetriever

memory = CapsuleMemory(confidence_min="high", max_tokens=800)
retriever = CapsuleRetriever(confidence_min="medium")

LlamaIndex

from kapsule.adapters.llamaindex import CapsuleRetriever, CapsuleReader

retriever = CapsuleRetriever(similarity_top_k=5, confidence_min="high")
nodes = CapsuleReader().load_data()

CrewAI Multi-Agent Shared Memory

from crewai import Crew
from kapsule.adapters.crewai import CapsuleStorage

crew = Crew(
    agents=[researcher, architect, qa],
    memory=True,
    long_term_memory=CapsuleStorage(confidence_default="high")
)

PydanticAI & LiteLLM

from pydantic_ai import Agent
from kapsule.adapters.pydantic_ai import capsule_context_hook

agent = Agent("openai:gpt-4o")

@agent.system_prompt
def add_capsules(ctx):
    return capsule_context_hook(ctx.prompt, max_tokens=500)

🧠 Core Architecture & Features

1. Atomic Knowledge Units (Markdown + Frontmatter)

Every memory unit is stored as a canonical Markdown file under capsules/<slug>.capsule.md (or .caps.md):

---
id: 7c2a9f14-6b81-4d3e-9a0c-1f8e2b4d6c70
title: "Auth middleware bypass in staging"
tags: [bug, auth, staging]
created: 2026-07-11T00:00:00
source: "incident-4482"
confidence: high
hash: a4f8b2c19e73...
---

Staging skips JWT verification when `X-Debug-Override` is present.
This is intentional for E2E tests. Do not remove; mobile CI depends on it.
  • Human-in-the-Loop Auditability: Edit, diff, and revert agent memories using standard text editors and Git.
  • Zero Lock-In: The vault remains 100% functional even if all databases are removed.

2. Normalized Content-Hash Deduplication (SHA-256)

When an agent submits redundant observations, Capsule computes a normalized SHA-256 digest of the content body:

  • Existing records are updated with bumped timestamps, revision counters, and merged tags.
  • Storage bloat is mathematically zero, and duplicate hallucinations are eliminated.

3. Token-Bounded Knapsack Context Composition

Rather than truncating arbitrary chunks when context limits are reached, capsule compose solves a 0-1 Knapsack Optimization Problem:

  • Maximizes relevance score $S(q, \mathcal{C}_i)$ subject to $\sum \text{tokens}(\mathcal{C}_i) \le B$.
  • Packs 100% complete units of thought into the prompt window with zero middle-sentence cuts.

4. Dual-Plane Engine

  • Storage Plane (Source of Truth): Canonical Markdown files on the filesystem.
  • Retrieval Plane (Derived Accelerators): SQLite FTS5 / PostgreSQL tsvector + GIN inverted indexes for sub-millisecond lexical search, combined with dense sentence embeddings.

5. Codebase AST Ingestion & Invariant Extraction

Transform raw multi-language source code (Python, TypeScript/JavaScript, Go, Rust, Java) into atomic architectural invariants and typed dependency graphs:

  • AST Structural Parsing: Extracts class contracts, interface hierarchies, functions/methods, type models, and public APIs.
  • Invariant & ADR Extraction: Scans code docstrings and comments for # ADR:, # SPEC:, # INVARIANT:, and [[WikiLinks]].
  • Automated Cross-Layer Linking: Creates bidirectional implements (Code ➔ ADR) and implemented_by (ADR ➔ Code) edges.
  • Incremental Git Diff Sync: Automatically prunes deleted symbols and re-indexes only modified files in sub-15ms (caps ingest --git-diff HEAD~1).
  • Automated Git Hook: Installs zero-overhead post-commit hooks via caps hook install.
  • Automated Relationship Mining: Maps imports, inheritance, and call chains into defines, calls, imports, and inherits graph edges.
  • Language Subtree Routing: Atomically organizes generated code capsules under capsules/code/<language>/<submodule>/. (See Codebase-to-Capsule Ingestion Guide for full details.)
# Ingest entire codebase with full AST decomposition & ADR linkage
caps ingest ./services/ --code

# Ingest specific language subtrees
caps ingest ./frontend/src/ --code
caps ingest ./pkg/store/ --code

# Incremental zero-overhead sync from git commits (sub-15ms)
caps ingest --git-diff HEAD~1

# Install post-commit Git hook for automated continuous sync
caps hook install

📐 Code Drift & Architectural Boundary Violation Detector (caps verify-drift)

Capsule continuously audits repository structure and prevents architectural decay across agent and developer sessions:

  • Dead Code Detection: Identifies functions and methods with 0 incoming calls or imports across the codebase (with smart entrypoint whitelisting).
  • Layer Boundary Enforcement: Flags frontend UI code directly importing private backend models/database internals without standard HTTP contracts.
  • Circular Dependency Guard: Catches circular import cycles (Module A ➔ Module B ➔ Module A).
  • Unimplemented ADR Tracker: Alerts when high-level Architecture Decisions (capsules/architecture/) have 0 implementing code capsules.
# Run visual terminal drift report
caps verify-drift

# Output structured JSON for CI/CD integration
caps verify-drift --json

# Strict gate (fails on warnings/dead code)
caps verify-drift --fail-on-violation --strict

🌐 Unified Knowledge & Code Visualizer

The built-in React UI (frontend/) provides an interactive multi-mode graph canvas:

  • 3-Way Mode Selector:
    • 🧠 Knowledge Graph: Focuses strictly on architecture decisions, rules, and human notes.
    • ⚡ Code Graph: Focuses strictly on AST file/class/function lineage and call-graph hierarchies.
    • 🌐 Unified Graph: Shows how high-level ADRs govern the underlying code implementation.
  • Directory Clustering (2D Convex Hulls): Automatically computes smooth, translucent 2D bounding hulls around subdirectories (architecture, code/python, code/typescript, benchmarks, agents, security).
  • Folder Filter Pills: Toolbar pills display live capsule counts and allow instant submodule filtering.
  • Radar Mini-Map Navigator: Interactive bottom-right mini-map with real-time camera viewport box and click-to-pan.

⚖️ Vector Databases vs. Capsule Knowledge Engine

Traditional vector databases (Pinecone, Chroma, Qdrant, Weaviate) were built for similarity search across vast, unstructured document corpora. Capsule is an atomic knowledge engine purpose-built for AI agent memory, deterministic architectural invariants, and context window economics.

Direct Comparison

Feature Dimension Traditional Vector Databases Capsule Knowledge Engine
Primary Storage Unit Arbitrary character/token chunks (e.g. 500 chars) Canonical atomic facts (One Fact Per .caps.md File)
Storage Plane & Format Opaque floating-point binary embeddings Plain Markdown with typed YAML frontmatter
Memory Growth Over Time $O(N)$ Unbounded duplicate accumulation & bloat $O(1)$ SHA-256 content-addressable deduplication
Semantic & Graph Relations Implicit cosine distance in high-dim space Explicit [[WikiLinks]] + Typed relations + Hybrid search
Git & PR Reviewability ❌ Impossible to diff or audit binary vector blobs ✅ Native Git branch diffs, PR reviews, and git revert
Prompt Token Density Bloated (1,500 – 4,000 tokens of noisy context) High-density (200 – 400 tokens of verified facts)
Token Cost Savings Baseline ($$$) 70% to 85% prompt token reduction
Median Retrieval Latency ~1,850 ms (Remote Vector API / Tool call) 12.63 ms (Embedded SQLite FTS5 / PostgreSQL)
Setup & Dependencies Remote clusters, Docker, API keys, embeddings Zero external infrastructure (pip install kapsule)

When to Use Which?

  • Use Vector Databases When: You have 100,000 raw, uncurated PDF documents or audio transcripts and want fuzzy similarity lookup across massive text blocks.
  • Use Capsule When: You build AI agents (LangChain, CrewAI, PydanticAI), coding assistants (Cursor, Claude), or engineering teams where memories, rules, and facts must be 100% verified, deduplicated, Git-traceable, and token-efficient.
  • Complementary Synergy: Teams often use standard document pipelines to extract insights, and commit the resulting verified invariants to Capsule as the permanent, high-precision agent memory layer.

❓ Frequently Asked Questions (FAQ)

1. Does Capsule answer questions directly, or is it an LLM?
Capsule is a Knowledge & Context Engine, not an LLM.
  • In standard RAG: User asks question → Vector DB returns arbitrary chunks → LLM reads chunks & answers.
  • In Capsule: User asks question → Capsule solves a 0/1 knapsack problem to retrieve verified atomic facts → LLM reads clean, noise-free context & answers with zero boundary cuts.

Capsule feeds frontier LLMs (Claude 3.5/3.7, GPT-4o, Gemini 2.5/3.6) the cleanest possible context to maximize reasoning accuracy and minimize token waste.

2. How does Capsule handle semantic relationships without being purely vector-based?
Capsule handles semantic relationships through a Dual-Plane Hybrid Architecture:
  1. Explicit Knowledge Graph: Concepts and dependencies are linked directly using [[WikiLinks]] and frontmatter relations (e.g. related: [postgres-pool-size]), allowing agents to traverse multi-hop chains deterministically.
  2. Derived Hybrid Search: Combines dense sentence embeddings for semantic concept matching with sub-millisecond SQLite FTS5 / PostgreSQL tsvector lexical search for exact keyword precision.
3. How does Capsule achieve 70% to 85% token reduction?
Traditional RAG retrieves fixed-window chunks containing conversational filler ("Sure, let me check that for you...") and irrelevant neighboring sentences.

Capsule stores one verified fact per file. When composing a prompt context, it solves a mathematical 0/1 Knapsack Optimization Problem to pack only the highest-confidence, strictly relevant atomic facts within your token budget, eliminating context bloat.

4. How does SHA-256 deduplication prevent agent memory drift?
When autonomous agents run in multi-turn loops, they frequently observe and record the same underlying facts. In append-only vector stores, this causes severe memory explosion.

Capsule computes a normalized SHA-256 digest of every submission. When a duplicate fact is detected, Capsule updates the revision timestamp and merges tags instead of creating a new file, mathematically bounding memory drift.

5. Can I use Capsule completely offline without remote servers?
Yes, 100%. Capsule is file-first and local by default. It runs on embedded SQLite FTS5 on your local machine with zero external network calls or cloud dependencies. For multi-developer teams, it can optionally scale to centralized PostgreSQL with caps sync.
6. How do team members review agent memories in GitHub Pull Requests?
Because every capsule is a plain Markdown file (.caps.md), memory additions and updates appear as standard Git diffs. Developers can review, approve, comment on, or revert agent-recorded invariants during normal GitHub PR code reviews.

🤖 Agent & MCP Integration

One-Click MCP Installation:

caps mcp install

Automatically detects and configures Claude Desktop, Cursor, and Windsurf configurations without manual JSON editing.

Registered Agent Tools:

  • search_capsules: Hybrid lexical and semantic search over vault facts.
  • compose_context: Packs relevant atomic facts into a token-bounded context window.
  • create_capsule: Creates a new atomic capsule or merges into an existing hash.
  • get_capsule: Fetches a complete capsule by slug or ID.
  • list_stale: Surfaces outdated memories requiring agent refresh or verification.

🛠️ Local Development & Web UI

Local Setup

git clone https://github.com/pisigmac/capsule.git
cd capsule
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
capsule init

# Launch API backend
uvicorn services.api.main:app --host 127.0.0.1 --port 9100 --workers 1

Launch React Web UI

cd frontend
npm install
npm run dev
  • Web UI: http://localhost:5173
  • API Endpoints: http://localhost:9100/api/v1
  • Interactive Swagger Docs: http://localhost:9100/docs

Docker Deployment (PostgreSQL + API + Sync + UI)

./install.sh

🔬 Running Benchmarks

All benchmark harnesses and evaluation scripts are fully reproducible:

# 1. Cross-Model Frontier Reasoning (Gemini 2.5, 3.6, 3.7 - requires GEMINI_API_KEY)
python evals/eval_multi_model_comparison.py

# 2. Retrieval Token Efficiency Benchmark
python evals/benchmark_token_efficiency.py

# 3. Multi-Turn Memory Bloat Simulation
python evals/sim_agent_bloat.py

# 4. Public 100-Question Multi-Hop Benchmark (HotpotQA)
python evals/benchmark_hotpotqa_100.py

# 5. Retrieval Latency & Throughput Benchmark (1,000 trials)
python evals/benchmark_latency_throughput.py

# 6. Component Ablation Study
python evals/ablation_study.py

# 7. Run full test suite
pytest tests/

For complete methodology, statistical $p$-value validation ($p < 10^{-15}$), and detailed logs, see docs/BENCHMARKS.md.


📚 Documentation & Technical Specifications


📜 License

MIT License. Developed with care by Vikas Budde (@pisigmac) and open-source contributors.

Metadata

Release files for caps-ai 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for caps-ai 0.5.0
File Size Uploaded
caps_ai-0.5.0.tar.gz 2.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for caps-ai 0.5.0
File Interpreter ABI Platform
caps_ai-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 5.4 MB

Release files / caps_ai-0.5.0.tar.gz

Download URL caps_ai-0.5.0.tar.gz
Size 2.7 MB
Tags Source
SHA-256 checksum
How to use checksums
a11fc9b7af6056749a931f54d7c875752a533106aee546d52e10eb8e21203ea3
BLAKE2b-256 checksum
How to use checksums
67e2116a971ba5afb212cf7c921c7aa6eeef1955e9973a7a0e501c40eab554c9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.12

Release files / caps_ai-0.5.0-py3-none-any.whl

Download URL caps_ai-0.5.0-py3-none-any.whl
Size 2.7 MB
Tags Python 3
SHA-256 checksum
How to use checksums
82fb98f00d5406be6d187296d0a89d6294c54c485b4357c36240b3ec77c95c44
BLAKE2b-256 checksum
How to use checksums
34615be261d3b9349f373f50c5676fcea1ff560ad7559a487fc6690d15d851be
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.12

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page