gaius
Ops memory lifecycle manager for AI coding agents.
Not another RAG chatbot memory — a production-grade system that extracts facts from Claude Code, Gemini CLI, Grok, and Codex sessions, ranks them into an inject-ready corpus, enforces behavioral gates, and prevents you from breaking prod at 3am.
Why gaius — three things most agent-memory tools skip:
- Runs unattended — extract → promote → inject with no human in the hot path; correction is optional.
- Prevents actions, not just recalls them — hard gates
exit:2on force-push, unconfirmed live-trade, prod-delete.- Fully offline — BM25 + sqlite-vec in one SQLite file. No API keys, no cloud.
# CLI (offline core). Extras: [semantic], [mcp]
pip install gaius-memory
# Or with uv / Claude Code MCP (see .mcp.json):
# uvx --from "gaius-memory[mcp]" gaius-mcp
#
# Tip: for unreleased mainline, pin git:
# pip install "gaius-memory @ git+https://github.com/jkubo/gaius"
⚠️ Install name is
gaius-memory, not baregaius. PyPI already has an unrelated package namedgaius(ImmobilienScout24 deploy client v145+). That package owns neither our import path nor our CLI forever, butpip install gaiuswill not install this project. Always:pip install gaius-memory→ importgaius, console scriptgaius.
gaius retire # scan sessions → stage summaries
gaius batch # (optional) review + correct — facts inject by default
gaius quiz # Leitner HITL loop — the only legitimate `confidence_source='human'`
gaius inject --task "what you're working on" # inject context into active session
What It Does
Engineers running Claude Code all day generate enormous amounts of institutional knowledge that vanishes when the context window closes. gaius captures it:
- Extract — scans Claude Code (and Gemini CLI) session JSONLs, extracts compact summaries with typed signals (knowledge, patterns, errors)
- Review (optional) — extracted facts are promoted and inject-eligible by default; review is a correction loop, not a gate before entry.
rejectdrops a bad fact,deferpunts it,agent-reviewclears it from the queue without touching its rank,confirmpins a verified one — while decay, dedup, and mnemosyne health checks keep quality up without a human bottleneck. - Index — promotes extracted facts into a hybrid keyword + semantic SQLite database (
facts.db) - Inject — at task start, retrieves relevant facts and loads skills context into the session
- Coordinate — cross-session claims, shared findings, and a claimable task pool (
gaius concord) so parallel sessions on one machine divide work instead of colliding
Benchmark
The demo below is a regression / smoke check, NOT a quality score: it confirms retrieval still works on a fresh clone. A perfect score on a hand-built 27-fact corpus (same author wrote the facts and the queries) proves the pipeline runs, nothing more. Real quality is the External evaluation section below, on a benchmark gaius didn't build.
bench_inject.py builds a throwaway SQLite corpus from benchmarks/demo_corpus/:
$ python3 benchmarks/bench_inject.py
gaius Injection Benchmark (25 queries, bundled demo corpus)
════════════════════════════════════════════════════════════
Recall: 25/25 (100%) regression check, NOT a quality score
Prec-warn: 3/25
Corpus: 27 demo facts (throwaway SQLite, no daemon, ~0.08s/query)
Semantic mode (embed daemon, real corpus):
Cold start: ~7s (first query, model load)
Daemon warm: ~8ms (Unix socket, model kept in memory)
Storage: sqlite-vec (384-dim, all-MiniLM-L6-v2) + BM25 in a single SQLite file
No API keys, no cloud, runs entirely offline.
External evaluation (LongMemEval-S)
The demo above only proves the pipeline runs. For an unbiased measure, gaius's retrieval is scored on
LongMemEval-S (ICLR 2025): a third-party benchmark whose
500 questions, ~48-session haystacks, and gold labels gaius has no hand in choosing (independent by
construction, not a self-graded demo). Reproducible: python3 benchmarks/bench_longmemeval.py --matrix.
Two metrics, not the same: R@k = per-question fractional recall (stricter; 65% of questions are multi-gold); hit@k = recall_any@k ("at least one gold session in top-k"), the metric published baselines report, so compare hit@k to those, not R@k.
gaius retrieval (all-MiniLM-L6-v2), 500 questions:
| config | MRR | R@5 | R@10 | hit@10 |
|---|---|---|---|---|
| session-semantic | 0.827 | 85.7 | 92.7 | 96.8 |
| turn-semantic | 0.906 | 92.8 | 96.6 | 98.8 |
| turn-hybrid (real inject path) | 0.912 | 92.8 | 95.2 | 98.6 |
On the matched metric (hit@k), gaius at its real regime (short-fact units + the hybrid inject ranking) scores hit@10 98.6%, at par with published same-model (all-MiniLM-L6-v2) session-level retrieval baselines (~96-99% recall_any@10; protocols differ; retrieval-only, not an end-to-end QA metric). Larger embedding models beat all-MiniLM outright, a deliberate tradeoff for gaius's 384-dim, no-GPU, fully-offline footprint. Weakest category: temporal-reasoning (~94% R@10 at turn-semantic).
This measures the retrieval engine. Whether proactive injection improves agent outcomes is a separate evaluation (in progress), and is not claimed here.
Chunked-embedding validation: feeding gaius a whole long session and letting it chunk internally recovers the recall naive whole-session embedding loses to the 256-token cap: chunk-granularity hit@10 98.4% vs naive whole-session 95.2% on a matched 250-question subset, at the turn-level ceiling.
What's Genuinely Novel
| Feature | Description |
|---|---|
| Review & correction loop | retire → stage → promote runs unattended and facts inject by default; reject/defer/confirm plus decay, dedup, and mnemosyne health let you correct the corpus without a human in the hot path. |
Leitner human review (gaius quiz) |
Spaced-repetition boxes over corpus facts. The only legitimate producer of confidence_source='human'; --report is the calibration view (disagreement-rate by domain). Agents are refused on the mutating path. |
| Multi-agent corroboration | Facts confirmed by multiple AI agents (Claude + Gemini) get a 1.5× score boost. Cross-model verification for higher confidence. |
| Session-type behavioral priming | Load different skill sets based on what you're doing: ops, trading, security, code review. |
| Hard enforcement gates | Memory that prevents actions. exit:2 blocks force-push, live trading without confirmation, critical resource deletion. |
| Mnemosyne health monitoring | Automated memory bloat prevention. Line-count thresholds, misclassification audit, split/prune proposals. |
| Live state injection | Domain files with kubectl/curl commands in frontmatter, TTL-cached. Memory that knows what's happening right now. |
| Hybrid sqlite-vec search | Keyword TF-IDF/BM25 + semantic embeddings in a single SQLite file. |
| Chunked embedding | Long facts are split into <=256-token chunks, each embedded; retrieval max-pools over a fact's chunks, so long content isn't silently truncated to its first ~256 tokens. |
| Temporal knowledge graph | Entity-relationship triples with validity windows. gaius kg timeline node shows what changed and when. |
| Obsidian-vault viewer | The knowledge graph exports [[wikilink]] "Related" blocks into your memory files (kg export-links), so the corpus opens directly as an Obsidian vault — browse entities and their links visually, no extra tooling. |
| Cross-session coordination | gaius concord — advisory claims with TTL + pid-liveness, shared findings with an adversarial review loop, a claimable task pool, and a live roster. Parallel sessions get ownership: each can go deep on its lane instead of defensively re-checking the whole world. |
| MCP server | 7 tools for mid-session memory access without leaving Claude Code. |
Quick Start
Install
# Option 1: Claude Code plugin (recommended — wires the hooks, skill, and MCP server for you)
/plugin marketplace add jkubo/claude-plugins
/plugin install gaius@jkubo-tools
/reload-plugins
# Option 2: development install (editable, reads from repo)
git clone https://github.com/jkubo/gaius
cd gaius
pip install -e ".[semantic]" # includes sentence-transformers + sqlite-vec
# Option 3: script install (no pip, uses system/venv python)
install -m 755 gaius_cli ~/.local/bin/gaius
The plugin bundles the gaius skill, the MCP server, and three hooks: corpus
injection at session start, a mnemosyne health check after memory-file edits, and
an optional per-prompt injection (GAIUS_PROMPT_INJECT=1, off by default because it
bills tokens every turn). It needs uv for the MCP
server, and it stays out of the way if you already run a standalone install — the
hooks yield to ~/.local/bin/gaius-* wrappers rather than injecting the corpus twice
(GAIUS_PLUGIN_HOOKS=force overrides).
Token budgets are deliberately conservative (2000 corpus + 1500 skills per session);
raise with GAIUS_INJECT_BUDGET / GAIUS_SKILLS_BUDGET.
You still run gaius init once after installing — the plugin wires Claude Code, not
your corpus.
Initialize
# Interactive setup — locates the packaged presets (gaius/presets/) and writes
# ~/.gaius/config.yaml. Works for pip installs; no repo checkout needed.
gaius init
# Non-interactive form:
gaius init --backend k8s --yes # for K8s clusters
# or: gaius init --backend default --yes
# Edit to set your sessions_dir and domain_dir
$EDITOR ~/.gaius/config.yaml
First Run
# Scan local sessions and stage summaries.
# Plain `retire` also auto-sweeps Grok (~/.grok/sessions) and Codex
# (~/.codex/sessions) when those CLI dirs exist; `--format <name>` scopes to one.
gaius retire
# Review staged summaries (optional — extracted facts already inject by default).
# Read each and, if you want, promote highlights into your domain/*.md files.
# Worth the loop for high-stakes domains (trading, prod ops) where a wrong fact is
# costly; skip it for low-stakes notes and let decay + dedup self-correct.
gaius batch # print all unreviewed in sequence
gaius next # print one at a time
gaius done <uuid> # mark a staged summary reviewed (queue hygiene, not an inject gate)
# Check corpus stats
gaius stats
# Inject context at session start
gaius inject --task "debug flannel networking" --budget 4000
MCP Server (Claude Code)
# Register the MCP server in Claude Code
claude mcp add gaius -- python3 -m gaius.mcp_server
# Or if using the script install:
claude mcp add gaius -- /path/to/gaius/gaius/mcp_server.py
7 tools available mid-session: gaius_search, gaius_kg_query, gaius_kg_timeline, gaius_stats, gaius_fact_add, gaius_prime_session, gaius_skill_recommend.
Cross-Session Coordination (concord)
When to use: more than one agent (or human + agent) on the same repo or incident at once — parallel sessions, a multi-responder incident, or a fan-out whose subtasks must not collide. Single-session work doesn't need it.
Run five Claude Code sessions against one repo and they will re-derive the same diagnosis
and overwrite each other's fixes. gaius concord is a local, offline-first coordination
sidecar (~/.gaius/concord.db — one SQLite file, no services) that gives parallel sessions:
- Advisory claims — atomic single-winner leases on shared resources
(
subsystem:storage,node:web-01,incident:IC) with TTL + holder-pid liveness, so a dead session can't squat a lease. Winning a claim retitles the terminal tab (⚑ storage · session-name) — ownership visible at a glance. Near-miss naming (subsystem:dbvssubsystem:db-migration) surfaces as an overlap warning. - Findings — discoveries published to sibling sessions, with an adversarial review
loop (
open → reviewing → confirmed/refuted). - Task pool — an incident commander seeds divided work once; each new session takes the next task atomically.
- Roster — live sessions merged from Claude (
~/.claude/sessions/) + Grok (~/.grok/active_sessions.json) registries, joined with the claims they hold.
gaius concord status # one-screen sitrep
gaius concord claim subsystem:db --note "schema migration"
gaius concord finding add --summary "replica lag is the root cause" --severity major
gaius concord task add "verify backups" --resource svc:backup
gaius concord task take # atomic — one winner per task
Claims are advisory by design: awareness is automated, action is never taken on a
peer's behalf — a sibling's message is an observation, not authorization. Hook wiring
(session-start briefs, per-prompt deltas, warn-on-conflicting-mutation, the kill-switch
pattern) is documented in docs/concord.md. Zero configuration
required; an optional remote bridge (gaius concord sync) federates multiple machines
against any server implementing the same heartbeat/finding contract.
Architecture
gaius/
├── gaius/
│ ├── _core.py # core logic (extraction, search, inject, dedup, decay)
│ ├── concord.py # cross-session coordination (claims / findings / task pool)
│ ├── kg.py # temporal knowledge-graph commands
│ ├── parsers.py # session / domain-file parsers
│ ├── record.py # session recorder (OpenAI-compatible endpoints)
│ ├── telemetry.py # prompt / injection event logging
│ ├── mcp_server.py # MCP server (7 tools)
│ ├── presets/
│ │ ├── k8s.yaml # Entity patterns for Kubernetes clusters
│ │ └── default.yaml # Minimal defaults for any project
│ ├── __init__.py # Public API surface
│ └── __main__.py # python -m gaius
├── benchmarks/
│ ├── bench_inject.py # injection regression check (bundled demo corpus)
│ ├── bench_longmemeval.py # LongMemEval-S external retrieval eval (--matrix)
│ └── bench_retrieval.py # legacy Recall@k/MRR harness (not for public numbers)
├── tests/
│ └── *.py # unit + integration suites (test_core.py, ...)
├── pyproject.toml
└── LICENSE # Apache 2.0
Storage: single ~/.gaius/facts.db SQLite file — facts table with BM25 virtual table + sqlite-vec embedding index. No external services required.
Embed daemon: optional systemd user service (gaius-embed-daemon) keeps all-MiniLM-L6-v2 loaded in memory (~8ms/query warm vs ~7s cold).
Configuration
Config file: ~/.gaius/config.yaml
# Sessions directory (Claude Code project JSONLs)
sessions_dir: ~/.claude/projects
# Domain memory directory (your *.md knowledge files)
domain_dir: ~/my-memory/domain
# Skills directory (prospective how-to guides)
skills_dir: ~/my-memory/skills
# Principal mapping — agents grouped for cross-agent scoring
principals:
default: operator
# Entity extraction — extend the built-in K8s baseline
entities:
preset: k8s # use built-in K8s patterns; set to "none" to disable
patterns:
service: '\b(?:my-api|my-worker|my-scheduler)\b'
namespace: '\b(?:prod|staging|dev)\b'
See gaius/presets/k8s.yaml for a full annotated example (installed with the package; gaius init copies it for you).
Commands
| Command | Description |
|---|---|
gaius retire |
Scan local sessions → stage new summaries (auto-sweeps Claude/Grok/Codex; --format to scope) |
gaius record |
Capture chat sessions into gaius JSONL (vLLM, any OpenAI-compatible endpoint) |
gaius s3-retire <agent> |
Retire from S3-archived agent sessions (rclone) |
gaius harvest |
Scan cold Gemini CLI sessions (.json format) |
gaius grok-retire |
Scan Grok CLI sessions (~/.grok/sessions/) → stage decision events |
gaius codex-retire |
Scan Codex CLI rollouts (~/.codex/sessions/) → stage decision events |
gaius next |
Print oldest unreviewed summary |
gaius batch |
Print all unreviewed summaries in sequence |
gaius done <uuid> |
Mark summary as reviewed |
gaius confirm / reject / defer <fact-id> |
Review-loop verdict on a pending fact (confirm → human/confidence=1.0; reject → excluded from inject; defer → re-surface in 7 days) |
gaius agent-review <fact-id> |
Mark a pending fact machine-reviewed — queue hygiene only; weighted ≤ auto, never boosts inject rank |
gaius rescan <uuid> |
Force re-extraction of a specific staged session |
gaius show |
List all staged summaries |
gaius stats |
Extraction and corpus statistics |
gaius inject |
Inject ranked corpus + skills into active session |
gaius kg query <entity> |
Query knowledge graph for an entity |
gaius kg timeline <entity> |
Show temporal changes for an entity |
gaius governor |
Cross-agent knowledge gap analysis |
gaius embed |
Build/rebuild semantic embedding index |
gaius index |
Rebuild memory index |
gaius landscape |
Show memory system landscape |
gaius skills |
List available skills with scores |
gaius concord <sub> |
Cross-session coordination — claims / findings / task pool (see below) |
gaius recent-roll |
Evict aged, done, pointered ## Recent State bullets from MEMORY.md into a non-injected archive changelog |
gaius reconcile |
Promote curated repo-doc facts into the corpus (flagged-unverified, insert-once) + dev↔mirror divergence sentinel |
gaius drift |
Check canonical cluster facts for cross-agent drift against a registry |
gaius decay |
Apply time-based score decay to all facts |
gaius completion <shell> |
Emit a shell completion script (bash/zsh/fish) for command names + global flags |
This table is a highlights subset — run
gaius --helpfor the full command list.
Documentation
| Doc | Covers |
|---|---|
| docs/getting-started.md | Install → init → first retire/inject walkthrough |
| docs/hard-gates.md | Hard enforcement gates — the exit:2 action blocks |
| docs/inject.md | Injection ranking, token budgets, session priming |
| docs/review-lifecycle.md | Fact review + correction loop (confirm / reject / defer, decay) |
| docs/kg.md | Temporal knowledge graph — entities, triples, timelines |
| docs/concord.md | Cross-session coordination — claims, findings, task pool, hook wiring |
| docs/session-jsonl-schema.md | Session JSONL format reference for parser authors |
Dependencies
Core (no extras): pyyaml>=6.0 — pure Python, no binary deps.
Semantic search (pip install "gaius-memory[semantic] @ git+https://github.com/jkubo/gaius"):
sentence-transformers>=2.7— local embedding model (all-MiniLM-L6-v2, 384-dim)sqlite-vec>=0.1— vector search extension for SQLite
MCP server (pip install "gaius-memory[mcp] @ git+https://github.com/jkubo/gaius"):
mcp[server]>=1.0
Without [semantic], gaius falls back to keyword-only BM25 search (no embeddings required).
License
Apache 2.0 — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file gaius_memory-0.2.0.tar.gz.
File metadata
- Download URL: gaius_memory-0.2.0.tar.gz
- Upload date:
- Size: 388.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
612cfdacdb49ec61e05c9ec67dcf07f1c4d184ef893379d432bda681ad5df42c
|
|
| MD5 |
1ed91b334a44d9988999af2bfd8a36e8
|
|
| BLAKE2b-256 |
f1472ff828969e79582dd5c233afaa218cbb868ef48260b42df3291ce7a72cb5
|
Provenance
The following attestation bundles were made for gaius_memory-0.2.0.tar.gz:
Publisher:
publish.yml on jkubo/gaius
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
gaius_memory-0.2.0.tar.gz -
Subject digest:
612cfdacdb49ec61e05c9ec67dcf07f1c4d184ef893379d432bda681ad5df42c - Sigstore transparency entry: 2582690206
- Sigstore integration time:
-
Permalink:
jkubo/gaius@3efcec915e31a6fae13dfc4621818c5eb8e93e2e -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/jkubo
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@3efcec915e31a6fae13dfc4621818c5eb8e93e2e -
Trigger Event:
release
-
Statement type:
File details
Details for the file gaius_memory-0.2.0-py3-none-any.whl.
File metadata
- Download URL: gaius_memory-0.2.0-py3-none-any.whl
- Upload date:
- Size: 281.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
242e484e723d56d06d2b38b15495b0ca07ab448c849a361a8a7b04736cbb89f5
|
|
| MD5 |
17fcb3fdeb07549d7e4523500efeb8b6
|
|
| BLAKE2b-256 |
d4eba25ff6f6bee3e2b9ddfeda743aefbdff9aca14bf442b8914c49ffc85d02f
|
Provenance
The following attestation bundles were made for gaius_memory-0.2.0-py3-none-any.whl:
Publisher:
publish.yml on jkubo/gaius
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
gaius_memory-0.2.0-py3-none-any.whl -
Subject digest:
242e484e723d56d06d2b38b15495b0ca07ab448c849a361a8a7b04736cbb89f5 - Sigstore transparency entry: 2582690217
- Sigstore integration time:
-
Permalink:
jkubo/gaius@3efcec915e31a6fae13dfc4621818c5eb8e93e2e -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/jkubo
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@3efcec915e31a6fae13dfc4621818c5eb8e93e2e -
Trigger Event:
release
-
Statement type: