Skip to main content

gaius

Ops memory lifecycle manager for AI coding agents.

License PyPI GitHub

Not another RAG chatbot memory — a production-grade system that extracts facts from Claude Code, Gemini CLI, Grok, and Codex sessions, ranks them into an inject-ready corpus, enforces behavioral gates, and prevents you from breaking prod at 3am.

Why gaius — three things most agent-memory tools skip:

  • Runs unattended — extract → promote → inject with no human in the hot path; correction is optional.
  • Prevents actions, not just recalls them — hard gates exit:2 on force-push, unconfirmed live-trade, prod-delete.
  • Fully offline — BM25 + sqlite-vec in one SQLite file. No API keys, no cloud.
# CLI (offline core). Extras: [semantic], [mcp]
pip install gaius-memory

# Or with uv / Claude Code MCP (see .mcp.json):
# uvx --from "gaius-memory[mcp]" gaius-mcp
#
# Tip: for unreleased mainline, pin git:
# pip install "gaius-memory @ git+https://github.com/jkubo/gaius"

⚠️ Install name is gaius-memory, not bare gaius. PyPI already has an unrelated package named gaius (ImmobilienScout24 deploy client v145+). That package owns neither our import path nor our CLI forever, but pip install gaius will not install this project. Always: pip install gaius-memory → import gaius, console script gaius.

gaius retire      # scan sessions → stage summaries
gaius batch       # (optional) review + correct — facts inject by default
gaius quiz        # Leitner HITL loop — the only legitimate `confidence_source='human'`
gaius inject --task "what you're working on"   # inject context into active session

What It Does

Engineers running Claude Code all day generate enormous amounts of institutional knowledge that vanishes when the context window closes. gaius captures it:

  1. Extract — scans Claude Code (and Gemini CLI) session JSONLs, extracts compact summaries with typed signals (knowledge, patterns, errors)
  2. Review (optional) — extracted facts are promoted and inject-eligible by default; review is a correction loop, not a gate before entry. reject drops a bad fact, defer punts it, agent-review clears it from the queue without touching its rank, confirm pins a verified one — while decay, dedup, and mnemosyne health checks keep quality up without a human bottleneck.
  3. Index — promotes extracted facts into a hybrid keyword + semantic SQLite database (facts.db)
  4. Inject — at task start, retrieves relevant facts and loads skills context into the session
  5. Coordinate — cross-session claims, shared findings, and a claimable task pool (gaius concord) so parallel sessions on one machine divide work instead of colliding

Benchmark

The demo below is a regression / smoke check, NOT a quality score: it confirms retrieval still works on a fresh clone. A perfect score on a hand-built 27-fact corpus (same author wrote the facts and the queries) proves the pipeline runs, nothing more. Real quality is the External evaluation section below, on a benchmark gaius didn't build.

bench_inject.py builds a throwaway SQLite corpus from benchmarks/demo_corpus/:

$ python3 benchmarks/bench_inject.py

gaius Injection Benchmark (25 queries, bundled demo corpus)
════════════════════════════════════════════════════════════
Recall:    25/25 (100%)  regression check, NOT a quality score
Prec-warn: 3/25
Corpus:    27 demo facts  (throwaway SQLite, no daemon, ~0.08s/query)

Semantic mode (embed daemon, real corpus):
Cold start:  ~7s    (first query, model load)
Daemon warm: ~8ms   (Unix socket, model kept in memory)

Storage: sqlite-vec (384-dim, all-MiniLM-L6-v2) + BM25 in a single SQLite file
No API keys, no cloud, runs entirely offline.

External evaluation (LongMemEval-S)

The demo above only proves the pipeline runs. For an unbiased measure, gaius's retrieval is scored on LongMemEval-S (ICLR 2025): a third-party benchmark whose 500 questions, ~48-session haystacks, and gold labels gaius has no hand in choosing (independent by construction, not a self-graded demo). Reproducible: python3 benchmarks/bench_longmemeval.py --matrix.

Two metrics, not the same: R@k = per-question fractional recall (stricter; 65% of questions are multi-gold); hit@k = recall_any@k ("at least one gold session in top-k"), the metric published baselines report, so compare hit@k to those, not R@k.

gaius retrieval (all-MiniLM-L6-v2), 500 questions:

config MRR R@5 R@10 hit@10
session-semantic 0.827 85.7 92.7 96.8
turn-semantic 0.906 92.8 96.6 98.8
turn-hybrid (real inject path) 0.912 92.8 95.2 98.6

On the matched metric (hit@k), gaius at its real regime (short-fact units + the hybrid inject ranking) scores hit@10 98.6%, at par with published same-model (all-MiniLM-L6-v2) session-level retrieval baselines (~96-99% recall_any@10; protocols differ; retrieval-only, not an end-to-end QA metric). Larger embedding models beat all-MiniLM outright, a deliberate tradeoff for gaius's 384-dim, no-GPU, fully-offline footprint. Weakest category: temporal-reasoning (~94% R@10 at turn-semantic).

This measures the retrieval engine. Whether proactive injection improves agent outcomes is a separate evaluation (in progress), and is not claimed here.

Chunked-embedding validation: feeding gaius a whole long session and letting it chunk internally recovers the recall naive whole-session embedding loses to the 256-token cap: chunk-granularity hit@10 98.4% vs naive whole-session 95.2% on a matched 250-question subset, at the turn-level ceiling.


What's Genuinely Novel

Feature Description
Review & correction loop retire → stage → promote runs unattended and facts inject by default; reject/defer/confirm plus decay, dedup, and mnemosyne health let you correct the corpus without a human in the hot path.
Leitner human review (gaius quiz) Spaced-repetition boxes over corpus facts. The only legitimate producer of confidence_source='human'; --report is the calibration view (disagreement-rate by domain). Agents are refused on the mutating path.
Multi-agent corroboration Facts confirmed by multiple AI agents (Claude + Gemini) get a 1.5× score boost. Cross-model verification for higher confidence.
Session-type behavioral priming Load different skill sets based on what you're doing: ops, trading, security, code review.
Hard enforcement gates Memory that prevents actions. exit:2 blocks force-push, live trading without confirmation, critical resource deletion.
Mnemosyne health monitoring Automated memory bloat prevention. Line-count thresholds, misclassification audit, split/prune proposals.
Live state injection Domain files with kubectl/curl commands in frontmatter, TTL-cached. Memory that knows what's happening right now.
Hybrid sqlite-vec search Keyword TF-IDF/BM25 + semantic embeddings in a single SQLite file.
Chunked embedding Long facts are split into <=256-token chunks, each embedded; retrieval max-pools over a fact's chunks, so long content isn't silently truncated to its first ~256 tokens.
Temporal knowledge graph Entity-relationship triples with validity windows. gaius kg timeline node shows what changed and when.
Obsidian-vault viewer The knowledge graph exports [[wikilink]] "Related" blocks into your memory files (kg export-links), so the corpus opens directly as an Obsidian vault — browse entities and their links visually, no extra tooling.
Cross-session coordination gaius concord — advisory claims with TTL + pid-liveness, shared findings with an adversarial review loop, a claimable task pool, and a live roster. Parallel sessions get ownership: each can go deep on its lane instead of defensively re-checking the whole world.
MCP server 7 tools for mid-session memory access without leaving Claude Code.

Quick Start

Install

# Option 1: Claude Code plugin (recommended — wires the hooks, skill, and MCP server for you)
/plugin marketplace add jkubo/claude-plugins
/plugin install gaius@jkubo-tools
/reload-plugins

# Option 2: development install (editable, reads from repo)
git clone https://github.com/jkubo/gaius
cd gaius
pip install -e ".[semantic]"   # includes sentence-transformers + sqlite-vec

# Option 3: script install (no pip, uses system/venv python)
install -m 755 gaius_cli ~/.local/bin/gaius

The plugin bundles the gaius skill, the MCP server, and three hooks: corpus injection at session start, a mnemosyne health check after memory-file edits, and an optional per-prompt injection (GAIUS_PROMPT_INJECT=1, off by default because it bills tokens every turn). It needs uv for the MCP server, and it stays out of the way if you already run a standalone install — the hooks yield to ~/.local/bin/gaius-* wrappers rather than injecting the corpus twice (GAIUS_PLUGIN_HOOKS=force overrides).

Token budgets are deliberately conservative (2000 corpus + 1500 skills per session); raise with GAIUS_INJECT_BUDGET / GAIUS_SKILLS_BUDGET.

You still run gaius init once after installing — the plugin wires Claude Code, not your corpus.

Initialize

# Interactive setup — locates the packaged presets (gaius/presets/) and writes
# ~/.gaius/config.yaml. Works for pip installs; no repo checkout needed.
gaius init

# Non-interactive form:
gaius init --backend k8s --yes       # for K8s clusters
# or: gaius init --backend default --yes

# Edit to set your sessions_dir and domain_dir
$EDITOR ~/.gaius/config.yaml

First Run

# Scan local sessions and stage summaries.
# Plain `retire` also auto-sweeps Grok (~/.grok/sessions) and Codex
# (~/.codex/sessions) when those CLI dirs exist; `--format <name>` scopes to one.
gaius retire

# Review staged summaries (optional — extracted facts already inject by default).
# Read each and, if you want, promote highlights into your domain/*.md files.
# Worth the loop for high-stakes domains (trading, prod ops) where a wrong fact is
# costly; skip it for low-stakes notes and let decay + dedup self-correct.
gaius batch          # print all unreviewed in sequence
gaius next           # print one at a time
gaius done <uuid>    # mark a staged summary reviewed (queue hygiene, not an inject gate)

# Check corpus stats
gaius stats

# Inject context at session start
gaius inject --task "debug flannel networking" --budget 4000

MCP Server (Claude Code)

# Register the MCP server in Claude Code
claude mcp add gaius -- python3 -m gaius.mcp_server

# Or if using the script install:
claude mcp add gaius -- /path/to/gaius/gaius/mcp_server.py

7 tools available mid-session: gaius_search, gaius_kg_query, gaius_kg_timeline, gaius_stats, gaius_fact_add, gaius_prime_session, gaius_skill_recommend.


Cross-Session Coordination (concord)

When to use: more than one agent (or human + agent) on the same repo or incident at once — parallel sessions, a multi-responder incident, or a fan-out whose subtasks must not collide. Single-session work doesn't need it.

Run five Claude Code sessions against one repo and they will re-derive the same diagnosis and overwrite each other's fixes. gaius concord is a local, offline-first coordination sidecar (~/.gaius/concord.db — one SQLite file, no services) that gives parallel sessions:

  • Advisory claims — atomic single-winner leases on shared resources (subsystem:storage, node:web-01, incident:IC) with TTL + holder-pid liveness, so a dead session can't squat a lease. Winning a claim retitles the terminal tab (⚑ storage · session-name) — ownership visible at a glance. Near-miss naming (subsystem:db vs subsystem:db-migration) surfaces as an overlap warning.
  • Findings — discoveries published to sibling sessions, with an adversarial review loop (open → reviewing → confirmed/refuted).
  • Task pool — an incident commander seeds divided work once; each new session takes the next task atomically.
  • Roster — live sessions merged from Claude (~/.claude/sessions/) + Grok (~/.grok/active_sessions.json) registries, joined with the claims they hold.
gaius concord status                                   # one-screen sitrep
gaius concord claim subsystem:db --note "schema migration"
gaius concord finding add --summary "replica lag is the root cause" --severity major
gaius concord task add "verify backups" --resource svc:backup
gaius concord task take                                # atomic — one winner per task

Claims are advisory by design: awareness is automated, action is never taken on a peer's behalf — a sibling's message is an observation, not authorization. Hook wiring (session-start briefs, per-prompt deltas, warn-on-conflicting-mutation, the kill-switch pattern) is documented in docs/concord.md. Zero configuration required; an optional remote bridge (gaius concord sync) federates multiple machines against any server implementing the same heartbeat/finding contract.


Architecture

gaius/
├── gaius/
│   ├── _core.py          # core logic (extraction, search, inject, dedup, decay)
│   ├── concord.py        # cross-session coordination (claims / findings / task pool)
│   ├── kg.py             # temporal knowledge-graph commands
│   ├── parsers.py        # session / domain-file parsers
│   ├── record.py         # session recorder (OpenAI-compatible endpoints)
│   ├── telemetry.py      # prompt / injection event logging
│   ├── mcp_server.py     # MCP server (7 tools)
│   ├── presets/
│   │   ├── k8s.yaml      # Entity patterns for Kubernetes clusters
│   │   └── default.yaml  # Minimal defaults for any project
│   ├── __init__.py       # Public API surface
│   └── __main__.py       # python -m gaius
├── benchmarks/
│   ├── bench_inject.py        # injection regression check (bundled demo corpus)
│   ├── bench_longmemeval.py   # LongMemEval-S external retrieval eval (--matrix)
│   └── bench_retrieval.py     # legacy Recall@k/MRR harness (not for public numbers)
├── tests/
│   └── *.py              # unit + integration suites (test_core.py, ...)
├── pyproject.toml
└── LICENSE               # Apache 2.0

Storage: single ~/.gaius/facts.db SQLite file — facts table with BM25 virtual table + sqlite-vec embedding index. No external services required.

Embed daemon: optional systemd user service (gaius-embed-daemon) keeps all-MiniLM-L6-v2 loaded in memory (~8ms/query warm vs ~7s cold).


Configuration

Config file: ~/.gaius/config.yaml

# Sessions directory (Claude Code project JSONLs)
sessions_dir: ~/.claude/projects

# Domain memory directory (your *.md knowledge files)
domain_dir: ~/my-memory/domain

# Skills directory (prospective how-to guides)
skills_dir: ~/my-memory/skills

# Principal mapping — agents grouped for cross-agent scoring
principals:
  default: operator

# Entity extraction — extend the built-in K8s baseline
entities:
  preset: k8s        # use built-in K8s patterns; set to "none" to disable
  patterns:
    service: '\b(?:my-api|my-worker|my-scheduler)\b'
    namespace: '\b(?:prod|staging|dev)\b'

See gaius/presets/k8s.yaml for a full annotated example (installed with the package; gaius init copies it for you).


Commands

Command Description
gaius retire Scan local sessions → stage new summaries (auto-sweeps Claude/Grok/Codex; --format to scope)
gaius record Capture chat sessions into gaius JSONL (vLLM, any OpenAI-compatible endpoint)
gaius s3-retire <agent> Retire from S3-archived agent sessions (rclone)
gaius harvest Scan cold Gemini CLI sessions (.json format)
gaius grok-retire Scan Grok CLI sessions (~/.grok/sessions/) → stage decision events
gaius codex-retire Scan Codex CLI rollouts (~/.codex/sessions/) → stage decision events
gaius next Print oldest unreviewed summary
gaius batch Print all unreviewed summaries in sequence
gaius done <uuid> Mark summary as reviewed
gaius confirm / reject / defer <fact-id> Review-loop verdict on a pending fact (confirm → human/confidence=1.0; reject → excluded from inject; defer → re-surface in 7 days)
gaius agent-review <fact-id> Mark a pending fact machine-reviewed — queue hygiene only; weighted ≤ auto, never boosts inject rank
gaius rescan <uuid> Force re-extraction of a specific staged session
gaius show List all staged summaries
gaius stats Extraction and corpus statistics
gaius inject Inject ranked corpus + skills into active session
gaius kg query <entity> Query knowledge graph for an entity
gaius kg timeline <entity> Show temporal changes for an entity
gaius governor Cross-agent knowledge gap analysis
gaius embed Build/rebuild semantic embedding index
gaius index Rebuild memory index
gaius landscape Show memory system landscape
gaius skills List available skills with scores
gaius concord <sub> Cross-session coordination — claims / findings / task pool (see below)
gaius recent-roll Evict aged, done, pointered ## Recent State bullets from MEMORY.md into a non-injected archive changelog
gaius reconcile Promote curated repo-doc facts into the corpus (flagged-unverified, insert-once) + dev↔mirror divergence sentinel
gaius drift Check canonical cluster facts for cross-agent drift against a registry
gaius decay Apply time-based score decay to all facts
gaius completion <shell> Emit a shell completion script (bash/zsh/fish) for command names + global flags

This table is a highlights subset — run gaius --help for the full command list.


Documentation

Doc Covers
docs/getting-started.md Install → init → first retire/inject walkthrough
docs/hard-gates.md Hard enforcement gates — the exit:2 action blocks
docs/inject.md Injection ranking, token budgets, session priming
docs/review-lifecycle.md Fact review + correction loop (confirm / reject / defer, decay)
docs/kg.md Temporal knowledge graph — entities, triples, timelines
docs/concord.md Cross-session coordination — claims, findings, task pool, hook wiring
docs/session-jsonl-schema.md Session JSONL format reference for parser authors

Dependencies

Core (no extras): pyyaml>=6.0 — pure Python, no binary deps.

Semantic search (pip install "gaius-memory[semantic] @ git+https://github.com/jkubo/gaius"):

  • sentence-transformers>=2.7 — local embedding model (all-MiniLM-L6-v2, 384-dim)
  • sqlite-vec>=0.1 — vector search extension for SQLite

MCP server (pip install "gaius-memory[mcp] @ git+https://github.com/jkubo/gaius"):

  • mcp[server]>=1.0

Without [semantic], gaius falls back to keyword-only BM25 search (no embeddings required).


License

Apache 2.0 — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gaius_memory-0.2.0.tar.gz (388.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gaius_memory-0.2.0-py3-none-any.whl (281.5 kB view details)

Uploaded Python 3

File details

Details for the file gaius_memory-0.2.0.tar.gz.

File metadata

  • Download URL: gaius_memory-0.2.0.tar.gz
  • Upload date:
  • Size: 388.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for gaius_memory-0.2.0.tar.gz
Algorithm Hash digest
SHA256 612cfdacdb49ec61e05c9ec67dcf07f1c4d184ef893379d432bda681ad5df42c
MD5 1ed91b334a44d9988999af2bfd8a36e8
BLAKE2b-256 f1472ff828969e79582dd5c233afaa218cbb868ef48260b42df3291ce7a72cb5

See more details on using hashes here.

Provenance

The following attestation bundles were made for gaius_memory-0.2.0.tar.gz:

Publisher: publish.yml on jkubo/gaius

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gaius_memory-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: gaius_memory-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 281.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for gaius_memory-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 242e484e723d56d06d2b38b15495b0ca07ab448c849a361a8a7b04736cbb89f5
MD5 17fcb3fdeb07549d7e4523500efeb8b6
BLAKE2b-256 d4eba25ff6f6bee3e2b9ddfeda743aefbdff9aca14bf442b8914c49ffc85d02f

See more details on using hashes here.

Provenance

The following attestation bundles were made for gaius_memory-0.2.0-py3-none-any.whl:

Publisher: publish.yml on jkubo/gaius

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page