Skip to main content

Ops memory lifecycle manager for AI coding agents

Project description

gaius

Ops memory lifecycle manager for AI coding agents.

License GitHub

Not another RAG chatbot memory — a production-grade system that extracts facts from Claude Code, Gemini CLI, Grok, and Codex sessions, ranks them into an inject-ready corpus, enforces behavioral gates, and prevents you from breaking prod at 3am.

Why gaius — three things most agent-memory tools skip:

  • Runs unattended — extract → promote → inject with no human in the hot path; correction is optional.
  • Prevents actions, not just recalls them — hard gates exit:2 on force-push, unconfirmed live-trade, prod-delete.
  • Fully offline — BM25 + sqlite-vec in one SQLite file. No API keys, no cloud.

Install from git — the gaius-memory name on PyPI is not published yet. Do not pip install gaius-memory from PyPI.

# CLI (offline core). Extras: [semantic], [mcp], [http]
pip install "gaius-memory @ git+https://github.com/jkubo/gaius"

# Or with uv / Claude Code MCP (see .mcp.json):
# uvx --from "gaius-memory[mcp] @ git+https://github.com/jkubo/gaius" gaius-mcp
gaius retire      # scan sessions → stage summaries
gaius batch       # (optional) review + correct — facts inject by default
gaius inject      # inject context into active session

What It Does

Engineers running Claude Code all day generate enormous amounts of institutional knowledge that vanishes when the context window closes. gaius captures it:

  1. Extract — scans Claude Code (and Gemini CLI) session JSONLs, extracts compact summaries with typed signals (knowledge, patterns, errors)
  2. Review (optional) — extracted facts are promoted and inject-eligible by default; review is a correction loop, not a gate before entry. reject drops a bad fact, defer punts it, agent-review clears it from the queue without touching its rank, confirm pins a verified one — while decay, dedup, and mnemosyne health checks keep quality up without a human bottleneck.
  3. Index — promotes extracted facts into a hybrid keyword + semantic SQLite database (facts.db)
  4. Inject — at task start, retrieves relevant facts and loads skills context into the session
  5. Coordinate — cross-session claims, shared findings, and a claimable task pool (gaius concord) so parallel sessions on one machine divide work instead of colliding

Benchmark

The demo below is a regression / smoke check, NOT a quality score: it confirms retrieval still works on a fresh clone. A perfect score on a hand-built 27-fact corpus (same author wrote the facts and the queries) proves the pipeline runs, nothing more. Real quality is the External evaluation section below, on a benchmark gaius didn't build.

bench_inject.py builds a throwaway SQLite corpus from benchmarks/demo_corpus/:

$ python3 benchmarks/bench_inject.py

gaius Injection Benchmark (25 queries, bundled demo corpus)
════════════════════════════════════════════════════════════
Recall:    25/25 (100%)  regression check, NOT a quality score
Prec-warn: 3/25
Corpus:    27 demo facts  (throwaway SQLite, no daemon, ~0.08s/query)

Semantic mode (embed daemon, real corpus):
Cold start:  ~7s    (first query, model load)
Daemon warm: ~8ms   (Unix socket, model kept in memory)

Storage: sqlite-vec (384-dim, all-MiniLM-L6-v2) + BM25 in a single SQLite file
No API keys, no cloud, runs entirely offline.

External evaluation (LongMemEval-S)

The demo above only proves the pipeline runs. For an unbiased measure, gaius's retrieval is scored on LongMemEval-S (ICLR 2025): a third-party benchmark whose 500 questions, ~48-session haystacks, and gold labels gaius has no hand in choosing (independent by construction, not a self-graded demo). Reproducible: python3 benchmarks/bench_longmemeval.py --matrix.

Two metrics, not the same: R@k = per-question fractional recall (stricter; 65% of questions are multi-gold); hit@k = recall_any@k ("at least one gold session in top-k"), the metric published baselines report, so compare hit@k to those, not R@k.

gaius retrieval (all-MiniLM-L6-v2), 500 questions:

config MRR R@5 R@10 hit@10
session-semantic 0.827 85.7 92.7 96.8
turn-semantic 0.906 92.8 96.6 98.8
turn-hybrid (real inject path) 0.912 92.8 95.2 98.6

On the matched metric (hit@k), gaius at its real regime (short-fact units + the hybrid inject ranking) scores hit@10 98.6%, at par with published same-model (all-MiniLM-L6-v2) session-level retrieval baselines (~96-99% recall_any@10; protocols differ; retrieval-only, not an end-to-end QA metric). Larger embedding models beat all-MiniLM outright, a deliberate tradeoff for gaius's 384-dim, no-GPU, fully-offline footprint. Weakest category: temporal-reasoning (~94% R@10 at turn-semantic).

This measures the retrieval engine. Whether proactive injection improves agent outcomes is a separate evaluation (in progress), and is not claimed here.

Chunked-embedding validation: feeding gaius a whole long session and letting it chunk internally recovers the recall naive whole-session embedding loses to the 256-token cap: chunk-granularity hit@10 98.4% vs naive whole-session 95.2% on a matched 250-question subset, at the turn-level ceiling.


What's Genuinely Novel

Feature Description
Review & correction loop retire → stage → promote runs unattended and facts inject by default; reject/defer/confirm plus decay, dedup, and mnemosyne health let you correct the corpus without a human in the hot path.
Multi-agent corroboration Facts confirmed by multiple AI agents (Claude + Gemini) get a 1.5× score boost. Cross-model verification for higher confidence.
Session-type behavioral priming Load different skill sets based on what you're doing: ops, trading, security, code review.
Hard enforcement gates Memory that prevents actions. exit:2 blocks force-push, live trading without confirmation, critical resource deletion.
Mnemosyne health monitoring Automated memory bloat prevention. Line-count thresholds, misclassification audit, split/prune proposals.
Live state injection Domain files with kubectl/curl commands in frontmatter, TTL-cached. Memory that knows what's happening right now.
Hybrid sqlite-vec search Keyword TF-IDF/BM25 + semantic embeddings in a single SQLite file.
Chunked embedding Long facts are split into <=256-token chunks, each embedded; retrieval max-pools over a fact's chunks, so long content isn't silently truncated to its first ~256 tokens.
Temporal knowledge graph Entity-relationship triples with validity windows. gaius kg timeline node shows what changed and when.
Obsidian-vault viewer The knowledge graph exports [[wikilink]] "Related" blocks into your memory files (kg export-links), so the corpus opens directly as an Obsidian vault — browse entities and their links visually, no extra tooling.
Cross-session coordination gaius concord — advisory claims with TTL + pid-liveness, shared findings with an adversarial review loop, a claimable task pool, and a live roster. Parallel sessions get ownership: each can go deep on its lane instead of defensively re-checking the whole world.
MCP server 7 tools for mid-session memory access without leaving Claude Code.

Quick Start

Install

# Option 1: Claude Code plugin (recommended — wires the hooks, skill, and MCP server for you)
/plugin marketplace add jkubo/claude-plugins
/plugin install gaius@jkubo-tools
/reload-plugins

# Option 2: development install (editable, reads from repo)
git clone https://github.com/jkubo/gaius
cd gaius
pip install -e ".[semantic]"   # includes sentence-transformers + sqlite-vec

# Option 3: script install (no pip, uses system/venv python)
install -m 755 gaius_cli ~/.local/bin/gaius

The plugin bundles the gaius skill, the MCP server, and three hooks: corpus injection at session start, a mnemosyne health check after memory-file edits, and an optional per-prompt injection (GAIUS_PROMPT_INJECT=1, off by default because it bills tokens every turn). It needs uv for the MCP server, and it stays out of the way if you already run a standalone install — the hooks yield to ~/.local/bin/gaius-* wrappers rather than injecting the corpus twice (GAIUS_PLUGIN_HOOKS=force overrides).

Token budgets are deliberately conservative (2000 corpus + 1500 skills per session); raise with GAIUS_INJECT_BUDGET / GAIUS_SKILLS_BUDGET.

You still run gaius init once after installing — the plugin wires Claude Code, not your corpus.

Initialize

# Create config dir
mkdir -p ~/.gaius

# Copy example config
cp presets/k8s.yaml ~/.gaius/config.yaml   # for K8s clusters
# or: cp presets/default.yaml ~/.gaius/config.yaml

# Edit to set your sessions_dir and domain_dir
$EDITOR ~/.gaius/config.yaml

First Run

# Scan local sessions and stage summaries.
# Plain `retire` also auto-sweeps Grok (~/.grok/sessions) and Codex
# (~/.codex/sessions) when those CLI dirs exist; `--format <name>` scopes to one.
gaius retire

# Review staged summaries (optional — extracted facts already inject by default).
# Read each and, if you want, promote highlights into your domain/*.md files.
# Worth the loop for high-stakes domains (trading, prod ops) where a wrong fact is
# costly; skip it for low-stakes notes and let decay + dedup self-correct.
gaius batch          # print all unreviewed in sequence
gaius next           # print one at a time
gaius done <uuid>    # mark a staged summary reviewed (queue hygiene, not an inject gate)

# Check corpus stats
gaius stats

# Inject context at session start
gaius inject --task "debug flannel networking" --budget 4000

MCP Server (Claude Code)

# Register the MCP server in Claude Code
claude mcp add gaius -- python3 -m gaius.mcp_server

# Or if using the script install:
claude mcp add gaius -- /path/to/gaius/gaius/mcp_server.py

7 tools available mid-session: gaius_search, gaius_kg_query, gaius_kg_timeline, gaius_stats, gaius_fact_add, gaius_prime_session, gaius_skill_recommend.


Cross-Session Coordination (concord)

When to use: more than one agent (or human + agent) on the same repo or incident at once — parallel sessions, a multi-responder incident, or a fan-out whose subtasks must not collide. Single-session work doesn't need it.

Run five Claude Code sessions against one repo and they will re-derive the same diagnosis and overwrite each other's fixes. gaius concord is a local, offline-first coordination sidecar (~/.gaius/concord.db — one SQLite file, no services) that gives parallel sessions:

  • Advisory claims — atomic single-winner leases on shared resources (subsystem:storage, node:web-01, incident:IC) with TTL + holder-pid liveness, so a dead session can't squat a lease. Winning a claim retitles the terminal tab (⚑ storage · session-name) — ownership visible at a glance. Near-miss naming (subsystem:db vs subsystem:db-migration) surfaces as an overlap warning.
  • Findings — discoveries published to sibling sessions, with an adversarial review loop (open → reviewing → confirmed/refuted).
  • Task pool — an incident commander seeds divided work once; each new session takes the next task atomically.
  • Roster — live sessions merged from Claude (~/.claude/sessions/) + Grok (~/.grok/active_sessions.json) registries, joined with the claims they hold.
gaius concord status                                   # one-screen sitrep
gaius concord claim subsystem:db --note "schema migration"
gaius concord finding add --summary "replica lag is the root cause" --severity major
gaius concord task add "verify backups" --resource svc:backup
gaius concord task take                                # atomic — one winner per task

Claims are advisory by design: awareness is automated, action is never taken on a peer's behalf — a sibling's message is an observation, not authorization. Hook wiring (session-start briefs, per-prompt deltas, warn-on-conflicting-mutation, the kill-switch pattern) is documented in docs/concord.md. Zero configuration required; an optional remote bridge (gaius concord sync) federates multiple machines against any server implementing the same heartbeat/finding contract.


Architecture

gaius/
├── gaius/
│   ├── _core.py          # core logic (extraction, search, inject, dedup, decay)
│   ├── concord.py        # cross-session coordination (claims / findings / task pool)
│   ├── kg.py             # temporal knowledge-graph commands
│   ├── parsers.py        # session / domain-file parsers
│   ├── record.py         # session recorder (OpenAI-compatible endpoints)
│   ├── telemetry.py      # prompt / injection event logging
│   ├── mcp_server.py     # MCP server (7 tools)
│   ├── __init__.py       # Public API surface
│   └── __main__.py       # python -m gaius
├── http_adapter.py       # FastAPI REST adapter (optional, for remote access)
├── presets/
│   ├── k8s.yaml          # Entity patterns for Kubernetes clusters
│   └── default.yaml      # Minimal defaults for any project
├── benchmarks/
│   ├── bench_inject.py        # injection regression check (bundled demo corpus)
│   ├── bench_longmemeval.py   # LongMemEval-S external retrieval eval (--matrix)
│   └── bench_retrieval.py     # legacy Recall@k/MRR harness (not for public numbers)
├── tests/
│   └── *.py              # unit + integration suites (test_core.py, ...)
├── pyproject.toml
└── LICENSE               # Apache 2.0

Storage: single ~/.gaius/facts.db SQLite file — facts table with BM25 virtual table + sqlite-vec embedding index. No external services required.

Embed daemon: optional systemd user service (gaius-embed-daemon) keeps all-MiniLM-L6-v2 loaded in memory (~8ms/query warm vs ~7s cold).


Configuration

Config file: ~/.gaius/config.yaml

# Sessions directory (Claude Code project JSONLs)
sessions_dir: ~/.claude/projects

# Domain memory directory (your *.md knowledge files)
domain_dir: ~/my-memory/domain

# Skills directory (prospective how-to guides)
skills_dir: ~/my-memory/skills

# Principal mapping — agents grouped for cross-agent scoring
principals:
  default: operator

# Entity extraction — extend the built-in K8s baseline
entities:
  preset: k8s        # use built-in K8s patterns; set to "none" to disable
  patterns:
    service: '\b(?:my-api|my-worker|my-scheduler)\b'
    namespace: '\b(?:prod|staging|dev)\b'

See presets/k8s.yaml for a full annotated example.


Commands

Command Description
gaius retire Scan local sessions → stage new summaries (auto-sweeps Claude/Grok/Codex; --format to scope)
gaius record Capture chat sessions into gaius JSONL (vLLM, any OpenAI-compatible endpoint)
gaius s3-retire <agent> Retire from S3-archived agent sessions (rclone)
gaius harvest Scan cold Gemini CLI sessions (.json format)
gaius grok-retire Scan Grok CLI sessions (~/.grok/sessions/) → stage decision events
gaius codex-retire Scan Codex CLI rollouts (~/.codex/sessions/) → stage decision events
gaius next Print oldest unreviewed summary
gaius batch Print all unreviewed summaries in sequence
gaius done <uuid> Mark summary as reviewed
gaius confirm / reject / defer <fact-id> Review-loop verdict on a pending fact (confirm → human/confidence=1.0; reject → excluded from inject; defer → re-surface in 7 days)
gaius agent-review <fact-id> Mark a pending fact machine-reviewed — queue hygiene only; weighted ≤ auto, never boosts inject rank
gaius rescan <uuid> Force re-extraction of a specific staged session
gaius show List all staged summaries
gaius stats Extraction and corpus statistics
gaius inject Inject ranked corpus + skills into active session
gaius kg query <entity> Query knowledge graph for an entity
gaius kg timeline <entity> Show temporal changes for an entity
gaius governor Cross-agent knowledge gap analysis
gaius embed Build/rebuild semantic embedding index
gaius index Rebuild memory index
gaius landscape Show memory system landscape
gaius skills List available skills with scores
gaius concord <sub> Cross-session coordination — claims / findings / task pool (see below)
gaius recent-roll Evict aged, done, pointered ## Recent State bullets from MEMORY.md into a non-injected archive changelog
gaius reconcile Promote curated repo-doc facts into the corpus (flagged-unverified, insert-once) + dev↔mirror divergence sentinel
gaius drift Check canonical cluster facts for cross-agent drift against a registry
gaius decay Apply time-based score decay to all facts
gaius completion <shell> Emit a shell completion script (bash/zsh/fish) for command names + global flags

This table is a highlights subset — run gaius --help for the full command list.


Dependencies

Core (no extras): pyyaml>=6.0 — pure Python, no binary deps.

Semantic search (pip install "gaius-memory[semantic] @ git+https://github.com/jkubo/gaius"):

  • sentence-transformers>=2.7 — local embedding model (all-MiniLM-L6-v2, 384-dim)
  • sqlite-vec>=0.1 — vector search extension for SQLite

MCP server (pip install "gaius-memory[mcp] @ git+https://github.com/jkubo/gaius"):

  • mcp[server]>=1.0

HTTP adapter (pip install "gaius-memory[http] @ git+https://github.com/jkubo/gaius"):

  • fastapi>=0.100, uvicorn[standard]>=0.23

Without [semantic], gaius falls back to keyword-only BM25 search (no embeddings required).


License

Apache 2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gaius_memory-0.1.0.tar.gz (268.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gaius_memory-0.1.0-py3-none-any.whl (213.8 kB view details)

Uploaded Python 3

File details

Details for the file gaius_memory-0.1.0.tar.gz.

File metadata

  • Download URL: gaius_memory-0.1.0.tar.gz
  • Upload date:
  • Size: 268.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.13

File hashes

Hashes for gaius_memory-0.1.0.tar.gz
Algorithm Hash digest
SHA256 f6acf03dfc18adeab3dcabf537631b2cf86848b3554a0863c6b65be84787a3c6
MD5 8107c989051057a35c28fce1cd8d71f8
BLAKE2b-256 626bc10cc756624779147055d5077258b3a1b85ecdc928f5c9dd4e614a49db76

See more details on using hashes here.

File details

Details for the file gaius_memory-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: gaius_memory-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 213.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.13

File hashes

Hashes for gaius_memory-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b39b5ac03012ba4b90d26ebcd70f4cf102d91f232b15a072cd427a5886188ea2
MD5 c3a134f22f7083d36b27bdd295939393
BLAKE2b-256 b608e48ecd1c3b9f08f920b69233e5a6acf1894cd0635fdb851f10daa522f71c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page