Skip to main content

Empirica

We Gave AI a Mirror. Now It Measures What It Believes.

Version PyPI Python License

Epistemic infrastructure for AI — measurement, memory, and calibration across sessions.

Empirica tracks what AI knows, gates what it does, and compounds learning across session boundaries. It measures the gap between what AI predicts and what's true — making AI agents measurably more reliable.

Training & Guides | CLI Reference | Architecture

Important: Empirica is an AI measurement framework. It has no cryptocurrency, token, coin, or blockchain component. Any token using the Empirica name (including "$EMPIRICA" on Solana) is unauthorized and not affiliated with this project or Empirica AI GmbH.


The Problem

AI coding agents today have no self-awareness about what they know:

  • Forgets between sessions — same questions, same dead ends, every time
  • Acts before understanding — edits your code without knowing the architecture
  • Can't tell you when it's guessing — no distinction between knowledge and confabulation
  • No audit trail — reasoning evaporates with the context window

What Empirica Does

Capability What You Experience
Measures before acting AI investigates your codebase before touching it. The Sentinel gate blocks edits until understanding is demonstrated
Remembers across sessions Findings, dead-ends, and learnings persist in a 4-layer memory system. Session 3 starts where Session 2 left off
Prevents confident mistakes The CHECK gate uses domain-aware thresholds scaled by criticality — cybersec/high is stricter than default/low
Shows confidence in real-time Live statusline in your terminal: [empirica] ⚡94% ↕70% │ 🎯3 │ POST 🔍92% │ K:95% C:92%
Calibrates against reality Three-vector model: self-assessed, observed (from deterministic checks), and AI-reasoned grounded state with rationale. Domain compliance loops iterate until all checks pass
Tracks your codebase Temporal entity model auto-extracts functions, classes, and imports from every file edit — the AI knows what's alive and what's stale
Works through natural language You describe tasks normally. The AI operates the measurement system automatically
Optional: coordinates with peer AIs Cross-Claude mesh via Cortex (opt-in) — peer AIs propose work, ECO accepts/declines, completion handshakes carry commit SHAs. A persistent listener wakes idle sessions on inbox events. Empirica core works standalone without this — see Cross-AI Mesh below for the ecosystem layer

How You Use It

You talk to your AI normally. Empirica works in the background:

You:      "Fix the authentication bug in the login flow"

Empirica: [AI investigates → logs findings → passes Sentinel gate → implements fix → measures learning]

You see:  ⚡87% ↕70% │ 🎯1 │ POST 🔍85% │ K:88% C:82% │ Δ +K

You direct. The AI measures.

Empirica's CLI has 280 commands spanning investigation, measurement, calibration, and memory — like a cockpit instrument panel. You don't need to learn any of them. The AI reads the instruments, operates the controls, and reports back in natural language. The statusline gives you the flight data at a glance.

For power users, direct CLI access is always available: empirica goals-list, empirica calibration-report, empirica project-search --task "...", and more.

Learn the full workflow: getempirica.com has interactive training, guides, and deep explanations of every concept.


Quick Start

Install + Claude Code (Recommended)

pip install empirica
empirica setup

Then just start working. The hooks, Sentinel, system prompt, statusline, and MCP server are all configured automatically.

empirica setup --harness picks the harness. It resolves from --harness, else $EMPIRICA_HARNESS, else claude-code. An unsupported harness is refused by name and nothing is written — previously it wrote ~/.claude/ regardless and reported success, which on any other harness configured a path that harness never reads.

Codex is refused deliberately: ecodex is self-provisioning (it vendors the plugin into its own binary), so there is nothing here to write, and the refusal points you at that pipeline. setup-claude-code remains as an alias and pins claude-code regardless of environment.

See Claude Code Setup for details — including a "What the hooks inject" section for Claude sessions that want to see the contract (which hook fires when, what it adds to the AI's context, source pointers for every emission) before agreeing to install.

Already have Claude Code configured? Use --force to replace your default Claude Code settings with Empirica's epistemic hooks. Without --force, setup only writes files that don't already exist — so if you've already used Claude Code, the default internals stay in place and Empirica's hooks won't activate.

empirica setup --force

--force replaces hooks in settings.json but only removes Empirica's own hooks — hooks from other plugins (Railway, Superpowers, etc.) are preserved.

Alternative Installation Methods

Homebrew (macOS)
brew tap empiricaai/tap
brew install empirica
empirica setup
Docker
# Security-hardened Alpine image (~276MB, recommended)
docker pull nubaeon/empirica:1.13.7-alpine

# Standard image (Debian slim, ~414MB)
docker pull nubaeon/empirica:1.13.7

# Run
docker run -it -v $(pwd)/.empirica:/data/.empirica nubaeon/empirica:1.13.7 /bin/bash
Manual / Other AI Platforms
pip install empirica
pip install empirica-mcp        # MCP Server (for Cursor, Cline, etc.)
cd your-project && empirica project-init

The CLI works standalone on any platform. The full epistemic workflow (epistemic transactions, Sentinel, calibration) requires loading the system prompt into your AI — the easiest path is empirica setup, which wires the lean prompt into ~/.claude/empirica-system-prompt.md and references it from your ~/.claude/CLAUDE.md. See Claude Code Setup for details.

First Session

empirica onboard   # Interactive walkthrough of the full workflow

Or just start working — with Claude Code hooks active, the AI manages the epistemic workflow automatically.


The Measurement Architecture

Empirica works through nested abstraction layers:

Plan
 └── Transaction 1 (Goal A)
      ├── NOETIC: investigate, search, read → findings, unknowns, dead-ends
      ├── CHECK: Sentinel gate → proceed / investigate more
      ├── PRAXIC: implement, write, commit → goals completed
      └── POSTFLIGHT: measure learning delta → persists to memory
 └── Transaction 2 (Goal B, informed by T1's findings)
      └── ...

Plans decompose into transactions — one per goal or Claude Code task. Each transaction is a noetic-praxic loop: investigate first (noetic), then act (praxic), with the Sentinel gating the transition. Along the way, the AI collects and reads artifacts (findings, unknowns, assumptions, dead-ends, decisions) while using semantic search to surface relevant epistemic patterns and anti-patterns from the project's history. Top artifacts are ranked by confidence and fed into each project's MEMORY.md as a hot cache.

The Epistemic Transaction Cycle

PREFLIGHT ────────► CHECK ────────► POSTFLIGHT
    │                 │                  │
 Baseline         Sentinel           Learning
 Assessment        Gate               Delta
    │                 │                  │
 "What do I      "Am I ready      "What did I
  know now?"      to act?"         learn?"

PREFLIGHT: AI assesses its knowledge state before starting work. CHECK: Sentinel gate validates readiness before allowing code edits. POSTFLIGHT: AI measures what it learned, creating a delta that persists.


Live Statusline

With Claude Code hooks enabled, you see the AI's epistemic state in real-time:

[empirica] ⚡94% ↕70% │ 🎯3 ❓12/5 │ POST 🔍92% │ K:95% C:92% │ Δ +K +C
Signal Meaning
⚡94% Overall epistemic confidence
↕70% Sentinel threshold (know gate) — user-facing only
🎯3 ❓12/5 Open goals (3), unknowns (12 total, 5 blocking)
POST 🔍92% Transaction phase + work state (🔍 investigating / 🔨 acting) with composite score
K:95% C:92% Knowledge and Context vectors (color-coded by gap to threshold)
Δ +K +C Learning delta (POSTFLIGHT only) — which vectors improved

The 13 Epistemic Vectors

These vectors emerged from 600+ real working sessions across multiple AI systems. They measure the dimensions that consistently predict success or failure in complex tasks.

Tier Vector What It Measures
Gate engagement Is the AI actively processing or disengaged?
Foundation know Domain knowledge depth
do Execution capability
context Access to relevant information
Comprehension clarity How clear is the understanding?
coherence Do the pieces fit together?
signal Signal-to-noise in available information
density Information richness
Execution state Current working state
change Rate of progress/change
completion Task completion level
impact Significance of the work
Meta uncertainty Explicit doubt tracking

Deep dive: Epistemic Vectors Explained


How It Works With Claude Code

Empirica doesn't replace or reinvent anything Claude Code already does. Claude Code owns tasks, plans, memory, and projects. Empirica adds the measurement layer on top:

Claude Code Does Empirica Adds
Task management Epistemic goals with measurable completion
Plan mode Investigation phase with Sentinel gating — no edits until understanding is verified
MEMORY.md Auto-curated hot cache ranked by epistemic confidence
Context window 4-layer memory that survives compaction and persists across sessions
Code editing Grounded calibration — was the AI's confidence justified by test results?
Subagent spawning Bounded autonomy with delegated work counting and budget tracking

The result: Claude Code's native capabilities, enhanced with measurement, gating, and calibration feedback that compounds over time.

18 skills ship with the plugin — the transaction discipline, graph gardening, the quality sweep, mesh messaging, and more. They are lazy: a skill does nothing until it loads, so knowing when each fires matters more than knowing what it holds. → Skills reference


Cross-AI Mesh (Optional Ecosystem Layer)

This section describes an optional layer. Empirica core — measurement, calibration, artifacts, goals, project-search, sentinel gating — works fully standalone. The mesh is an opt-in capability for users who run multiple Claude sessions across projects and want them to coordinate as peers. If you only use one AI in one repo, skip this section.

The mesh runs on top of Empirica Cortex (proprietary serving layer) plus an optional browser extension for ECO triage. At a high level:

empirica AI ── proposes work ──► ECO Accept/Decline ──► peer AI wakes + acts
                                                             │
                              completion handshake (commit SHA)
                                                             │
empirica AI ◄────────── outbox/completed event ──────────────┘
Capability What it does
Mesh proposals (two flavors) A noetic flavor is auto-accepted (FYI / question / discussion). Praxic flavors (code change / architecture / investigation) are ECO-gated — they wait for an Accept/Decline decision before the target AI acts
empirica mailbox reply One CLI verb closes the AI-to-AI handshake atomically — single-step completion ack instead of two
Persistent listener service systemd-user / launchd daemon holds a push stream open. Idle sessions wake the moment a peer's proposal is decided, not on next user prompt
Canonical loops (opt-in) Wake-on-event is the standing trigger, so nothing is scheduled by default. Inbox polling (30s adaptive) exists for harnesses that cannot do wake-on-event, and daily housekeeping is a cron loop — both are opt-in, registered with empirica loop register. See Trigger Model

The browser-side ECO surface (Accept/Decline, inbox triage, publish review) lives in the proprietary Empirica Extension. The full API surface for proposals, listener events, and the trust pipeline is documented at getempirica.com.


Mesh + Shared Epistemic Record (1.11.0)

Requires Empirica Cortex (proprietary). Everything in this section — mesh proposals, the persistent listener, the Shared Epistemic Record, and the empirica mesh command cluster — is a Cortex-served layer. It is not available in empirica core on its own; without Cortex, empirica is a single-AI measurement layer.

The cross-AI coordination layer. Practitioners in different practices coordinate not via text-only chat but via epistemic envelopes that carry calibrated state, source-tagged provenance, noetic/praxic intent, and workflow position.

  • Practitioner / practice framing — practices are calibrated epistemic specializations that persist; practitioners (the LLMs) are fungible. See MESH_CONCEPTS.md.
  • Shared Epistemic Record (SER) — cortex-resident shared-state object for coordination across ≥2 practitioners. Goals stay per-practitioner; SER carries the joint state (coordination_state, role-tiered participants, escalate-on-silence). Three actions: create_ser / transition_ser / ser_ack. Spec at empirica-cortex/docs/architecture/SHARED_EPISTEMIC_RECORD.md.
  • empirica mesh command cluster (1.11.0) — unified diagnostic + control surface across listener instances + the optional cortex bridge:
    empirica mesh status              # per-instance health (local + cortex bridge)
    empirica mesh diagnose <ai_id>    # deep diagnostic + suggested fix command
    empirica mesh restart <ai_id>     # systemd/launchd restart + verify
    empirica mesh on|off <ai_id>      # install + start | stop the listener
    empirica mesh tail [<ai_id>]      # live-tail loop_fires.log
    
  • Listener self-heal — in-process watchdog terminates stale curl streams (TCP-zombie detection at 120s by default); HTTP 429 detection applies long backoff with catch-up poll continuing during the window.
  • Mesh Routing Protocol v0 locked four-way with cortex + extension + mesh-support. L1/L2/L3 trust model, server-stamped layer annotation, participant-scoped thread reads.

Without Cortex, empirica is a single-AI measurement layer — the proposals, listener, SER, and empirica mesh cluster above are all Cortex-dependent. Core does ship a minimal local empirica message-* git-notes primitive for passing notes between your own sessions, but that is note-passing, not the coordinating mesh. Everything that makes empirica valuable on its own — measurement, calibration, artifacts, goals, project-search, sentinel gating — works fully standalone.


Practice Model + Entity Graph (1.10.0)

Empirica's workspace stores entities (projects, contacts, organisations, engagements, users) in entity_registry with typed edges in entity_memberships. The Practice Model frames this consistently:

Term Maps to
Practitioner the AI working on the project (you)
Practice the empirica project itself
Agent a subagent spawned during the work

Four CLI verbs query the graph without raw SQL:

empirica entity-list [--type project|contact|organization|engagement|user]
empirica entity-show <type:id>          # full record + incoming/outgoing edges
empirica entity-walk <type:id> --depth 3 # BFS membership graph, cycle-safe
empirica entity-search "query" [--type T]

All read-only, all support --output json. Backs cross-project orchestration, CRM workflows, and the entity-aware POSTFLIGHT retrospective.


Platform Support

Two harnesses are supported. Everything else is untested — prompt and rules files may exist for other tools, but their presence is not support, and we would rather say so than let you find out mid-project.

Harness Status What you get
Claude Code Supported Full integration — plugin, hooks, Sentinel gate, skills, agents, statusline, MCP
Codex (via ecodex) Supported Self-provisioning: ecodex vendors the plugin into its own binary and loads hooks natively. empirica setup refuses codex by name and points you at ecodex's pipeline — that refusal is correct, not a gap
Antigravity (Google), Vibe (Mistral AI) Not yet Possible future support. Not now
Everything else Untested The CLI works anywhere Python does, but no harness integration is verified

MCP is a fallback, not a second path

empirica-mcp exists for harnesses that cannot run the CLI integration — Claude Desktop, Gipitee-style clients. Use it when you have no alternative.

The reason matters: an MCP surface cannot enforce the Sentinel gate the way a blocking pre-tool hook does. So an MCP-only deployment gives you the measurement layer with a weaker guarantee — the noetic/praxic firewall becomes advisory. That is a real difference in what the system promises, and it should be a deliberate choice rather than something you infer from a feature table.


Documentation & Training

Resource What It Covers
getempirica.com Training course, interactive guides, deep explanations
Natural Language Guide How to collaborate with AI using Empirica
Getting Started First-time setup and concepts
CLI Reference All 280 commands documented
Architecture Technical reference for contributors
Claude Code Setup Install + system prompt + plugin wiring
Changelog Full release history — every version since 1.0
Upgrade to 1.11 Migration guide rolling up 1.10.5+1.10.6+1.11.x — bead v0 → SER, mesh substrate hardening, MESH_CONCEPTS framing

The Empirica Ecosystem

Project Description Status
Empirica Core measurement system — epistemic transactions, Sentinel, calibration, 13 vectors Open source (MIT)
Empirica Iris Epistemic browser automation with SVG spatial indexing — Sentinel gating for visual interactions Open source (MIT)
Docpistemic Epistemic documentation coverage assessment — know what your docs know Open source (MIT)
Breadcrumbs Survive context compacts with git notes — dead simple session continuity Open source (MIT)
Ecodex Codex-based Rust harness — the epistemic firewall and pattern hunt, native to Rust/cargo (clippy, cargo check/test/audit) Open source (Apache-2.0)
Eat the Broccoli Portable quality-and-pattern audit — deterministic tooling plus a learned-pattern hunt for the failure classes that pass every test and still ship broken Open source (MIT)
Empirica Cortex Cross-project intelligence layer — serves verified predictions and accumulated learnings to condition future work Proprietary
Empirica Workspace Entity Knowledge Graph, Epistemic Prompt Engine, CRM, portfolio dashboard Proprietary
Empirica Extension Chrome extension — desktop face of the mesh. ECO Accept/Decline, inbox/outbox triage, publish review, conversation extraction from Claude.ai / ChatGPT / Gemini / Grok Proprietary
Empirica Outreach Voice-aware outreach + publishing — prosody-matched content generation and multi-channel dispatch Proprietary

Building something with Empirica? Open an issue to get listed.


The Empirica Foundation

The Empirica Foundation stewards the open ecosystem — the public projects above and the community growing around them — keeping the commons healthy as it scales.

The open-source projects are free for everyone. What the Foundation adds is a seat at the table: contributors who help build the community get to use the rest of the ecosystem too — the otherwise-paid layers (Cortex, Workspace, the Extension) and the collaborative mesh that lets practitioners coordinate as peers. Build the commons, use the whole thing.

Want in? Open an issue or a PR with your reasons — that's the whole application. Everyone who wants to help shape where this goes is welcome.


What's New in 1.13.7

  • §REPORTING in the system prompt: report D, not A → B → C → D. If the work went A→D, the user gets D — A–C is one sentence at most, because everything cut is already in the artifacts. The reply points at a graph that holds the detail rather than retelling it, and if a paragraph would make a good artifact, it IS one: log it and cut it. Every reply now ends with what is still to do, as a list — the one thing a human cannot reconstruct from the graph, and the thing that sets direction. In a multi-practice environment their attention is the scarce resource, far scarcer than the practitioner's.
  • --publish is tag-and-push; CI owns every channel. release.yml already defined build, both PyPI packages, Docker, Homebrew and the GitHub release, so the local path was redundant — and running both made every release race on the GitHub release (a release with the same tag name already exists, recovered with --clobber, three times in one day). CI_CD.md set the bar at "verified for a release or two"; 1.13.4, 1.13.5 and 1.13.6 all published cleanly through CI. --local-artifacts restores the old path for when CI is down — an escape hatch, not an alternative, since running both is what caused the race. A test asserts release.yml still defines every channel job, so removing one fails there rather than at the next release.
  • setup reports the prompt's CONTENT, not just its version. It now prints a sha256 digest, word count and source path beside the version, so two seats can compare what was installed instead of comparing claims about it. A version string is provenance, never evidence of content — a distinction that cost three separate diagnoses in one day (PyPI's .info.version and .releases both lagging behind the simple index; an interpreter path read as a code version; a prompt header read as proof of the prompt's text).
  • doctor: Presence coverage. Since the heartbeat emitter became practice-scoped, a practice with live sessions and no running listener goes dark on the mesh — previously any listener carried it. Zero-impact today (23 records, 22 stale, the one live record has its listener) but that is circumstance, not construction. Also catches label drift, where a record written as workspace while its listener runs as empirica-workspace is orphaned just as effectively.
  • emitter_id on the practitioner heartbeat (<ai_id>:<pid>). The payload carried machine and session_id and nothing identifying the SENDER, so N listeners posting one session every 60s was indistinguishable, receiver-side, from one emitter posting every 60/N — an ambiguity that produced a reported interval defect which did not exist. Additive and optional; an unknown field is ignored by the handler.

What's New in 1.12.35

  • docs-assess tech_docs rewarded name-dropping. _check_if_documented counted a feature as documented if its NAME substring-matched the concatenated markdown, so the metric was satisfiable by dumping class names into a .md and unmovable by real documentation — inverting the EU AI Act Art. 11 / ISO 7.5 intent it is framed against. Measured by empirica-workspace: 137 accurate docstrings moved coverage 0%, while a generated file listing 256 names took it to 100%. A feature now counts when it has a substantive docstring OR the docs carry real prose about it (the mention's line must retain ≥8 words once list/table/heading punctuation and the name are stripped). The OR is deliberate: gating on docstrings alone would penalise practices that document in markdown. check_docstrings now also returns documented_symbols — it already computed that truth to count documented_items but never named them, so the docstring half had nothing to consult.
  • A listed source could 404 on /content. The sources LIST reads the whole per-practice DB while the content lookup was still scoped to one project_id, so a source with a drifted id appeared in the pane and failed the instant it was opened (10 of 50 on one practice). Both are practice-scoped now.
  • docs-assess crashed on a stale project pointer. _auto_detect_project_config called iterdir() on a path taken from resolver state without checking it exists — resolver state outlives the directory it names, and a pointer to a deleted project took the whole command down.
  • Review stamping could break the audit. The new cadence opened a session DB unconditionally, so an environment without one (CI, a bare checkout) failed the check instead of simply not recording.
  • empirica setup — harness-neutral name for setup-claude-code, which stays as an alias. The old name leaked "claude-code" into model-facing prose (hook errors, skill docs) that other harnesses vendor verbatim; the instruction text is swept too. One parser, one handler, two names — no new capability.
  • sources-check sees LOCAL file rot. It gated on http(s), so file-backed sources were skipped entirely — one practice reported "all probed source links resolve" while 25 of its 50 sources could not be served. Non-URL sources are now classified ok / missing / not_a_locator (the column holds a title, needing a re-point rather than a file hunt) / out_of_scope (a mailto:/ftp:/doi: URI the disk check cannot speak to). Path resolution mirrors the daemon's, pinned by a test.
  • source-update --url — re-point a source whose file MOVED. Gardening is prune and replant, but the CLI could only re-fetch, never re-target, so a moved doc had to be archived and re-added — losing its id and with it every sourced_from edge pointing at it. --url retargets in place, recomputes content identity, and records a repointed event in lifecycle_audit_log so the move is traceable.
  • A review cadence. sources-check now records a timestamped verdict per source (last_reviewed_at, review_verdict) — previously nothing ever wrote those columns (0 of 63 reviewed), so "unchecked since X" was unanswerable and a source could never be more than an assertion with a date on it.

Privacy & Data

Your data stays local:

  • .empirica/ — Local SQLite database (gitignored by default)
  • .git/refs/notes/empirica/* — Epistemic checkpoints (local unless you push)
  • Qdrant runs locally if enabled

No cloud dependencies. No telemetry. Your epistemic data is yours.


Community & Support


License

MIT License — see LICENSE for details.


Author: David S. L. Van Assche Version: 1.13.7

Turtles all the way down — built with its own epistemic framework, measuring what it knows at every step.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

empirica-1.13.7.tar.gz (2.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

empirica-1.13.7-py3-none-any.whl (2.5 MB view details)

Uploaded Python 3

File details

Details for the file empirica-1.13.7.tar.gz.

File metadata

  • Download URL: empirica-1.13.7.tar.gz
  • Upload date:
  • Size: 2.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for empirica-1.13.7.tar.gz
Algorithm Hash digest
SHA256 462ca93c27cc3055e8649c04213e2a63299a1e08ceee935046f8b18abb7f177c
MD5 748f86158e3bc7d094cb534fee289884
BLAKE2b-256 62854a8d7dd256e335246a6412f78726ffadf177af8f743207c9947d8f8b7e72

See more details on using hashes here.

Provenance

The following attestation bundles were made for empirica-1.13.7.tar.gz:

Publisher: release.yml on EmpiricaAI/empirica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file empirica-1.13.7-py3-none-any.whl.

File metadata

  • Download URL: empirica-1.13.7-py3-none-any.whl
  • Upload date:
  • Size: 2.5 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for empirica-1.13.7-py3-none-any.whl
Algorithm Hash digest
SHA256 68cdf7422f5992973216daeb0ea193a40c13405177ab9d4c003539b840d93063
MD5 43b433fe5d9c0c4e5fd704b567ca8a13
BLAKE2b-256 9f2eefde7a7eaea7a2aa76e699fa26b606cde61bfc957611afdac55be7b55bfc

See more details on using hashes here.

Provenance

The following attestation bundles were made for empirica-1.13.7-py3-none-any.whl:

Publisher: release.yml on EmpiricaAI/empirica

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page