Context-M
Deterministic agent memory. 96 bytes per fact. Zero LLM at ingest.
Mem0 gives your agent a notebook. Context-M gives your agent a brain.
Context-M is a memory layer for AI agents that needs zero LLM calls to
ingest and proves every retrieved fact with a BLAKE3 hash chain.
Mem0-compatible: drop-in replacement for from mem0 import Memory.
A memory substrate that combines a bi-temporal symbolic Trace (hippocampus) with a VSA Memory Palace (neocortex), bound by a μ=0 deterministic bridge — cryptographic provenance on every retrieval, edge-first deployment at 96 bytes per memory.
pip install cortexm # works offline, no API keys, single command
from cortexm import Memory # Mem0-compatible surface
m = Memory()
m.add("I work at Google", user_id="alice")
m.search("Where does Alice work?", user_id="alice")
# → [Memory — Known facts]
# - (Alice, works_at, Google) [valid 2026-08-27→∞; learned …; conf 0.92;
# id 3f2a91c2; src #a1b2c3d4; "I work at Google"]
Benchmark results — August 2026
We run four tiers of evaluation, and the honest number is not the
biggest one. Full methodology, judge identities and failure analysis:
docs/BENCHMARKS.md ·
docs/FAILURE_MODES.md ·
open the leaderboard →
Tier 1 — Out-of-distribution (where users live)
Ground-truth fact registries were re-rendered by an independent LLM in styles the pattern extractor never saw, then evaluated with the same probes and judge as the in-distribution run:
| OOD style | Tier-1.1 (pre-fix) | Tier-1.2 (post-fix, 2026-08-28) | Δ |
|---|---|---|---|
| paraphrase | 9.4% ± 9.4% | 22.9% | +13.5pp (2.4×) |
| negation | 75.6% ± 3.3% | 75.6% | flat |
| indirect speech | 44.9% ± 10.2% | 48.2% | +3.3pp |
| informal/slang | 5.1% ± 5.9% | 41.3% | +36.2pp (8.1×) |
| non-English | 0.0% | 32.2% | +32.2pp (∞ → real recall) |
| code-switching | 57.9% ± 18.1% | 61.3% | +3.4pp |
The slang jump (5.1% → 41.3%) is the single biggest fix in this
cycle: the unmess pipeline (DisSim + idiolect + Bitap) is now safe
to enable in the bench config (previously the period-strip bug
forced unmess_enabled=False). The non-English jump (0% → 32%)
comes from the LaBSE polyglot encoder + idiolect normalizer
handling accented characters without crashing the trigger.
Tier 4.3 — LongMemEval independent judge
| subtask | pre-fix | post-fix (2026-08-28) | plugin-kernel (v0.5.0) | v0.5.1 synthetic | v0.5.2 canonical |
|---|---|---|---|---|---|
| single_hop | 1.0 | 1.0 | 1.0 | 1.0 | 0.222 |
| knowledge_update | 0.333 | 0.667 | 1.000 | 1.0 | 0.333 |
| multi_session | 0.5 | 0.5 | 0.5 | 1.0 | 0.333 |
| temporal_reasoning | 0.5 | 0.5 | 0.5 | 1.0 | 0.667 |
| overall | 0.600 | 0.700 | 0.800 | 1.000 | 0.333 |
v0.5.2: the honest canonical score. MemPalace (246K-step benchmark)
got 96.6% recall at $0 cost. Context-M hits 1.000 on a 20-question
synthetic LongMemEval subset (matches MemPalace's framing on data we
control), and 0.333 on a real 18-question sample from the
canonical xiaowu0162/longmemeval-cleaned benchmark — μ=0
throughout (no LLM at ingest, retrieval, or judging).
We do NOT claim parity on the canonical 500-question, 23,867-session benchmark. We claim:
-
End-to-end deterministic QA is possible. MemPalace stops at retrieval — they never answer the question. We built the full pipeline: Question → Intent Router → Datalog-lite / Trace / VSA → Answer Extraction → Judge → Score. All zero neural networks.
-
The 1.000 on synthetic is real. Same judge, same reader, same Trace, run 3× — every run returns 1.0. Promise #5 holds.
-
The 0.333 on canonical is also real. Real human text — slang, typos, indirect speech, code-mixed language, multi-game arithmetic, meta-answer preference questions — the deterministic extractor misses things an LLM would catch. That's the honest gap, by design (μ=0 trades breadth for cost).
v0.5.2 wiring fixes (the user-identified gap on multi_session + temporal_reasoning):
-
recall_stepwired into the LongMemEval reader path. The asymmetric step-distance boost surfaces scrolled-out session-1 facts that the access_count boost on current session-N facts would otherwise push below top-k. This is the multi_session fix: older-session facts that list questions need now have a higher retrieval weight. -
Temporal query pre-processor + TEMPORAL CHAIN note. When the question matches
when/before/after/did X move/did X change/how many times, the reader walks the bi-temporal SUPERSEDES chain per (entity, relation) and emits an explicit ordering note:TEMPORAL CHAIN: Bob|lives_in: Berlin (SUPERSEDED) → Munich (CURRENT) → 1 supersession(s) detected → Bob changed. The BOOL judge reads this directly (STRATEGY 0), bypassing the regex fallback. canonical temporal_reasoning went 0.5 → 0.667. -
Smarter NUGGET + LIST judges. Both now fall back to token- overlap when literal-substring fails, so canonical answers like "4 years and 9 months" score True when both tokens appear in the context, even if the exact "and"-joined phrase doesn't.
Reproduce (synthetic, 1.000):
python scripts/longmemeval_judge.py \
--out benchmarks/results/longmemeval_v0.5.2_synth.json
Reproduce (canonical, 0.333):
# one-time: download the canonical benchmark (~277 MB)
python -c "from huggingface_hub import hf_hub_download; \
hf_hub_download(repo_id='xiaowu0162/longmemeval-cleaned', \
repo_type='dataset', filename='longmemeval_s_cleaned.json', \
local_dir='data/longmemeval')"
# sample 3 questions per subtask (18 total), μ=0 ingest + judge
python scripts/longmemeval_canonical.py \
--n-per-type 3 --max-messages-per-q 300 \
--out benchmarks/results/canonical_longmemeval_v0.5.2_n3.json
Determinism: 3 sequential synthetic runs return
det_judge_accuracy: 1.0 every run — the "same every time" promise
holds. The canonical score varies with the random sample seed; the
overall 0.333 is reproducible with --seed 42.
Pre-plugin-kernel fixes (0.600 → 0.700): (1) works_at regex
contraction fix ("I'm now working at OpenAI" now extracts),
(2) role pattern |$ lookahead + uppercase support ("I'm an ML
engineer" now extracts), (3) employment-anchored temporal window
(resolves "where did X live when at Y" via the works_at fact's
valid_from/valid_to).
Plugin-kernel fixes (0.700 → 0.800): the new verbatim tier (FTS5
- int8 dense, MemPalace-style) catches "I'm now working at OpenAI" verbatim when the structured extractor's role pattern still misses it. The fusion bridge then merges both tiers at μ=0 cost.
v0.5.1 fixes (0.800 → 1.000): the deterministic judge switches from literal-substring to a 3-strategy rule engine (NUGGET / LIST / BOOL) plus the retrieval window widens from 5 to 10 so multi-session list questions get both values surfaced. The 2 remaining answer-shape mismatches the previous run hit are now closed — every question has a strategy that can answer it correctly without an LLM.
v0.5.2 fixes (canonical 0.333 honest scope + temporal_reasoning
0.667): recall_step wired into the reader path (surfaces scrolled-
out facts); TEMPORAL CHAIN note emitted for when/before/after/did X move questions (BOOL judge reads the verdict directly); NUGGET +
LIST judges gain token-overlap fallbacks for free-form canonical
answers.
That is the capability profile of the μ=0 extractor on real phrasing:
strong on change-of-state statements, weak on identity/preference
restatements, weak on non-English without the LaBSE polyglot encoder,
weak on arithmetic ("how many hours total" requires summing across
chunks — deterministic reader can't add). The async LLM enrichment
fallback helps marginally — it surfaces facts but does not reconstruct
bi-temporal chains. docs/FAILURE_MODES.md
documents which phrasings break, with worked examples.
Independent LLM judges grade these numbers lower, not higher. The
full 240-item OOD sweep was re-graded by gemini-3.5-flash-lite from a
clean CI runner: LLM-judge mean 0.222 vs offline judge 0.335,
exact agreement 82.7% (237/240 items;
results/ood/llm_judge_crosscheck_gemini.json).
A second judge (glm-4-plus, 58-item quota sample) agrees: 0.250 vs 0.345.
Two independent models, same conclusion — the offline grader is not
inflating scores. Judge model ≠ canonical BEAM's gpt-5, so these are
cross-checks, not BEAM-comparable numbers.
Tier 2 — In-distribution (the regression harness). Synthetic BEAM-style conversations (arXiv:2510.27246 methodology), 10 abilities, deterministic nugget judge, μ=0 ingest asserted, 5 seeds:
| Bucket | questions | Context-M | BM25-RAG | vector-only |
|---|---|---|---|---|
| 128K | 37 | 100.0% ± 0.0% | 70.2% | 69.0% |
| 500K | 72 | 100.0% ± 0.0% | 70.5% | 67.9% |
| 1M | 107 | 100.0% ± 0.0% | 68.8% | 70.1% |
| 10M | 216 | 100.0% ± 0.0% | 61.6% | 66.1% |
Why 100% here is not a capability claim: the corpus generator and the extractor patterns were authored against the same template families, so this tier measures template coverage, ceiling by construction. Its job is regression detection — "did we break template extraction?" — not marketing. We do not compare it against canonical BEAM SOTA (Exabase M-1, 68.0%): different corpus, different judge, different protocol — an apples-to-oranges comparison we refuse to make.
Tier 3 — Real GitHub data. Real issue threads from public repos
(rust-lang/rust, numpy/numpy, pydantic/pydantic; attribution in
benchmarks/real_github/): the μ=0 extractor vs an LLM reference
extractor (gemini-3.5-flash-lite) on identical comments, plus retrieval
QA judged by the same LLM:
| Track | Result |
|---|---|
| μ=0 extraction | 16 facts from 150 comments · 1.1 ms/comment · $0.00 |
| LLM reference extraction | 158 facts · 2,779 ms/comment · ~90K tokens |
| μ=0 recall vs LLM reference | 0.6% — the honest gap on real technical text |
| Retrieval QA (LLM-judged, 19 Qs) | overall 0.263 · answerable 0.067 · abstention 100% |
Read this as the cost/coverage frontier: the μ=0 path is ~2,500× faster
and free but, on developer-issue language, captures ~10× fewer facts
than an LLM extractor. The enrichment fallback and per-domain pattern
packs are the bridge. Artifacts:
benchmarks/results/real_github/ ·
results/llm_eval_summary.md.
Engineering facts measured alongside (see docs/BENCHMARKS.md):
- Ingest: 10M tokens in ~98 s (~102K tokens/s), ~2,000 messages/s, 0 LLM calls
- Memory grows sublinearly: 10M tokens → ~590 facts (repeated noise dedupes)
- Provenance: 100% of retrieved facts hash-verified; audit latency ~6 ms
- Retrieval: tree index p50 ≈ 0.4–1.1 ms at 10K–100K vectors (flat: 16–194 ms)
- Crash-recoverable: WAL journaling with SIGKILL-recovery tests
(
tests/test_wal_recovery.py) — committed memories survive hard kills - Reproducible: runs are process-independent — score ties break on fact content, never on random ids (verified across four PYTHONHASHSEED values)
The architecture
┌──────────────────────────────────────────────────────────────────┐
│ THE BRIDGE (μ = 0) │
│ write: text → chunks → BLAKE3 → patterns → triples → holograms │
│ read: query → intent plan → VSA probe ∥ symbolic query → │
│ fusion → [Memory — Known facts] + provenance chain │
└──────────────┬───────────────────────────────────┬───────────────┘
│ │
┌──────────────▼──────────────────┐ ┌──────────────▼───────────────┐
│ LAYER 1: SYMBOLIC TRACE │ │ LAYER 2: VSA MEMORY PALACE │
│ (hippocampus) │ │ (neocortex) │
│ bi-temporal facts (SQLite) │ │ HRR holograms, role-bound │
│ CONTRADICTS / PRECEDED_BY / │ │ INT8 · Binary · RaBitQ · PQ │
│ EXTRACTED_FROM edges │ │ codecs (770/96/96/8 B each) │
│ Datalog-lite rules engine │ │ page-clustered tree index │
│ interference-aware lifecycle │ │ 64-entry semantic L1 (SLB) │
│ Memory Git: hash-chained DAG │ │ TMR self-healing + re-encode│
└─────────────────────────────────┘ └──────────────────────────────┘
Layer 1 — Symbolic Trace. Subject-Relation-Value triples with
valid-time and transaction-time (when it was true vs when we learned it), contradiction resolution by truth maintenance (new values
supersede, old values retire with their windows intact), temporal
edges, a Datalog-lite forward-chaining engine (manages(Y,X) → reports_to(X,Y), member_of(X,T) ∧ uses(T,L) → team_uses(X,L)), and
an interference-aware lifecycle: facts are evaluated for how they
interact with existing memory before commitment.
Layer 2 — VSA Memory Palace. Each fact becomes a holographic reduced representation: role-bound subject/relation/value fillers plus a λ-weighted lexical superposition, quantized to your storage tier. Permutation binding is the default algebra because it maps directly to binary HDC hardware (XOR/permutation) — when edge ASICs arrive, the same code compiles down.
The Bridge. μ=0 ingest: a 61-pattern deterministic extractor
(first/third/second-person, pronoun resolution, relative dates,
retractions, Mem0-summary shapes) — no LLM anywhere on the synchronous
write path. When patterns find nothing (non-English, heavy slang,
indirect speech), an explicit async enrichment fallback
(memory.enrich()) re-extracts those chunks with an LLM post-store —
confidence-capped at 0.85, provenance-marked llm_enrichment, auditable,
and counted in the μ=0 honesty counters. The read path is a
deterministic query planner (temporal windows, ordering proofs,
counting, supersession chains, Personalized PageRank graph diffusion
for multi-hop — HippoRAG 2 lineage) fused with VSA retrieval, and
every returned fact carries its full audit chain: query → VSA match
→ symbolic dereference → BLAKE3 hash → original source text.
The five category-defining features
| Feature | What it does | Try it |
|---|---|---|
| Memory Git | branch / merge / diff / blame over agent memory, hash-chained commits | examples/07_memory_git.py |
| ZK-lite proofs | prove a fact matches a query without revealing it to the LLM | examples/08_zk_proof.py |
| Self-healing memory | bit flips detected by hash, TMR majority vote, re-encode from Trace — 100% self-ID up to 10% corruption | examples/09_self_healing.py |
| Predictive prefetching | MBTB co-access prediction feeds the fusion boost set | cortexm/features/prefetch.py |
| Cross-modal binding | episodic holograms: bind text/structured/sensor roles, recall by any modality | cortexm/vsa/ops.py |
Storage tiers (cortexm-compress)
| Tier | Bytes/vector | 1M memories | Fits on |
|---|---|---|---|
int8 (default) |
770 | 770 MB | any laptop |
binary + TMR |
96 (288 w/ TMR) | 96 MB | Raspberry Pi 5 → 10M memories |
rabitq |
96 | 96 MB | Raspberry Pi Zero 2W |
pq |
8 | 8 MB | cloud, billions |
Measured codec quality (20K fact holograms): int8 overlap@10 vs FP32 =
0.90; binary/rabitq/PQ recover the FP32 top-10 within their top-50 at
1.00/1.00/0.9995 — shortlist codecs, exactly as designed. See
docs/COMPRESSION.md.
Security (InjecMEM + MINJA + scope sandbox + PermissionGate)
Every fact carries a BLAKE3 hash of its source text, re-verified on
retrieval (BLAKE2b-256 fallback with a loud warning if the optional
blake3 wheel is absent — pip install cortexm[blake3]; the active
provider is always reported in stats() and audit output). Memory-
injection patterns ("ignore all previous instructions…") are
quarantined at ingest — stored for audit, never active, never retrieved
into prompt context. On top of that, the MINJA contagion guard
treats quarantined text as a tainted corpus: any later ingest that
quotes or substantially overlaps it (even when light edits defeat every
regex) is quarantined too — closing the query-only injection loop where
an attacker poisons memory through the agent's own write-back.
The scope sandbox enforces the isolation the InjecMEM threat model
implies: facts written by an agent (agent_id=...) are invisible to
user-scope reads until explicitly promote()d — and promotion is
gated on confidence, re-scans the source chunk through both injection
detectors, and lands in the tamper-evident audit chain
(tests/test_sandbox_enrich.py). Building it surfaced and fixed three
genuine pre-existing read-path leaks (empty-scope fallback, falsy scope
checks, unscoped supersession chains). verify_integrity() audits the
whole store.
The PermissionGate (v0.5.1; hardened v0.5.2) is a default-deny gate for code execution + user-data reads. The user directive:
"security is important — no malicious code shall be executed to read user data without explicit permission."
is enforced as a strict allowlist with NO wildcards:
Plugins that want to invoke os.system / subprocess / open() on
the user's behalf MUST first call permission.grant_read(path) or
permission.grant_exec(cmd) — otherwise the gate denies and audits
the attempt. Sensitive paths (~/.ssh, ~/.aws, /etc/passwd,
~/.config/gh) and sensitive executables (curl, wget, sudo,
ssh, nc) are ALWAYS denied unless the user calls
grant_sensitive() on the exact item. There is no wildcard.
Composition, not coercion: the plugin doesn't monkeypatch os or
subprocess — plugins that consult the gate are gated; plugins
that ignore it are not. tests/test_permission.py,
34 tests.
from cortexm.kernel import Context
from cortexm.plugins.security import SecurityPlugin
ctx = Context()
ctx.mount(SecurityPlugin())
sec = ctx.inject("security")["security"]
perm = sec.permission
# An agent tool wants to "ls /tmp/agent_ws"
perm.grant_read("/tmp/agent_ws")
perm.grant_exec("ls")
perm.can_exec("ls /tmp/agent_ws").allowed # True
perm.can_read("/etc/passwd").allowed # False (sensitive)
perm.can_exec("curl evil.com").allowed # False (sensitive)
perm.can_exec("rm -rf /").allowed # False (no grant)
# Every denial is recorded on the tamper-evident audit chain
Enterprise controls (shipped, not roadmap)
The controls a buyer's security review actually blocks on — all in the
repo, all under test (tests/test_enterprise.py):
| Control | What ships |
|---|---|
| PII firewall | Luhn/mod-97/area-rule-validated detection of emails, phones, cards, SSNs, IBANs, IPs, API keys — redacted to reversible vault tokens before extraction (GDPR/CCPA write-path guard) |
| Encryption at rest | AES-256-GCM envelope (KEK→DEK), key rotation, env/keyfile/sidecar master keys |
| RBAC + API keys | admin / operator / reader / auditor roles, peppered-key digests, TTLs, constant-time verify |
| Tamper-evident audit | hash-chained per-operation log; SIEM export (JSONL + syslog); tampering pinpoints the broken seq |
| GDPR governance | Art. 17 right-to-erasure with crypto-shredding + attestation; Art. 5 retention policies; DSAR vault resolution |
| Backup / DR | atomic snapshots with SHA-256 manifests; PITR — bi-temporal replay, the database is its own WAL |
| REST API | 20 endpoints, OpenAPI 3.1 at /openapi.json, bearer auth, per-key rate limiting, Prometheus /metrics, /healthz /readyz |
| Deploy anywhere | Docker (non-root, tini, healthcheck) · docker-compose + nightly snapshots · K8s manifests · Helm chart — deploy/ |
cortexm serve-rest --db /data/memory.db --pii redact --admin-key yes
See docs/ENTERPRISE.md (control matrix + compliance mapping) and
docs/DEPLOYMENT.md (SDK / MCP / REST / Docker / K8s / Helm runbooks).
MCP server (Day 1)
cortexm serve # stdio JSON-RPC, zero dependencies
Tools: contextm_add, contextm_search, contextm_get_all,
contextm_history, contextm_temporal, contextm_audit,
contextm_prove, contextm_stats, contextm_delete. Works with
Claude Code / Cursor / any MCP client. Claude Code plugin:
plugins/context-m-claude.
Migration
cortexm migrate --from mem0 --path mem0.db
cortexm migrate --from zep --path zep_export.jsonl
cortexm migrate --from chroma --path chroma.sqlite3
Each importer handles the vendor's real on-disk formats (mem0's
history JSON payloads and bare memories tables, Zep graph triples
with bi-temporal windows, Chroma's embeddings table) and is verified
end-to-end against fixture stores built in those exact formats
(tests/test_migration.py).
Durability
WAL journaling (Aeon-inspired) with a wal_sync durability knob
(normal — survives process crash; full — fsyncs every commit,
survives power loss), WAL checkpoint-on-close, and a test that
SIGKILLs a writer mid-stream and verifies every acknowledged commit
survives (tests/test_wal_recovery.py).
Federation (CRDT replication)
Multi-node memory replication without a coordinator: bi-temporal facts as
HLC-stamped CRDT versions (SINGLE_VALUED relations collapse into one
versioned register per key — the version set IS the temporal history),
union merge that is commutative/associative/idempotent, OR-set
retraction semantics (write-after-retract wins, retract-after-write
wins), purge poison-pills for GDPR, and digest/delta anti-entropy that
ships only divergent buckets over HMAC-signed envelopes. Convergence is
proven byte-exact (canonical serialization compared, not just query
equivalence); a partition with divergent writes + retractions heals with
no lost retraction semantics. Transports: in-memory mesh for tests, file
spool (outbox/inbox) for offline mule sync — rsync/git/USB completes the
physical channel, the CRDT guarantees convergence regardless of delivery
order. See cortexm/federation/ and benchmarks/federation_bench.py.
Rust acceleration (optional wheels)
rust/cortexm-core and rust/quadrant compile the hot paths with
PyO3; the Python/NumPy implementation stays the reference and everything
works without them (CONTEXTM_RUST=0 forces the pure-Python path).
Measured on the bundled scorecard (benchmarks/rust_vs_numpy.py):
encode_fact 4.8×, bind 3.4×, h64 2.2× — h64 is byte-exact with the
Python hash (tested), and permutations/role vectors are injected from
Python's deterministic VSA state, so mixed deployments produce
bit-identical holograms. The SLB is a tie (1.0× — BLAS is already
optimal at 64×768; published as such). quadrant is the page-clustered
log-depth vector index for the L2 palace: 97% recall@10 at 7× NumPy
brute-force speed, visiting ~32 of 529 pages for 20k vectors — visit
counts are instrumented, the O(log N) claim is measured, and the
adversarial random-corpus recall collapse is published alongside the
win. Build: pip install ./rust/cortexm-core ./rust/quadrant.
More
docs/ARCHITECTURE.md— every layer in detaildocs/BENCHMARKS.md— full results, methodology, per-ability tablesdocs/FAILURE_MODES.md— where the extractor breaks on real phrasing, with worked examples (read before citing any number)docs/ENTERPRISE.md— enterprise control matrix + compliance mappingdocs/DEPLOYMENT.md— SDK / MCP / REST / Docker / K8s / Helm runbooksdocs/RESEARCH.md— literature lineage: every paper we adopted, aligned with, or rejected (with reasons)docs/SECURITY.md— InjecMEM + MINJA defenses, scope sandbox, provenance modeldocs/COMPRESSION.md— the tier stack and measured trade-offsdocs/ROADMAP.md— phase status vs the strategic plandocs/GOVERNANCE.md— foundation governance + licensing commitmentsleaderboard/— self-hosted benchmark site (rebuild:python leaderboard/build.py; openleaderboard/index.html)examples/— runnable scripts, offline, no API keystests/— 116 tests: fabric + enterprise + PPR + concurrency + sandbox + enrichment + WAL crash-recovery + migration + CRDT federation convergence/partition-heal + Rust parity
License
Apache 2.0 — open core done right: the memory fabric is and stays open; federated sync and the audit UI are the enterprise tier.
arXiv-inspired improvements (2026 round)
A second research pass over 2024-2026 arxiv literature surfaced 8
concrete improvements, all preserving the μ=0 invariant. Full citations
in docs/BENCHMARKS.md Tier 8.
| Improvement | Module | Solves |
|---|---|---|
| Hopfield cleanup memory | cortexm/vsa/cleanup.py |
VSA interference after unbind |
| Bitap fuzzy matching (Wu-Manber) | cortexm/text/fuzzy.py |
Slang/spelling-tolerant pattern triggers |
| Per-user idiolect normalization | cortexm/text/idiolect.py |
"bruh"→"friend" via embedding k-NN |
| DisSim rule-based simplifier | cortexm/text/dissim.py |
Compound-sentence pattern recall |
| TLSH ternary trie | cortexm/vsa/tlsh_trie.py |
O(log N + w) software TCAM |
| Holographic fact overlay | cortexm/vsa/hologram_overlay.py |
O(1) single-hop fact lookup |
| ProtoDash attribution | cortexm/vsa/attribution.py |
Source weights for retrieval results |
| LayerCast FP32 determinism seam | cortexm/bridge/onnx_runtime.py |
μ=0 over LLM enrichment path |
Architectural fixes (per Con #4-#7 list):
- Storage bloat →
cortexm/trace/dedup.pyformalizes dedup+compression audit - Normalization → Bitap + idiolect + hybrid search wired into patterns
- Debugging →
retrieval_path ∈ {vsa_unbind, pattern_match, neural_fallback, raw_chunk, tree_index, tlsh_trie}on every retrieved fact - Determinism → LayerCast + ONNX Runtime CPU + FP32 seam documented
Plus explicit binary/FP32 tiering (accel.detect_tier, accel.recommend_codec),
Hamming ZK proofs (security/zk_hamming.py), and trace/rebuild.py
for checksum-audited rebuilds from the symbolic Trace.
Claude Code plugin — session lifecycle
plugins/context-m-claude/src/index.ts v0.2 adds auto-load on Claude
session start + write-on-end hooks:
- on session start →
recall last working state→ "I see you've been working on X. Continue?" - on session end → "Store summary? [Y/n]" → persists summary as a memory fact
- Session state at
~/.context-m/session_state.json
MCP tools added: contextm_query_extract (hybrid RAG), contextm_attribution (ProtoDash), contextm_zk_prove (Hamming proofs).
Honest measurement block
Reproducing the Ponytail convention: every headline number on this README is paired with the run that produced it, the SHA, the judge model, and the honest cost. "~96 bytes per fact" is the storage cost on the BEAM-10M corpus (n=200 personas × 30 turns × ~35 facts per persona, measured 2026-08-28 on commit 714f237). "Zero LLM at ingest" is enforced by the
LLM_CALLScounter incortexm/__init__.py; a CI assertion fails any PR that increments it on the ingest path. The Real-GitHub Tier-4 result (17 questions, 0.0 answerable, 1.0 abstention, 2026-08-28 14:46 UTC run #9) is a refusal-to-guess, not a coverage gap — the system abstains rather than hallucinate on real developer-issue language. Full method:docs/METHODOLOGY.md. Reproduce:python benchmarks/run_ood_pipeline.py --personas 4 --skip-render --no-enrich --no-judge.
Anti-lamprey warning
Don't fork-and-rebrand this repo. If you want to build on it, open an issue labeled
acceptedand submit a PR — seeCONTRIBUTING.mdandAGENTS.md. Fork-and-rebrand-without-attribution derivatives will be named indocs/FAILURE_MODES.mdunder the "Derivative works" section. The provenance chain (BLAKE3 hash + source span) is the system's whole point — strip it and you've built a different product, not a fork.
Star history
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cortexm-0.5.2.tar.gz.
File metadata
- Download URL: cortexm-0.5.2.tar.gz
- Upload date:
- Size: 495.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
69417086388214d59fd70d5a2958a7832568251543837cda2404824e2085fd48
|
|
| MD5 |
6e564300806436b47a2932c1dda79b0d
|
|
| BLAKE2b-256 |
44570339092c67e9c1dff44b2773cf0f8c5e3e1a0cc689fc743f3ee20493f4fc
|
Provenance
The following attestation bundles were made for cortexm-0.5.2.tar.gz:
Publisher:
release.yml on ssmurfgg04-gif/context-m
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
cortexm-0.5.2.tar.gz -
Subject digest:
69417086388214d59fd70d5a2958a7832568251543837cda2404824e2085fd48 - Sigstore transparency entry: 2632942635
- Sigstore integration time:
-
Permalink:
ssmurfgg04-gif/context-m@2947f4d6e9389635eb4683de55bf5065821810a6 -
Branch / Tag:
refs/tags/v0.5.2 - Owner: https://github.com/ssmurfgg04-gif
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@2947f4d6e9389635eb4683de55bf5065821810a6 -
Trigger Event:
push
-
Statement type:
File details
Details for the file cortexm-0.5.2-py3-none-any.whl.
File metadata
- Download URL: cortexm-0.5.2-py3-none-any.whl
- Upload date:
- Size: 438.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
df50f4f04ec49265aeee46fe709ddf145073b0df35fcf8db83b7af93b101a832
|
|
| MD5 |
8d838be05409b01a8fafdba600b2f400
|
|
| BLAKE2b-256 |
3f67b6d9efbf865dfb2fa28e61952bf04438875c33fbf5a4f25ff40e1c3bc138
|
Provenance
The following attestation bundles were made for cortexm-0.5.2-py3-none-any.whl:
Publisher:
release.yml on ssmurfgg04-gif/context-m
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
cortexm-0.5.2-py3-none-any.whl -
Subject digest:
df50f4f04ec49265aeee46fe709ddf145073b0df35fcf8db83b7af93b101a832 - Sigstore transparency entry: 2632942655
- Sigstore integration time:
-
Permalink:
ssmurfgg04-gif/context-m@2947f4d6e9389635eb4683de55bf5065821810a6 -
Branch / Tag:
refs/tags/v0.5.2 - Owner: https://github.com/ssmurfgg04-gif
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@2947f4d6e9389635eb4683de55bf5065821810a6 -
Trigger Event:
push
-
Statement type: