Skip to main content

Mountain of Helicon

Your AGENTS.md is lying to your coding agent. It points at files that moved, commands that no longer exist, paths the repo reorganized away. Your agent loads those rules as fact at the start of every session. Nothing tells you until something breaks.

pip install mountain-of-helicon
helicon review .

Needs Python 3.10 or newer. On a Mac, the python3 that Apple's Command Line Tools install is 3.9, and its pip then says only No matching distribution found. On a Mac, run brew install python@3.12, then python3.12 -m pip install mountain-of-helicon. Or run uv tool install mountain-of-helicon, which fetches a matching Python for you.

What it printed on our own public repo, MorkeethHQ/world-relay, from a fresh clone of main on 2026-09-03:

  ✗ 1 graded claim contradicted by repo evidence.

    ✗ AGENTS.md:19  points at src/__tests__/e2e-api.test.ts  — not in this repo

  GRADE B   ·   35 claim checks, 1 contradiction
  1 instruction claim conflicts with observed repo evidence.

The test moved to src/lib/__tests__/e2e-api.test.ts; the rule that tells the agent how to run it did not move with it. The same command on anthropics/anthropic-cookbook, same day, prints GRADE A, 6 references checked, 0 broken. An earlier build of this tool graded that repo D on four pointers that all resolved; the fix and its fixtures are in docs/HELICON-ON-OUR-OWN-BOARD-2026-09-03.md.

No API key. No config file. No upload. It reads the agent rules your repo already commits (AGENTS.md, CLAUDE.md, .cursorrules) and checks every pointer against the tree on disk. External host paths and create-on-demand outputs are shown as unverified and excluded from the grade; they cannot pad or lower it. Exit code is non-zero when a graded repo-local claim fails, so it drops straight into CI.

What the install pulls. pip install mountain-of-helicon installs this package and no third-party packages. The standard library is all the review needs. The web app, the model-backed commands and the memory lab sit behind extras: [web], [model], [retrieval], [embeddings], or [all]. A command that needs one says which, in one line: helicon serve needs the web extra: pip install "mountain-of-helicon[web]".

Second command: helicon witness grades your agent's last session instead of your repo: which of its "done / passes / fixed" claims have tool evidence in the same trace. On the author's own last 20 sessions the median verified-claims share was 0.42 (measured 2026-09-02; 10 of the 20 carried a checkable claim).

The other two doors: your agent's memory, and its claims

The review above reads a repo. The same engine reads a memory directory and a session transcript, with no key and no config:

helicon truth ~/.claude --recursive   # which of your agent's documents are lying, and why

No API key. No database. No config at run time. It reads a directory and returns a ranked report, and every row cites the line it fired on. On this machine, that first run reads:

1182 files scanned · 628 carry a staleness/rot signal · 554 clean
  1   39   GOLDEN_RULES.md
        +22  expired dated claim (108d past, 2026-05-15)  -> "Close-Out - May 15"
helicon witness           # your last agent session: every claim vs its evidence
helicon setup --audit     # local Claude, Cursor and Codex setup evidence

For a specific project, run helicon setup --audit --project /path/to/repo.

The Setup screen also includes an independent, read-only ZUP project-state review. It compares local event receipts, canonical board identities, and queue actions. Duplicate identities, missing receipts, inconsistent state, and actions on settled phases are findings. Missing stores are unmeasured. A clean comparison does not prove every project is current or that memory improves outcomes. The source is ZUP_HOME (default ~/.zen); no private records are published.

The Setup page also includes an index-and-memory operating review: live sources, scan errors, live embedding coverage and model mix, exact duplicate hashes, recorded retrieval use, and live keyword-search smoke probes. Expand each check for its rows, query and next action. Relevance, contradiction quality and causal benefit are explicitly unmeasured here; a working search is not a correct answer. For recorded results, helicon outcomes shows accepted, rework, rollback and missing acceptance decisions, with the recorded-run denominator. It does not grade all agent work. --json exports the helicon.outcomes/1 contract for ZUP. Use helicon outcomes --save-baseline /local/path/baseline.json to freeze a reading in a new file; later use --baseline /local/path/baseline.json to compare the same run IDs. New runs are excluded; missing old runs mark the comparison incomplete. Acceptance changes are not evidence of causal setup benefit.

Add --json for the helicon.setup-audit/1 contract consumed by ZUP. The audit reads instruction and skill files, reports their hashes, and checks inspectable Claude hook routes and post-compaction vault references. It neither executes hooks nor changes configuration. Reports contain local paths and hashes, not instruction bodies or configuration secrets; review paths before sharing.

Discovered files are candidates, not proof they reached a model. Effective context, plugin activation, cloud settings, and skill benefit remain unmeasured. The older helicon setup census remains available, but its limited file count no longer passes as a measurement of total loaded context.

One real catch, from a real transcript, in under a minute:

[NO-EVIDENCE ] L1127: "I ran the full pytest suite and all tests pass."
               → no tool call in this run could support it

Your agent said it. The trace doesn't back it. Now you know before you merge.

Your CLAUDE.md is lying to your agent, and nothing tells you.

It says "see docs/architecture.md" after that file moved. It routes to a directory the monorepo split apart. It warns about a rail that was retired six weeks ago. Your agent loads all of it as fact at the start of every session, acts on it, and neither of you finds out.

Mountain of Helicon runs the claims in those files against the repository in front of it, before the work starts. Every contradiction it reports carries the command and the stdout that proved it — so you are ruling on evidence, not on a model's opinion about your docs.

One line, no clone

Review any repo's agent setup in one command — no clone, no key, no LLM:

uvx --from mountain-of-helicon helicon-review ~/your-repo
# or:  pipx run --spec mountain-of-helicon helicon-review ~/your-repo

It reads that repo's CLAUDE.md / AGENTS.md / .cursorrules, checks every pointer, command, and version claim against the actual tree, and prints a graded verdict with the exact file:line of each broken reference. Exit code is non-zero when the setup lies to its agent, so it drops straight into CI.

Runs the existence + version tier by default. The execute-and-compare wedge — it runs a documented test/build command and grades the doc's "this passes" claim against the real exit code — is opt-in (HELICON_EXECUTE=1), because running a stranger's code should be a choice.

From a clone (the full lab)

git clone https://github.com/Morkeeth/mountain-of-helicon.git
cd mountain-of-helicon
python3 scripts/check_python.py     # not optional; see below
python3 -m pip install -e ".[web,model,retrieval]"   # or plain `-e .` for the review alone

helicon review ~/your-repo          # the front door: graded, evidence-backed review
helicon ci --path ~/your-repo       # your repo. no key, no config, no init.

That last line is the product. It reads the CLAUDE.md / AGENTS.md / .cursorrules / .clinerules / copilot-instructions your repository commits, probes each claim against your code, and prints every contradiction with the command and the stdout that proved it. Nothing is uploaded, no key is asked for, and no configuration file is written.

It is fast because there is nothing heavy in the path — no model download, no build step, no index to warm. The scan is git and subprocess calls, so on a cold clone it finishes in seconds; two different machines measured 3.5s and 2s for the scan itself. Time it on yours rather than trusting either number.

If you would rather watch it work on a known-convicted repository before pointing it at your own:

bash scripts/demo.sh                # a real gate firing, with evidence, no key

The demo scans named public repositories, shows contradictions with the probe and stdout that proved each one, installs the preflight hook into a throwaway settings file, and records an explicit override. It touches neither your real Claude settings nor any memory store.

Local-first. No API key for the deterministic checks. It warns by default and never blocks unless you opt in, because a preflight that wedges a terminal gets uninstalled. It can also audit longer-lived memory, let a human rule on contradictions, and compile those rulings into policy your agents query over CLI or MCP.

What it does not do: it settles what the filesystem can settle — a named path that is gone, a quoted command's output, a retired capability. It cannot tell you whether a sentence is true. Silence is not a clean bill of health; it means no executable probe could bind. helicon review does not read .clinerules yet, so paths and @imports in that file are not checked.

The measured finding

We ran the frozen 591-repository corpus and scored 577 current default branches; 14 exclusions are named. After removing projection duplicates, enforcing the existing-file evidence invariant, and hand-verifying all 30 mechanical survivors, 6 repositories (1.04%) contained a sendable doc-vs-code contradiction. Finding-level precision was 9/30. The rejected rows and reasons are part of the result, not a footnote.

The frozen corpus, commands, stdout, exclusions, and hand-verification ledger are in docs/agent-context-report-2026-08.md. Release gates and intentionally deferred work are tracked in LAUNCH_ROADMAP.md.

Build notes, night-run logs and review prompts live in docs/archive/ — kept, not deleted, just out of the front door.

Installing on a Mac

It took six cold clones before one ran clean. Not six errors found by reading the code — six real clones from GitHub into /tmp, installed into an empty HOME, driven with the stock macOS interpreter, following the block above exactly. Attempts 1, 3 and 5 each surfaced something the previous one had not: output that greeted a stranger by the author's name, pip install blaming the repo for a missing setup.py that was really a three-year-old pip, and a server that failed to bind and still printed open http://127.0.0.1:8420. The suite was green through all of it, because none of those are things a test suite is looking at. If you hit a seventh, that is a bug and worth an issue.

Run check_python.py first; it is not a formality. Python 3.10+ is required, and a stock Mac fails this two different ways, neither of which names the real cause. /usr/bin/python3 is 3.9 and dies on a PEP 604 annotation (TypeError: unsupported operand type(s) for |). Its bundled pip is 21.2.4, which predates PEP 660 and refuses with ERROR: File "setup.py" or "setup.cfg" not found — that one reads as this repo is missing a file rather than your pip is three years old. The preflight checks both and prints the exact command for each.

Bring your own Qwen key (BYOK). Get one free on the Alibaba Cloud Model Studio free tier, set QWEN_API_KEY or put it in ~/.helicon/config.json. helicon init keeps configuration and the SQLite store under ~/.helicon/, never inside the installed package. Keyless degrade: without a key every deterministic test still runs; only the two LLM-judged tests (Contradiction, Grounding) switch off -- the battery says so instead of faking a verdict.

What ships

  • Doorway: helicon sweep checks agent-rules files against repository reality; helicon doorway install adds a reversible Claude Code preflight.
  • Memory governance: a 13-class rot exam, human rulings, receipts, undo, and Golden Rules.
  • Agent access: local MCP exposes 25 tools, plus an authenticated remote endpoint with a narrower allowlist.
  • Connectors: Claude Code, Cursor and Cursor Cloud exports, git, Obsidian, agent rules, ChatGPT exports, Mem0, Letta, Graphiti, and LifeOS adapters.
  • Dashboard: Doorway, Rulings, governed runs, memory health, and the deeper Lab surfaces.

Visual demo

The web bundle is generated, never committed stale. From a source checkout:

helicon demo

On first run this installs/builds the dashboard with npm, seeds a labelled 19-memory demo under ~/.helicon/demo, and serves it only on http://127.0.0.1:8420/#findings. No personal connector runs and no API key is required.

Use it on your own repo

One command. No key, no config, no init. Run it inside a repository you own:

helicon ci

Use --path <dir> to check a repository you are not standing in. It reads that repo's committed CLAUDE.md / AGENTS.md / .cursorrules / .clinerules / copilot-instructions, probes each claim against the repository in front of it, and prints every contradiction with the command and the stdout that proved it.

  CONTRADICTED  CLAUDE.md:9   [path]
     claim   "Read `MASTER_GUIDE.md` for the operating model."
     probe   $ git ls-files -- MASTER_GUIDE.md
     output  (no output)
     why     the doc names MASTER_GUIDE.md; git tracks no such file and it is not on disk

If it prints nothing, no executable probe could bind. That is not a clean bill of health, and the section above says why.

Then, if you want the rest

The store, the dashboard, the rulings and the memory audit need a configuration file, so they start with helicon init. Running helicon scan before init will tell you the config is missing rather than do anything.

helicon init                       # writes ~/.helicon/config.json
helicon scan                       # extract memory from your configured sources
helicon doctor                     # PATH, config, key, DB, last scan
helicon audit
helicon check "what am I working on"
helicon serve

Qwen is optional and BYOK. Without a key, deterministic checks continue and the two LLM-judged checks report themselves unavailable rather than fabricating a verdict. Semantic embeddings are an optional install; the core remains slim.

CI for agent memory (GitHub Action)

The same helicon ci from the section above also runs in CI, so a pull request that drifts your agent's instruction files fails the build — CI for memory, literally. It scans a repo's committed CLAUDE.md / AGENTS.md / .cursorrules / .clinerules / copilot-instructions, runs 13 documented failure classes through the 13-class deterministic exam (no key, no torch, no LLM), emits GitHub annotations + a job-summary table, and exits non-zero on rot. R13 goes further than reading: it runs a probe against the repo's own running code and reports which sentences the system contradicts.

# .github/workflows/memory-ci.yml
name: memory-ci
on: [push, pull_request]
jobs:
  rot-exam:
    runs-on: ubuntu-latest
    steps:
      - uses: Morkeeth/mountain-of-helicon@main
        with:
          fail-on: rot   # or 'none' for report-only

Locally it's the same one command: helicon ci. This repo dogfoods the exam in report-only mode (--fail-on none) so known R6 findings remain visible without making unrelated pull requests permanently red. Teams that have ruled their baseline clean should use the Action's default fail-on: rot.

The live doorway (a warning backed by executable proof)

Everything above produces a verdict. The doorway puts it where work begins: a Claude Code UserPromptSubmit hook warns when the repository disproves loaded instructions, naming the offending lines and executable evidence in the terminal. Warning is the default because a preflight that wedges a terminal gets removed. Teams that explicitly want enforcement can set HELICON_GATE_MODE=block.

helicon board                 # every repo under ~/CODE and what it loads into an agent
helicon board --repo <name>   # every loaded line, with its probe verdict
helicon doorway install       # wire the gate into ~/.claude/settings.json (backup + diff + confirm)
helicon doorway install --uninstall   # remove exactly what it added
helicon hook --print-config   # the settings.json snippet (never auto-installed)
helicon receipt <session>     # did the harness actually RECEIVE the injection?

The gate a stranger installs is keyless and config-free: helicon doorway install writes one UserPromptSubmit hook (shown as a diff, backed up first, idempotent, and exactly reversible), and the hook — python3 -m helicon doorway gate — needs no config.json to run. On the next prompt in any repo whose loaded docs its own code disproves, the warning appears in your terminal and the run continues. Warnings and explicit overrides log into your configured store (so they show up in helicon runs / the dashboard), and fall back to a standalone ~/.helicon store for a stranger who has no config.

Three rules it obeys, all from the same law:

  • Machine-evidenced. A CONTRADICTED verdict came from a probe that executed and disagreed. The operator can correct/demote the line, continue after the warning, or explicitly record an override reason.
  • Cold lines never block. Demoting a line keeps it forever and loads nothing, so it cannot poison a run and must not stop one. --demote is a real exit, not advice.
  • Fail open, loudly. Any error lets the prompt through; a doorway that bricks a terminal gets uninstalled, and then it governs nothing. Logged warning/override events make intervention inspectable — the absence of one proves nothing.

helicon receipt is the honest half of delivery. Every other step can be satisfied by a row Helicon itself wrote; this one opens the transcript the harness wrote and looks for a content-derived token, ruling RECEIVED / NOT_FOUND / UNVERIFIABLE. UNVERIFIABLE is a verdict, never rounded up. Injections are checked against the ~32k context-rot onset first and trimmed if they would exceed it — a memory tool that quietly causes the rot it detects is the joke writing itself.

Headline Features

  • helicon snapshot -- regression tests for retrieved context. Capture what a task retrieves today; snapshot check fails when tomorrow's retrieval drifts. CI for memory.
  • helicon check "<task>" -- context-quality battery on what a task retrieves: Relevance, Freshness, Redundancy, Thinness, Expiry (deterministic) + Contradiction, Grounding (judged live by Qwen). Verdict: HEALTHY / DEGRADED / BROKEN. Every verdict prints the age of the last scan, because a DEGRADED verdict is uninterpretable if the scan itself is stale. --json for scripts and CI.
  • helicon reconcile -- timely forgetting. Re-scans sources and retires memories reality no longer contains (dry-run by default, never touches human decisions). On the live DB it retired 20 superseded memories in its first run.
  • helicon fix-skills -- write-back: Qwen writes missing descriptions into your agent skill files (dry-run by default, .bak backups). It fixed 7 of this project's own skills.
  • helicon doctor -- five checks (PATH, config, key, DB, last scan), exit 1 on failure. The front door to a daily loop.
  • helicon rule "<natural language>" -- prompted rules. Qwen compiles your sentence to a restricted predicate (whitelisted fields, never code); before approval you see coverage, samples, empirical precision against YOUR past decisions, and conflicts with other rules. One approved rule governs hundreds of items; applied rules are never counted as human evidence.
  • The regret ledger -- killed memories become a ghost list (LeCaR cache-eviction mechanics). When retrieval wants one back, a time-decayed regret event blames the exact decision that killed it, and FINDINGS shows "you retired this, retrieval wanted it 2x since -- restore?". Wrong forgetting is measured, not assumed.
  • helicon_flag over MCP -- point-of-use correction. Injected memories carry id + last_verified + used_count; the agent (or you, through it) flags stale/wrong/useful in one call. Flags become findings the human confirms -- the agent proposes, it never deletes.

Three Layers

Layer 1 -- Extraction. Pluggable connectors cover Claude Code, Cursor and Cursor Cloud exports, Obsidian, git history, ChatGPT exports, agent rules, LifeOS, Letta MemFS, Graphiti, and Mem0. Rewritten and expiring Mem0 memories carry their temporal fields into freshness tests. Agent rules files (CLAUDE.md, AGENTS.md, .cursorrules) are split into section-level memories so regression catches one section drifting. Every item becomes a HeliconCube: a versioned memory unit with source, confidence, content hash, review status, and decay parameters. A novelty gate prevents redundant storage.

Layer 2 -- Review pattern learning. Weibull forgetting curves with per-type shape (cliff decay for code, long tail for decisions). Auto-triage derives kill/approve rules from HUMAN reviews only -- its own decisions are excluded so it cannot reinforce its own echo. On its first run it handled 585 of the 1,268 memories the store held at that time autonomously. Spin detection, kill prediction, Helicon Score.

Layer 3 -- Meta-audit. The system audits its own stored patterns: temporal staleness ("this week" in a 27-day-old file), factual contradictions (Qwen-judged), decay, pattern staleness, anti-confabulation challenges. The human reviews the memory review.

Qwen Cloud API usage (where the LLM is load-bearing)

Tier Model Used for
fast qwen3.6-flash Memory summarization, novelty gate, skill descriptions
default qwen3.6-plus Battery judging (Contradiction, Grounding), factual audit, Next Moves
deep qwen3.7-max Consolidation synthesis, optimization reports
retrieval text-embedding-v4 Dense vectors (1024-dim) for hybrid + semantic search
retrieval qwen3-rerank Two-stage rerank over RRF-fused candidates

All calls go through the OpenAI-compatible endpoint https://dashscope-intl.aliyuncs.com/compatible-mode/v1 with a per-call SQLite response cache and per-operation cost tracking (/api/tokens). The two subjective battery tests are judged live and tagged (qwen) in output; if the judge call fails, the battery falls back to deterministic-only rather than fabricating a verdict.

MCP Server (25 tools)

Agents audit their own memory mid-conversation. Add to .claude.json:

{
  "mcpServers": {
    "helicon": { "command": "helicon", "args": ["mcp"], "cwd": "/path/to/mountain-of-helicon" }
  }
}
Tool Description
helicon_health Memory score and stats
helicon_stale Decayed memories below threshold
helicon_search Hybrid FTS5 + semantic search
helicon_contradictions Active factual conflicts
helicon_recent_reviews What the human approved/killed
helicon_patterns Learned behavioral patterns
helicon_guard Check a proposed claim against the compiled law before writing it: blocked / warn / clean
helicon_ask Guarded retrieve — what is safe to believe about a topic: the ruled-true answer + retrieved context split into safe vs. ruled-wrong
helicon_brief The morning brief — all five pillars in one call: truth, continuity, direction, reflection, calm
helicon_portrait Grounded portrait of what the record shows about the person, plus its health
helicon_context Proactive memory injection for a task -- every memory carries its id, last_verified, used_count
helicon_flag Point-of-use correction: flag a memory stale/wrong/useful by id; stale/wrong become findings the human confirms
helicon_playbook Task playbooks from review patterns
helicon_compile Compile reviewed memory to injectable files
helicon_triage Trigger auto-triage
helicon_prompt_gate Gate an execution prompt through a Wager -- approves only after a human accepted a BUILD or REPAIR move, else abstains
helicon_capture_launch Freeze the acceptance test and context packet before implementation starts
helicon_capture_closeout Close a run with real artifacts and a real verification receipt
helicon_workgraph_trace Join one work card to its task run, context, memory, skills and evidence
helicon_workgraph_attention Name the missing graph edge -- link_run, freeze_context, attach_artifact, choose_move
helicon_workgraph_learning Withhold recommendations until real resolved outcomes accumulate
helicon_workgraph_review_skill Record the skill version actually loaded, hashed over its bytes
helicon_consolidate Run a consolidation (sleep) cycle
helicon_context_packet_inspect Inspect a named run's local source-packet receipt; does not deliver content
helicon_context_packet_consume Return the exact reviewed project sources to the named local run and record receipt, not compliance

The full JSON-RPC 2.0 handshake (initialize, tools/list, tools/call) is exercised in the receipts; helicon mcp runs the server on stdio, so the bare CLI never silently becomes a server.

Remote MCP for cloud agents

helicon serve also exposes a stateless MCP endpoint at /mcp when a dedicated token is configured:

export HELICON_MCP_TOKEN="$(python3 -c 'import secrets; print(secrets.token_urlsafe(32))')"
HELICON_CONFIG=/path/to/config.json helicon serve

Configure the remote client with https://your-helicon-host/mcp and send Authorization: Bearer <token>. The endpoint refuses to start with a token shorter than 32 characters, rejects requests over 1 MiB, serializes access to the SQLite connection, and returns Cache-Control: no-store.

Remote access deliberately exposes the agent workflow, including context, guard, ask, and point-of-use flags, but not helicon_compile, helicon_triage, or helicon_consolidate. Those maintenance tools can write host files or mutate the store in bulk and remain available only through the local stdio transport.

The built-in server does not terminate TLS. Put it behind HTTPS on a private network or an authenticated reverse proxy; never send the bearer token over plain HTTP and never expose a personal memory store directly on a public IP. HELICON_PASSWORD protects dashboard API routes and is intentionally separate from HELICON_MCP_TOKEN.

Feed Cursor Cloud runs back into memory

The cursor-cloud connector reads local export bundles containing index.json and per-run transcript.json, diff-metadata.json, and events.json files:

{
  "connectors": {
    "cursor-cloud": {
      "enabled": true,
      "export_dir": "~/Downloads/cursor-cloud-agent-transcripts",
      "include_text": false
    }
  }
}

It selects the newest export for each stable cloud-agent id and produces one idempotent session summary: repository, branch, model, status, message/tool counts, observed tool failures, events, and diff/PR outcome. Metadata-only is the default because raw exports may contain private prompts, reasoning, terminal commands, file contents, diffs, and credentials.

Set include_text to true only when conversation prose belongs in the memory store. That opt-in includes bounded user and final-assistant text with common token patterns redacted. Reasoning, tool arguments, terminal output, file contents, search results, and diffs are never ingested.

CLI (79 commands)

init scan reconcile fix-skills serve demo triage review fix route score-runs runs run hook receipt judge-bench bench attribute move leaderboard snapshot lens taste check checkin report read audit consistency registry checkouts volatility unreviewed fleet queue guard ask teach brief board sweep doorway repair ci policy evolve wager capture lift resolve watch alias rule doctor export mcp score stack setup outcomes witness skills-review optimize eval embed playbooks reflect compile consolidate eval-consolidation complaints overboard ledger measure magnet measurement-bench review-queue science truth

helicon fix prints safe path rewrites as a dry run and writes only with --apply.

Four of them answer to a second name, kept working so older muscle memory doesn't break: battery = check, rot = audit, heal = repair, gold = policy. Aliases, not extra commands, so they are not counted above.

helicon truth is the stranger-facing cold path: point it at any agent memory directory (Claude Code, Cursor, Cline, Obsidian vault) and get a ranked, evidence-cited staleness+rot report — no Helicon DB, no API key, no LLM. Deterministic; every row cites the line it fired on.

helicon route turns output-verification into a model-routing recommendation: it reads the eval store — the verified verdicts review-queue --terminals produced — and ranks models by Wilson-scored verified-pass-rate per task-class, with sample size and confidence attached. The model is attributed from the git co-author trailer of the commits that produced the output; the outcome is a real reality-check, never a guess. Below a sample threshold it says insufficient evidence, never a fabricated number. helicon route --record --run builds the evidence first. See docs/ROUTE.md.

helicon score-runs and helicon runs score whole RUNS, the same output-verification one level up and made cost-aware: score = verified yield / cost - damage. Cost comes from the real transcript token usage, yield from the review-queue --terminals verdicts, damage from an incident flag. Every term traces to a real source; nothing is vibed. score-runs --card cuts one run card, runs renders the scored history, runs --suggest reads what to run next off it. See docs/RUNS.md.

helicon judge-bench benchmarks Qwen as the memory-rot judge against the operator's own human rulings, and (with an OpenRouter key) against GPT/Claude on the same probes. Real result (run #2, 26 probes, 13 real contradictions + 13 controls, ground truth = the operator's own rulings): qwen3.6-plus ties anthropic/claude-sonnet-5 at 0.962 accuracy and beats openai/gpt-5 (0.808), at $0.00444 per run and in 54s against Sonnet's 144s and GPT's 245s. qwen3.6-flash holds 0.923 for $0.00167, 8x cheaper than qwen3.7-max for the same specificity. And every model, Qwen and Claude and GPT alike, missed unit-drift: some rot is a domain ruling rather than a logical contradiction, and no judge at any price catches it. That is what the human-ruling layer exists for. Reproduce: helicon judge-bench --set all --save (needs openrouter_api_key for the competitors); the Judge tab reads the saved run, and renders an unrun bench as unrun. See docs/ROUTE.md.

helicon move is the context-mover: read memory from one platform, VERIFY each item (freshness, and with --verify-contradictions the Qwen judge), and render the survivors into another platform's native format (claude-code / cursor / markdown). Memory moves verified, never blindly; held-back items are listed with why. Dry-run by default; --apply backs up the target first.

helicon leaderboard is the population-scale version: it reads git history across many repos (where multiple models and harnesses actually co-authored commits) and ranks models by how often their commits SURVIVE vs get REVERTED, Wilson-scored. Execution-free (git only, so it is bounded and cannot freeze a machine); the revert is the honest failure signal. On 927 attributed commits across 25 local repos it already separates opus-4.6 / opus-4.8 / fable-5 / cursor by reliability.

helicon repair runs the self-healing audit loop — the thing no retriever can do. It scores the four truth gates (freshness / volatility / consistency / retrieval) on a store, surfaces each drift with its cross-source evidence, proposes a repair (retire the stale memory, move a fast fact to the live layer) as a diff you accept, applies the accepted ones, and re-scores so the gates visibly move. helicon repair --demo runs it on a seeded, universally-legible store (the classic "I told my agent I'm vegetarian, then started eating chicken again — it never updated" contradiction, plus a stale goal and a fast fact); --apply closes the loop.

helicon audit runs the rot exam: the 13 documented memory-failure classes in ROT.md checked live against your real store -- deterministic, zero LLM calls, free to run daily. On this repo's own store it currently finds rot in several of 13 classes and says so — and as of Jul 5 all classes are fully tested, 0 partial.

helicon watch makes the exam ambient: scan + selectors + rot exam on a timer (helicon watch --install writes the crontab line, every 6h), diffed against the last run. You get a macOS notification and a drift-report.md only when something NEW rots — no news, no noise. First run baselines silently.

helicon policy compiles GOLDEN RULES: the stack's law, built from your rulings, dismissal precedents, approved triage rules, declared renames, canonical sources and standing feedback — every rule with its provenance (a rule without provenance is a vibe). --inject writes it to ~/.claude/GOLDEN_RULES.md (dry-run default, .bak kept) so every session can obey it. helicon evolve is the night command: scan, every selector, the exam, a gold recompile, and the morning delta — what your stack learned while you slept.

helicon checkin asks one harness for the live shared-state revision, exactly five open task IDs, and one next action. It scores the answer against CURRENT.md on a fixed 100-point denominator, reports Codex memories as review candidates rather than silently merging them, and reads the existing Transcripto views in read-only mode.

helicon report prints a MemoryAgent Compliance Report: the track's four sub-goals (efficient storage/retrieval, timely forgetting, recall under limited context windows, cross-session accuracy) scored live from your real memory, thresholds printed with the numbers. Any memory stack a connector can scan could be graded by the same exam.

helicon fleet is one screen for the whole fleet, and it is project-first because that is the unit a person thinks in. Every field is DERIVED and none is typed: where a project stands comes from git (branch, HEAD, uncommitted, unpushed, commits in 24h), needs you from runs that finished with no verdict, friction from the complaint log, next only from a prompt you actually accepted and that matches this project — otherwise it says unmeasured and names what would fill it. A hand-typed "next step" field is stale in a week, and then you have rebuilt the dashboard problem one layer down.

It opens with IDLE, because that is the only line that costs money while you read it: terminal-hours sitting silent right now. Reported as a floor, not a total — a session thinking hard writes nothing and reads as idle — and counted only for sessions a human actually typed into. The first version counted every transcript and claimed 241 idle hours across 74 "terminals"; 68 were headless judge processes that had already exited.

It ends with YOU CAN, which is the section that is not about state. On the day this was built, eight terminals sat idle for two hours while prompts were hand-written for a human to carry between them, because no agent instance remembered that ListAgents and SendMessage address every peer session directly. A capability nobody recalls is not a capability, so the screen states it: an agent should leave this screen holding an ability it walked in without.

helicon complaints is the complaint log: every time you pushed back on an agent, recovered from your own transcripts and grouped by kind. Every other signal in this repo is the machine grading itself — battery verdicts, judge runs, self-scored retrieval. A correction is a human saying no, that is wrong with nothing to gain by lying, and it is destroyed the moment the terminal closes.

It rests on two authorship gates, both measured before they were written. A type: user entry is not necessarily the human: across ~600 local transcripts, 46% of non-tool user turns were programmatic judge runs, task notifications, messages from other agent sessions, or injected skill files. Count those as feedback and an agent is grading itself on its own prose. And even a turn you typed is not always your writing — pasted lane prompts are long and any correction-shaped phrase inside them sits deep in the body, so position separates them. Raw user turns score 28% precision; both gates ~81% (self-graded, the author read all 46). Storage is a cube of type='complaint', so it inherits FTS, decay and the review loop — no new table.

On this repo's own store the largest category is not a wrong fact. It is wrong-plan: the agent choosing the wrong next thing to do.

Audit a store you don't own

The exam is not limited to your own memory. Any repo with a committed agent-rules file (AGENTS.md, CLAUDE.md, .cursorrules, ...) is a memory store someone's agent obeys every session — so it can be examined:

bash scripts/demo_public_store.sh          # default: openai/codex AGENTS.md

This replays the file across its REAL git history (no staging): ingests an old commit, snapshots retrieval, replays to HEAD, reconciles, runs the rot exam. On openai/codex (27 real commits of AGENTS.md edits, cited by SHA in the output): 5 sections retired as drifted, 1/1 retrieval snapshot regressed — R10 and R8, live, on a store we don't own. Reproducible by anyone.

The life-OS benchmark — scored against human-labeled rot

On Jul 5 a 5-agent manual audit swept the operator's real second brain (Obsidian vault + Claude Code memory dir), archived 33 stale docs and stamped 21 drifting docs with dated > **LOUPE correction banners. Those banners are a labeled dataset of real memory rot. The benchmark ingests the same corpora with the banners stripped (the answer key never enters the input) and scores the deterministic detectors against them:

python3 scripts/rot_bench_lifeos.py    # read-only on sources, throwaway DB, zero LLM

Honest numbers from the first run (232 files, 1,667 section memories): 6/16 file-level catches, 4/16 strict facet-match — the output labels the difference itself. What it caught: both merge-status flips (audit doc still said 'NOT patched' after the fix merged), a stale dashboard doc, a dead 7-week-old plan. What it found that the humans missed: a win-count fight (9 vs 10) living in the resume and two application drafts, and 35 files still asserting a dead project name post-rebrand. Named misses, on the roadmap: overlapping-date-range drift (Aug 14-22 vs Aug 15-24 overlap, so interval semantics reads agreement), living-doc supersession without a declared rename, and content-based staleness (a young file asserting old facts).

Access & trust model (read this before connecting your vault)

A tool that audits your memory reads your memory. That access is scary, so here is exactly what Mountain of Helicon does with it — from the code, not a promise:

Reads (always read-only): your configured sources — Claude Code transcripts, Obsidian vault, git repos, rules files, memory stores via adapters. Connectors never write to a source. The life-OS benchmark and the rot exam open the store read-only.

Writes, exhaustively:

  • its own SQLite DB and data/ (findings, verdicts, drift reports, compiled context)
  • helicon fix-skills and other write-backs: dry-run by default, --apply required, .bak written next to every file before modification, second run is a no-op
  • helicon watch --install: one tagged line in your crontab, removed by --uninstall
  • helicon compile: compiled context files under data/compiled/ (--output redirects). It writes nothing into ~/.claude/: an auto-inject path exists in the source (compiler.inject_into_claude_code) but no command calls it, so the pull path (helicon_context over MCP) is the working half of that loop and the push half is unwired. Our own store still carries pre-rename glaze-* skill files from an older injector, which the skills audit flags: a live example of why write-backs need lifecycle discipline
  • helicon policy --inject (alias helicon gold --inject): ~/.claude/GOLDEN_RULES.md, dry-run by default, .bak kept
  • your vault: never. Corrections are memories in Helicon's store, not edits to your files. You stay the only writer of your second brain.

Leaves your machine: nothing, unless you configure a Qwen key — then excerpts of candidate memories (truncated content) go to the model for judging, and the response is cached locally. Keyless mode runs every deterministic check with zero egress and says so instead of degrading silently.

Decisions: every destructive or state-changing action (kill, retire, resolve, dismiss, rule application) is either made by you or made by a written rule you previewed and approved — and automated decisions are quarantined from the learning loop (rot class R9), so the tool cannot launder its own output into your evidence.

Your domain, your lexicon (config, not code)

The claim-conflict detectors ship with built-ins (win counts, episode numbers, merge status, decision status) and take the rest from config.json — an enterprise wiki or research vault declares its own counted things and polar statuses, and gets the same conflict machinery, evidence receipts and resolve loop:

"claims": {
  "metrics":   {"headcount": "\\b(\\d{2,5})\\s+employees\\b"},
  "statuses":  {"contract": {"live": "contract (is )?live", "expired": "contract (is )?expired"}},
  "canonical": {"wins": "mindmap.md"}
}

canonical encodes the single-source-of-truth rule: declare WHERE a fact's truth lives, and a conflict files as "Drift from canon: canon says 9; 8, 10 asserted elsewhere" — the human confirms a pre-decided direction instead of adjudicating from scratch.

Doc honesty is enforced: python3 -m helicon.docdrift compares this README's numeric claims against counts computed from source, and it runs in the test suite — stale docs fail the build. (It caught this very README claiming 20 commands the hour the 21st landed.)

Everything destructive is dry-run by default and takes --apply.

Honest eval numbers

None of the four numbers below reproduces on a fresh clone, and that is a gap, not a footnote. They were computed against the author's own populated store — roughly 6,900 memories carrying real human kill/approve rulings — which is personal data and is not in this repository. On a clean install helicon eval has no store to read and says so rather than printing anything. Treat these as measured once, on a corpus you cannot see, which is weaker evidence than the doorway numbers above, every one of which you can reproduce on your own repo in seconds.

  • Composite: ~67 (live, as of 2026-07-13 — run helicon eval to recompute; retrieval P@3 + MRR + decay-AUC; audit axis excluded -- no labeled ground truth).
  • Retrieval: P@3 0.615, MRR 0.596. Small internal benchmark (n=13, one label per query) -- disclosed, not hidden.
  • Decay predicts human kills at rank-AUC 0.78 (mean confidence of killed memories 0.14 vs approved 0.27). A real, independent signal.
  • Consolidation: ~9-10x fewer tokens; Qwen-judged quality favors synthesis (self-graded, shown as direction, not proof).
  • The public demo store is 19 labelled planted memories and contains no personal data. Live scans read only the sources each user configures.

Built on established patterns, extended

Mountain of Helicon's capabilities stand on well-understood memory-systems patterns and take each one further. The lineage, stated honestly — the second column is the established idea, the third is our own build on top of it:

Capability Established pattern How Mountain of Helicon extends it
Versioned memory units Structured memory units, not raw text HeliconCube: source, hash, valid_from, confidence, decay per type
Multi-axis audit Temporal/factual/logical consistency checks 13-class rot exam, each with a receipt and a never-twice guard
Weibull decay Non-uniform forgetting curves Per-type kappa, and decay rank-predicts human kills (AUC 0.78)
Novelty gate ADD/NOOP/MERGE at ingestion Gate + provenance, so a merge never loses the source it came from
Anti-confabulation Challenge claims against evidence Grounding check + R12 phantom-association catch
Retrieval learning Track surfaced vs acted-on Q-value ranking rewarded by human rulings only — no self-echo
Identity & phantom coherence (ours — no store or prior system does this) R11 fork detection, R12 phantom catch, rulings compiled to law

Architecture

Mountain of Helicon architecture: the store, retrieve, output, attribute, rule, law loop

  • Backend: Python 3.12, FastAPI (154 endpoints), SQLite + FTS5 (43 tables declared across the memory and correction stores). Qwen-native retrieval when a Model Studio key is configured: text-embedding-v4 (1024-dim) dense vectors + FTS5, fused by Reciprocal Rank Fusion, then a qwen3-rerank two-stage pass, the whole retrieve→rerank stack on Alibaba Cloud (falls back to local MiniLM + linear fusion, FTS-only, when no key)
  • Frontend (optional): React 19, TypeScript, Vite. Four surfaces — Next Moves (memory state → cited next prompts/goals, generated by Qwen, every move citing the memory it came from), Memory (sources, review coverage, health), Needs Ruling (every failed check with why/evidence/action, grouped Drift / Stale / Smartness), Golden Rules (rulings compiled with provenance, injectable). The dashboard is one of three interfaces (CLI · MCP-in-IDE · dashboard)
  • AI: Qwen Cloud API via OpenAI-compatible SDK (see table above)
  • Distribution: BYOK + local-first. No hosted personal-store service is advertised for v0.1; public hosting waits for HTTPS, sessions, configured CORS, rate limits, and backups.

License

MIT

Metadata

Release files for mountain-of-helicon 0.2.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mountain-of-helicon 0.2.4
File Size Uploaded
mountain_of_helicon-0.2.4.tar.gz 1.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for mountain-of-helicon 0.2.4
File Interpreter ABI Platform
mountain_of_helicon-0.2.4-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / mountain_of_helicon-0.2.4.tar.gz

Download URL mountain_of_helicon-0.2.4.tar.gz
Size 1.0 MB
Tags Source
SHA-256 checksum
How to use checksums
5d706f280dca52d1ccc300d2667f082fb265d1f0e7db0fd592f139f71c59da20
BLAKE2b-256 checksum
How to use checksums
d49c2919ea6a793ac5c636565f54af574f0d5809dd508d7d70aa82caa54e74a0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.5

Release files / mountain_of_helicon-0.2.4-py3-none-any.whl

Download URL mountain_of_helicon-0.2.4-py3-none-any.whl
Size 767.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
eaddf400229183cede161519ddc060039375b9332a3182545dc28bc89d311ced
BLAKE2b-256 checksum
How to use checksums
9adc9cc8eef84d777ec41d8275e9bc77036ed53aaea133bdc3b7dd16c05799b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.5

Release history Release notifications | RSS feed

This release

0.2.4 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page