Latent Compass
A flight recorder and offline proving ground for strategic decisions made by coding agents.
Release history: CHANGELOG.md.
Coding agents make consequential choices all day: which hypothesis to test, which file to change, which failure to pursue, when to stop, and when to ask for help. Most systems preserve the final patch. They do not preserve the decision that produced it.
That creates a dangerous blind spot. A successful task does not prove that the chosen strategy was sound. A failed task rarely tells us whether another available direction would have worked. If an agent later “learns” from those outcomes without strict evidence and authority boundaries, it can end up grading its own work, rewriting its own history, and promoting its own policy.
Latent Compass was created to explore a harder question:
Can an agent learn from strategic decisions without gaining the authority to declare those decisions correct?
The project starts with evidence, not intelligence. It captures decision episodes, validates their contracts, records them in tamper-evident local stores, and evaluates fixed policies offline. A deterministic external judge remains authoritative. A human remains the only actor allowed to promote.
The stable interface validates and records. It does not steer the agent and emits no operational advisory.
An isolated active-diagnosis lab now computes which
observation to acquire next from an explicit finite model, budget and horizon.
It includes bounded source reads,
host-executed observations it admits and re-checks,
justification memory, and
routing, evaluation and retirement dry-runs.
Run uv run python examples/lab_source_demo.py for a synthetic end-to-end example.
This experimental API has no execution authority. An optional passive hook
adapter can record non-authoritative routing advice from ordinary Codex and
Claude tool events without storing prompts, arguments or results and without
creating another model call; it never changes the route the host takes.
Removing an index still requires a real comparison, pilot and verified migration;
local correctness tests do not establish those outcomes. See ADR 0011
and the operational scope, which names the one
pilot perimeter, the read-only inventory and every open unknown.
The problem it addresses
Imagine an agent facing three plausible directions:
- patch the visible symptom;
- inspect the contract that produced it;
- stop and request a missing product decision.
The agent chooses one. Hours later, the task is green or broken. What is usually missing?
- the alternatives available before the choice;
- the evidence attached to each alternative;
- the logging propensity of the selected direction;
- the cost, violations, information gain, and reversibility observed later;
- a trustworthy record of what was known before the outcome;
- a separate authority capable of saying “continue” or “kill.”
Without those pieces, “learning from experience” is mostly storytelling after the fact. Latent Compass exists to make that story falsifiable.
Why it was built this way
The first design question was not “Which model should we train?” It was “What must a learning loop never be allowed to do?”
That led to five non-negotiable boundaries:
- Capture before outcome. Candidate evidence belongs to the moment before a direction is selected, not to a retrospective explanation.
- Refuse ambiguity. Missing versions, malformed probabilities, invented fields, and unsupported evidence fail closed.
- Separate observation from authority. The component that records a verdict cannot grant itself permission to act on that verdict.
- Keep holdout evidence scarce. Validation can be replayed; claim-bearing holdout evidence cannot be spent repeatedly until it says what we want.
- Make “not enough evidence” a valid result. Abstention and
KILL_DISCOVERYare safer than manufacturing a winner.
The name reflects that role: a compass can describe direction without taking the helm.
How it works
Agent reaches a decision point
│
▼
Latent Compass validates the episode and candidate evidence
│
▼
Separate memory and reconciliation stores preserve pre- and post-action facts
│
▼
Offline benchmark and prospective diagnostics inspect sealed recorded data
│
▼
External judge evaluates evidence ─── Human retains promotion authority
Latent Compass never calls the external judge, never writes back to it, and never turns a benchmark or collection result into an authorization.
What exists today
Demonstrated — implemented and covered by repository tests
- Strict, versioned episode contracts. Unknown fields, type coercions, non-finite values, non-canonical timestamps, invalid propensity distributions, and overstated observability are refused.
- A fail-closed authority boundary. Even a locally consistent positive
result is refused as
untrusted_evidencewithout an external trust root. - Host-bound append-only stores. Atomic append, replay, structural redaction, integrity verification, and durable anchors make silent mutations detectable within the documented threat model.
- A reproducible offline benchmark. Four pre-registered baseline policies run under the same budget on sealed validation data. Verification re-executes them instead of trusting a supplied checksum.
- Holdout discipline. The validation runner and pre-holdout planner do not open the holdout split. Consumption is atomic and keyed to the holdout corpus.
- Judgeable pre-action capture. Candidate-specific evidence is captured as an immutable sidecar without leaking the selected action or outcome.
- Opt-in strategic decision memory. A separate host-and-family-bound store
accepts only explicit
STRATEGIC_HIGH_IMPACT,NON_SENSITIVErecords. It supports revisions, revocation, expiry, tombstones, and sealed read-only transfer while never selecting, scoring, executing, or authorizing a route. - Post-action reconciliation. A separate journal binds observations to the exact pre-action revision. Authorization and execution remain distinct; missing values stay explicit unknowns; replay produces no score, ranking, causal claim, or authority.
- Preregistered prospective shadow collection. A sealed plan fixes source identity, population, strata, calendar, stop rules, producer separation, and exact-binomial assumptions before enrollment. Its terminal report is a missingness and inclusion diagnostic, not a policy evaluation.
- Confinement. Durable writes stay beneath an explicit root, refuse silent overwrite, and do not follow replacement links.
Experimental — real, but not yet evidence of value
- The eight-family metric taxonomy and its thresholds.
- IPS with a support floor, SNIPS diagnostics, and the weighted empirical tail cost quantile introduced in benchmark contract 1.1.0.
- The synthetic corpus and walkthrough. They prove that the pipeline executes and refuses bad inputs; they are not evidence that any policy is good.
- Tombstone redaction and local durable anchors.
- An externally evidenced keyless labeler. Its private workflow is bound by GitHub OIDC and Sigstore, but its conservative rubric produced abstention, not useful directional supervision.
Projected — not implemented and not claimed
- A trained pairwise ranker or contextual bandit. That was HOK-182's intended lane; the issue was canceled before training because its data gate did not pass.
- Calibration for a live coding agent.
- Canary evaluation or integration with Semctx.
- Automatic promotion or execution authority.
- Any demonstrated improvement in agent outcomes. No such claim is made.
The first research epoch ended in KILL_DISCOVERY. No ranker was trained. The
project lacked enough eligible directional supervision, calibrated outcomes,
and claim-bearing holdout evidence to authorize the next step. The refusal is
part of the result.
HOK-253 remains a real-data gate. A responsible owner must supply the target
population, source bindings, strata, null and alternative rates, alpha, target
power, clustering inflation, calendar, exclusions, and producer identities.
The package has no statistical defaults and schedules no collection. Synthetic
values in examples/ cannot unlock training, canary execution, activation, or
promotion.
A concrete example
Suppose an agent is debugging an authorization failure. Before it acts, a producer records candidate evidence. The memory stores that exact pre-action projection. Later, reconciliation records the observed result for the one executed direction while leaving every unavailable dimension explicitly unknown. A prospective journal may include the case only when it matches a plan sealed before collection began.
This preserves what was known, what happened, and what remains unknown. It does not identify the best counterfactual, establish causality, or authorize action.
Quick start
Requires Python 3.13 and uv.
This source line is the v0.3.0 candidate. It adds the latent-compass host
installation workflow while keeping all observations local, passive and
non-authoritative. The public v0.2.0 GitHub release remains a historical
artifact at its original commit; the new PyPI candidate does not reuse that
version. The GitHub release page and PyPI registry remain the authorities for
whether v0.3.0 has actually been published.
The public v0.1.0 assets are historical artifacts pinned by SHA-256. An
independent audit found that they were built from a Windows working tree rather
than uploaded by CI, so they are not claimed as reproducible or
provenance-attested builds. Future tags use the attested release workflow and
test the extracted source distribution before publication.
git clone https://github.com/hoklims/latent-compass
cd latent-compass
uv sync --all-groups
Install this checkout as an isolated, persistent tool and preview the Codex integration before changing host configuration:
uv tool install .
latent-compass host install --host codex --project-root . \
--project-alias latent-compass --dry-run --json
latent-compass host install --host codex --project-root . \
--project-alias latent-compass --json
latent-compass host status --host codex --project-root . --json
Repeat --host to target both hosts. Every selected host is preflighted before
any file is changed. A malformed host file or a project alias/root collision
refuses the whole operation. Repeating an installation is idempotent. Remove one
registration without disturbing the others with latent-compass host remove --host codex --project-alias latent-compass --dry-run --json, then repeat
without --dry-run after reviewing the plan. The legacy python -m latent_compass.shadow_install ... entry point remains available.
If install or remove was interrupted, both commands refuse with
recovery_required and include the exact recovery command in their JSON.
Preview and apply the repair separately, then rerun the original setup:
latent-compass host recover --dry-run --json
latent-compass host recover --json
Recovery restores only transaction-owned bytes. A concurrent third-party
change returns pending_transaction_conflict; its file and backup are left
untouched for inspection.
The release workflow is configured for PyPI Trusted Publishing. The package
name must first be bound to this repository, workflow and pypi environment in
PyPI. Do not replace . with the registry package name in installation guidance
until a tagged artifact has been published and installed successfully from
PyPI.
Run the self-contained synthetic walkthrough in a new output root:
uv run python examples/walkthrough.py --root ./synthetic-run
The walkthrough validates the shipped episode and projection, derives a blinded pair, writes decision memory and reconciliation stores, and closes one synthetic prospective collection. It refuses an existing root. Every result is labeled synthetic and carries no empirical or authority meaning.
Validate the complete repository gate:
uv run ruff format --check . && uv run ruff check . && uv run mypy && uv run pytest -q && uv build
Each step is independently runnable. pytest proves behavioral contracts;
Ruff and mypy do not.
See whether Latent Compass is being used
The passive adapter stays silent during normal work. The installed status command makes that behavior visible without displaying or retaining prompts, tool arguments or tool results:
latent-compass-status --project-root .
latent-compass-status --project-root . --json
latent-compass host status --project-root . --json
For each Codex or Claude host it reports whether all three hooks are present,
whether the current project is registered, how many privacy-minimised events
and sessions were observed, the latest observation time, and the counts of
ADVICE and ABSTAIN records. OBSERVING means records exist; it does not mean
the host followed the advice. Hook trust remains UNKNOWN until reviewed in the
host itself. On Windows, an existing event directory is reported as
OBSERVATION_UNKNOWN with null derived metrics because this release does not
enumerate it without a handle-bound directory API. The footer always restates that Latent Compass has no execution
authority, did not influence host routing and recorded no content.
Independent proof status
The current status remains PROOF_WEAK/BLOCK. The first independent audit
verified six high-impact invariants with red/green mutants and found no foreign
release payload, but it also found material release and proof-contract defects.
The public independent-audit protocol now replaces
the unreachable private gate. Schema v4 defaults to strict separate-account
review; operators may explicitly choose isolated-session for an independent
review on the same account, with observed isolation and stated limits. A fresh
receipt bound to the exact candidate is required. Changes to the proof mechanism
must also be admitted by the unchanged external N-1 evaluator; the candidate gate
cannot approve itself. This does not authorize a pilot, hook execution, or index
removal.
Record and replay an episode
uv run latent-compass init --root ./store --store-id store-alpha \
--host-id host-alpha --agent-family claude --epoch LC-2026-E1
uv run latent-compass validate --episode examples/synthetic-episode.json
uv run latent-compass append --root ./store --episode examples/synthetic-episode.json
uv run latent-compass verify --root ./store
uv run latent-compass replay --root ./store
uv run latent-compass export --root ./store --out ./store/snapshot.json
Capture a judgeable pre-action projection:
uv run latent-compass pairwise capture --projection examples/synthetic-projection.json \
--root ./run --out ./run/captures/synthetic-decision-0001.json
--out must remain inside --root. Omit it to write only to stdout. Exit codes
are part of the contract: 0 success, 2 usage, 3 refusal, 4 integrity
failure, and 5 store or filesystem failure.
Run the offline benchmark
The committed corpus is synthetic demonstration data. The validation runner uses it as a measuring instrument; the policies never reach a live agent.
uv run latent-compass benchmark manifest verify \
--corpus-dir ./corpus/synthetic-v1 --manifest ./corpus/synthetic-v1/manifest.json
uv run latent-compass benchmark run \
--spec ./corpus/synthetic-v1/spec.json --protocol ./corpus/synthetic-v1/protocol.json \
--manifest ./corpus/synthetic-v1/manifest.json --corpus-dir ./corpus/synthetic-v1 \
--root ./run --out ./run/report.json
uv run latent-compass benchmark verify --report ./run/report.json \
--spec ./corpus/synthetic-v1/spec.json --protocol ./corpus/synthetic-v1/protocol.json \
--manifest ./corpus/synthetic-v1/manifest.json --corpus-dir ./corpus/synthetic-v1
uv run latent-compass benchmark holdout plan --report ./run/report.json \
--spec ./corpus/synthetic-v1/spec.json --protocol ./corpus/synthetic-v1/protocol.json \
--manifest ./corpus/synthetic-v1/manifest.json --corpus-dir ./corpus/synthetic-v1 \
--baseline least-uncertainty --root ./run --out ./run/holdout-plan.json
The benchmark runner, report verifier, and phase-one planner never read the holdout split. Manifest authoring and verification do read all three split files to derive and check their seals.
Architecture
episode.py versioned decision episode contract
pairwise_capture.py bounded pre-action projections and blinded pairs
decision_memory/ durable pre-action strategic records
decision_reconciliation/ durable post-action observations and replay
prospective_collection/ sealed enrollment journal and diagnostics
benchmark/ sealed offline baseline benchmark
authority.py refusals, transitions, reproduced evidence
ledger.py append-only episode store, chain, anchor, replay
cli.py bounded command-line validation and recording
canonical.py / contracts.py canonical serialization and strict primitives
Benchmark results, reconciliations, and prospective diagnostics are evidence. None is an authorization.
What the evidence proves — and what it cannot
Latent Compass distinguishes consistency, integrity, authenticity, and authority because they are different claims.
- A valid contract proves that an input has the expected shape.
- A recomputed seal proves that content matches a known preimage.
- A hash chain and anchor detect tampering within their threat model.
- A Sigstore bundle identifies a workflow and immutable input.
- None of those facts proves that an outcome is true, that evidence predates a decision, or that an actor is authorized to promote.
An administrator who can rewrite both a store and its local anchor can forge a history that verifies. Detecting that attack requires an external anchor this package does not have. The limitation is documented and tested.
Documentation map
| Document | Purpose |
|---|---|
| Authority boundary | actors, capabilities, lifecycle, provenance, threat model |
| Episode contract | the decision episode schema |
| Evaluation protocol | pre-registration and holdout discipline |
| Benchmark protocol | baselines, estimators, metrics, and limits |
| Pairwise supervision | labels, outcomes, calibration, and the data gate |
| Judgeable projection | pre-action sidecars and blinded pair derivation |
| Independent labeler | external workflow identity, proof, rotation, and limits |
| Pairwise corpus readiness | current refusal and exit criteria |
| Decision memory | opt-in pre-action records, transfer, and trust limits |
| Decision reconciliation | post-action observations, unknowns, and replay |
| Prospective shadow collection | preregistration, enrollment, diagnostics, and real-data gates |
| Corpus provenance | how the synthetic corpus was produced and what it cannot show |
| Ledger | chain, anchor, replay, redaction, and honest limits |
| Architecture decisions | why the project chose its current boundaries |
| Governance | data and project governance |
| Contributing | contribution rules and the required gate |
| Security | vulnerability reporting |
Licence
Apache-2.0. See LICENSE and NOTICE.
The two workflow badges above are backed by public exact-SHA GitHub Actions runs. No coverage percentage or aggregate quality-score badge is claimed.
Metadata
Release files for latent-compass 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| latent_compass-0.3.0.tar.gz | 763.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| latent_compass-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.1 MB
Release files / latent_compass-0.3.0.tar.gz
| Download URL | latent_compass-0.3.0.tar.gz |
|---|---|
| Size | 763.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5c208edca9de4ae6b2d0207fb6547a5b154a7fc72f174b54064308bce704c489
|
|
BLAKE2b-256 checksum How to use checksums |
818c0da8b8b38b7b44f07e60873bf440781938e9ba1e2e802082d1c6acbab026
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / latent_compass-0.3.0-py3-none-any.whl
| Download URL | latent_compass-0.3.0-py3-none-any.whl |
|---|---|
| Size | 317.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
19356fbd327c27f11267df71368648a09668b0c8860524f1caabf6769321392e
|
|
BLAKE2b-256 checksum How to use checksums |
dc1be151db4aa06234b0e64f05448e63a2bc4fc71094af608082c2fd8064db6e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log