Skip to main content

handoverai

Handover of in-context learning state across session boundaries — a reference implementation of spec 2.0-icl-handover.

When a task outlives a single LLM session, the next session gets only what you pass on. This library defines what that record must contain, how big it needs to be, how to check it, and how to tell whether a post-handover failure came from missing information, from the memory budget, or from the continuation procedure.

The record is not judged by how much it resembles the transcript. It is judged by what the continuation can do with it. A record that reproduces earlier wording but drops a binding constraint is a failure; a record that shares no wording with the transcript but preserves what determines the next decision is a success.

pip install handoverai

Or from source:

git clone https://github.com/abshek/handoverai && cd handoverai
pip install -e .

Two minutes

handoverai write draft.json --out handover.json     # build a record
handoverai validate handover.json                   # ten blocking checks
handoverai render handover.json --later-input "..."  # prompt for the next session
from handoverai import (
    ExactEntry, ExactPart, Source, TaskSpec, LaterInputLaw,
    WriterEngine, WriterContext, ReaderEngine,
)

exact = ExactPart()
exact.add(ExactEntry(
    id="c1", kind="constraint",
    statement="Never write to the production database.",
    status="adopted", source=Source(ref="turn-14"), blocking=True,
))

task = TaskSpec(
    task_id="migration-42",
    goal="Finish the schema migration",
    target_form="unified diff",
    scoring_rule="tests pass",
    later_input_law=LaterInputLaw(description="follow-ups about the migration"),
)

record, report = WriterEngine().generate(
    WriterContext(exact=exact), task, limit_tot=8000,
)
prompt, state, log = ReaderEngine().resume(record.to_dict(), later_input="what next?")

What makes this different from compaction

Most compaction shortens a prompt and hopes. This library makes four things checkable that a summarizer leaves implicit.

A record has three parts, and they obey different rules.

Part Rule
H_exact Zero loss. Verbatim, append-only, hashed. Decisions, constraints, prohibitions, rejected options and why, open issues, permissions — everything whose alteration changes which actions are permitted.
H_stat Admissible only with an explicit relation to the task loss: sufficient (preserves the conditional law exactly) or bounded (carries a stated error bound). There is no third grade.
H_residual The observations the statistics do not capture — a rare failure, a counterexample, an unresolved error — selected by a stated rule, by marginal value per bit.

External storage is not free. b_act (in-prompt), b_ext (files, indexes) and b_tot are reported separately, always. A short prompt achieved by pushing text into files is not a small handover, and comparisons are made on b_tot.

Repeated compaction only loses. Summarizing a summary is monotonically lossy. So H_exact is carried forward by copy, never re-derived, and a status change is a new revision that supersedes the old entry rather than an edit.

Refusals are load-bearing. If the constraints do not fit the budget, the writer raises instead of truncating them. If a quantized block cannot support its error bound, the grade is refused rather than downgraded. If an external id does not resolve, the record is incomplete — not degraded.

Guarantees, and where they come from

The parametric path is exact. For Gaussian linear regression, (G_n, b_n) = (X'X, X'y) is predictively sufficient for every later input, and its size is d(d+1)/2 + d scalars regardless of how many demonstrations there were:

from handoverai.stat.gaussian import sufficient_scalar_count
sufficient_scalar_count(64)     # 2144 scalars — whether n was 200 or 200,000

Quantize it and the block declares the predictive-KL bound it actually achieves, with the precondition checked; if the precondition fails, the writer refuses the bounded grade rather than shipping an unbounded approximation.

The nonparametric path grows, and the sizing rules say by how much. Memory is driven by the target accuracy and the function class — not by transcript length, and not by n. The sample floor and the memory floor combine by maximum, and reducing intrinsic dimension is usually a larger win than any coding improvement.

Measuring it honestly

The library is opinionated about what a number is allowed to be called.

  • Only differences of ideal risks are information loss. With a real model as decoder you measure something that mixes in the continuation gap, and every such number is returned labelled end2end, not oracle.
  • The three terms are separable only against a known ideal decoder. The Gaussian harness supplies one, so three_term_split reports (i) budget, (ii) writer, (iii) continuation gap separately. Elsewhere, (ii) and (iii) mix.
  • The deliverable is a risk curve, not a point. budget_sweep and risk_curve exist because a single operating point hides where curves cross.
  • Savings are net of the writer, and quoted at matched quality. The writer is usually itself a model call over the full context, and it is the largest overhead. savings.py computes the break-even call count and refuses to let the compression ratio stand in for a saving.
from handoverai.harness import gauss
episodes = gauss.make_episodes(seed=0, count=32, d=4, n=64)
split = gauss.three_term_split(episodes, my_writer, bit_budget=2048)
# {"i_budget": ..., "ii_writer": ..., "iii_continuation": None}   <- None without a model runner

The leak detector

H-BERN implements the pre-query penalty: a writer committing before the query cannot beat max{0, 1 - B/m} bits of loss. If your measured loss falls below that, it is impossible — something in the pipeline let the writer see the query.

handoverai harness bern --m 16

This catches the most common way a handover evaluation becomes meaningless: a fixture that builds the record after sampling the later input.

Using it from a coding agent

See agent/ — a Claude Code skill, an AGENTS.md for Codex and OpenCode, and an MCP server.

# Claude Code
mkdir -p .claude/skills/handoverai && cp agent/SKILL.md .claude/skills/handoverai/SKILL.md

# Codex / OpenCode
cat agent/AGENTS.md >> AGENTS.md

# MCP
claude mcp add handoverai -- python -m handoverai.mcp_server

The split is deliberate: the model routes, the scripts compute. Deciding what counts as a decision versus an open issue is judgement; computing G_n, sizing a budget, or hashing a payload is not, and a record whose numbers were produced in-context cannot pass the conformance checks.

Command line

Command Purpose
write draft.json Build a record; refuses rather than truncating constraints
validate record.json Ten deterministic checks (§9.4); blocking
render record.json Build p_H, binding record first and verbatim
hop record.json Next-hop draft with H_exact copied forward
check-hops a.json b.json ... M1 constraint adherence across the chain
budget record.json --context-size N M5 and M6, with b_act/b_ext/b_tot
harness gauss|cell|bern|sep Run an analytic harness
arms table.json Check a comparison table is budget-matched
migrate v1.json Lift a 1.0-KatoKato payload (refuses on ungraded statistics)
schema Print the payload schema

Exit codes: 0 ok, 2 refused (a spec-mandated failure), 3 invalid.

Conformance

python conformance/runner.py

Five reference vectors reproduced from their inputs — including the hand-checkable gauss-d2-n5 from the spec's Appendix D, down to the measured predictive KL of 2.2877e-4 against its bound of 5.2190 — and eight malformed payloads that must each be rejected by a named check. conformance/canonicalization.md fixes the serialization so hashes are portable across implementations.

Testing

pytest                      # P1-P14 invariants, engines, conformance

The property table is the spec's §12.3 in executable form. The highest-value single test is P3: the matrix record and the exact synthetic sufficient demonstrations carry identical information, so any measured difference between them under a fixed model is provably a decoder gap, never information loss. That makes the presentation swap the one ablation with an unambiguous attribution.

Scope and honesty

  • sigma2 is treated as known in the parametric block. Unknown noise variance needs extra retained quantities and is not covered; the writer says so rather than claiming a grade it cannot support.
  • The beta > 1 local-polynomial extension is flagged as an extension, not covered by the main theorem.
  • Theorem 5.10's constants are not supplied by the theorem. The floors are reported as shapes with constant_calibrated: false; a measured error below them is not a violation.
  • Cell-statistic risk constants C1..C4 default to 1 and must be calibrated on H-CELL before any epsilon derived from them is published.

References

The full specification is in docs/handover_spec_v2.md. It derives from Masahiro Kato & Taka Kato, Handover of In-Context Learning State Across Session Boundaries, arXiv:2608.14528v1.

License

Apache-2.0.

Metadata

Release files for handoverai 2.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for handoverai 2.0.2
File Size Uploaded
handoverai-2.0.2.tar.gz 99.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for handoverai 2.0.2
File Interpreter ABI Platform
handoverai-2.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 197.9 kB

Release files / handoverai-2.0.2.tar.gz

Download URL handoverai-2.0.2.tar.gz
Size 99.1 kB
Tags Source
SHA-256 checksum
How to use checksums
ea7585e541fce2515c454640486dce02308aa177d9faf32c0250e205f732b5b5
BLAKE2b-256 checksum
How to use checksums
6fd38264fdd06b990ca7dd46918adbe4121e6fb68c14baf5cb11a8309a3eac36
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.1

Release files / handoverai-2.0.2-py3-none-any.whl

Download URL handoverai-2.0.2-py3-none-any.whl
Size 98.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0e16e5894a8c649dc1967a97893453abe275a937720152c06bf782740ab8111d
BLAKE2b-256 checksum
How to use checksums
e0d3ebf0b1321ba016ae3bba4bd2eec7bfe4d94de7677ecb8cdd2dd857b49339
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.1

Release history Release notifications | RSS feed

This release

2.0.2 This release

2 release files

2.0.1

2 release files

2.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page