Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Corollary

An agent runtime where the unit of state is a belief, not a message.

Every conclusion your agent reaches carries its proof. Correct one fact, and everything that followed from it updates itself.

CI Status: pre-alpha License: MIT Python 3.10+ Typed PRs welcome

Documentation · Getting started · Examples · Contributing


Animated diagram: beliefs linked by what they follow from. Q2 revenue is corrected, every conclusion that depended on it goes OUT, and only those are re-derived, while an independent risk belief stays untouched.

⚠️ Corollary is pre-alpha. The v0.1 kernel and agent runtime are implemented and tested, and most of v0.2 and v0.3 are in. Interfaces may still change before 1.0. If the idea resonates, star the repo, open an issue, or help shape the belief schema.

Why "Corollary"

In mathematics, a corollary is a result that follows directly from something already proven. It stands only as long as the theorem behind it stands.

That is the contract Corollary enforces for AI agents: every conclusion must follow from evidence, and when the evidence falls, the conclusion falls with it.

The problem

Every agent framework today stores state the same way: as a message log. A chat transcript, growing step by step.

That single design choice is why long-running agents rot:

  • A hallucination at step 3 is just text. At step 40 the agent reads it with the same trust as a verified tool result.
  • Nobody can answer why the agent believes something without re-reading the entire transcript.
  • When one input turns out to be wrong, the only fix is to re-run everything, or hope a human catches every downstream mistake.
  • Two sources disagree, and the model silently picks one. You never learn there was a conflict.

Better models reduce these failures. They cannot eliminate them, because the problem isn't the model. It's the data structure.

The idea

Corollary replaces the message log with a belief base.

A belief is a claim plus everything needed to trust it:

Belief
├── claim          "Q3 revenue grew 4.65% over Q2"
├── confidence     0.92
├── source         tool:sec_filings  |  model:any-llm  |  human:alice
├── valid_until    2026-12-31
└── follows_from   [ belief:revenue_Q2, belief:revenue_Q3, rule:growth ]

Underneath sits a Truth Maintenance System (Doyle, 1979): every belief records what supports it, and when that support is withdrawn, dependent beliefs are retracted automatically. A message log still exists, but only as a view derived from the belief base, never as the source of truth.

Truth maintenance never took off for general reasoning because humans had to write every justification by hand. LLMs can now propose them at scale. That was the missing piece.

What this unlocks

Retraction cascades. A tool returns a corrected figure. Every conclusion derived from the old value is retracted and re-derived, surgically. You fix one fact, not a transcript.

Contradictions as first-class objects. When evidence conflicts, Corollary raises a Conflict carrying both support chains instead of letting the model guess. Resolve it by policy, by source rank, or by asking a human.

Proof-carrying answers. Every answer ships with its proof: the graph of beliefs it follows from. A deterministic verifier checks it. Cited spans exist, arithmetic re-executes, dates are consistent, and every leaf traces back to a tool, a document, or a person. Trust moves from the model to the proof.

Beliefs that expire. A stock price is valid for a minute; a company's headquarters for a year. Stale beliefs trigger re-verification instead of silent reuse.

Trust that is earned. Every source, the model included, is trusted as much as its track record justifies. When a source is shown to be wrong, everything it reported counts for less. Independent sources that agree reinforce each other. See Confidence.

Auditable memory. New sessions inherit verified beliefs with their proofs attached, not raw transcripts.

Multi-agent by argument (planned). Agents exchange beliefs with their support. The receiver can accept, reject, or demand proof.

A taste of the API

from corollary import Agent, BeliefBase, tool

kb = BeliefBase()


@tool(trust="high")
def get_revenue(quarter: str) -> float:
    """Quarterly revenue in USD from SEC filings."""
    ...


agent = Agent(model="anthropic:claude-opus-5-5", beliefs=kb, tools=[get_revenue])

report = agent.run("Compare Q2 and Q3 revenue and assess the growth trend.")

print(report.answer)
print(report.proof)  # the graph of beliefs the answer follows from
report.verify()  # deterministic checks over the proof

Now an upstream figure is corrected:

kb.retract("revenue:Q2", reason="restated in 10-K/A")
kb.assert_("revenue:Q2", 4.1e9, source="tool:get_revenue")

for change in agent.repair():
    print(change)
# OUT  revenue:Q2          (retracted: restated in 10-K/A)
# OUT  growth:Q3_vs_Q2     (lost support: revenue:Q2)
# OUT  trend:Q3            (lost support: revenue:Q2, growth:Q3_vs_Q2)
# OUT  answer              (lost support: revenue:Q2, growth:Q3_vs_Q2, trend:Q3)
# IN   revenue:Q2          (asserted by tool:get_revenue)
# IN   growth:Q3_vs_Q2     (re-derived: 9.7561)
# IN   trend:Q3            (re-derived: 'strong')
# IN   answer              (re-derived: 'Q3 revenue grew 9.76% over Q2: strong growth.')

print(report.answer)  # the repaired answer

No re-run. No transcript archaeology. A diff of what changed. During repair, the model sees only the current inputs of each belief it re-derives. The retracted figure never reaches it again.

The kernel works without a model too. In examples/self_repairing_report.py, a 20-conclusion financial report repairs itself after a restatement: 12 conclusions change, 7 are never touched, the growth trend is recomputed, found unchanged, and stops the cascade, all in 13 rule evaluations and zero model calls.

How it works

Corollary is not a system prompt asking a model to "track its reasoning." A prompt is a request; a model can ignore it. Every guarantee in Corollary is enforced by code outside the model.

The key inversion: the model doesn't own the state, the runtime does. The model is a stateless proposer. It never sees a transcript and never writes state directly. It receives a context built by the runtime and can only respond with structured claims, which the runtime validates before accepting.

flowchart LR
    U[Task] --> P[LLM proposer]
    P -->|claims + dependencies| C[Contract validator]
    T[Tools] -->|executed by runtime| PR[Premises]
    PR --> K
    C --> K[(Kernel: belief graph + TMS)]
    K -->|IN beliefs only| V[Projector]
    V --> P
    K --> X[Conflict detector]
    X --> R[Resolver: policy / human]
    R --> K
    K --> A[Answer + proof]
    A --> VER[Verifier]

1. The kernel (pure code, no LLM). A dependency graph plus a labeling algorithm. Each node is a belief; each edge is a justification. The kernel computes which beliefs are IN or OUT from the current graph. It is deterministic, fast, and testable like any data structure. Retraction walks only the affected region, so cost scales with the change, not the size of the base.

2. The contract (structured output, enforced). The model must answer in a schema: the claim, its value, and the IDs of the beliefs it follows from. The runtime rejects anything that doesn't parse or that depends on a belief that doesn't exist or is currently OUT. Tool calls go through the runtime, which executes them and records the results as premises itself. The model cannot fabricate a tool result.

3. The projector (what makes retraction real). Every model call receives a context assembled from IN beliefs only. When a belief is retracted it doesn't just get a flag; it disappears from everything the model can see. There is no transcript for a wrong fact to leak from. Forgetting is structural, not requested.

4. The verifier. Answers are exported with their proof and checked deterministically: arithmetic re-executes, citations are span-matched, temporal claims are checked for consistency.

What is LLM and what is code

The LLM does Code does
Proposes claims Records premises from tools, documents, humans
Chooses which tools to call Executes tools and stores their results
Re-derives conclusions after a retraction Validates every claim against the contract
Verifies arithmetic, citations, dates
Computes IN / OUT for every belief
Builds every context window
Cascades retractions

Corollary doesn't make the model more honest. It makes the model's honesty irrelevant to whether bad facts propagate.

The hard part, stated honestly

A model may declare that a conclusion depends on A when it also used B. Corollary handles this in layers:

  • Conservative default. A claim is assumed to depend on everything in the context it was generated from (dependencies="conservative"). The projector guarantees the model could not have used anything else.
  • Narrowing by ablation. agent.narrow(key) regenerates a claim with one belief removed from context at a time. If the answer doesn't change, that dependency is pruned.
  • Deterministic checks on numeric and cited claims: formulas are re-executed, quotes are span-matched, and numbers that appear from nowhere are flagged.

The resulting guarantee is over-retraction, never under-retraction. You may occasionally re-derive something unnecessarily, but a retracted fact can never silently survive in a conclusion.

What code can't catch is listed too: see limits, stated honestly.

Core concepts

Concept What it is
Belief A claim with confidence, source, and revision. Its status (IN / OUT) is computed by the kernel.
Justification A link from a set of antecedent beliefs (and optionally a rule or formula) to a conclusion.
Premise A belief grounded directly in a tool, document span, or human. No antecedents.
Rule A deterministic derivation, re-run automatically when its inputs change.
Conflict Two IN beliefs that cannot both hold, with both support chains attached.
Constraint An invariant over beliefs; a violation raises a Conflict.
Proof The justification subgraph behind an answer. Exportable, verifiable, diffable.
Projector Builds every model context from IN beliefs.
TrustPolicy How much each source type is believed, and what is hidden from the model.
TrustLedger Learns how reliable each source really is from its track record.

Read more in Core concepts.

How Corollary differs

Message-log agents Memory layers Prompted confidence scoring Corollary
Unit of state Message Fact or summary Annotated message Belief with justifications
Knows why it believes X ❌ Source only Self-reported ✅ full derivation
Fix one input, update dependents Re-run Manual ❌ ✅ automatic cascade
Retracted facts removed from context ❌ ❌ ❌ ✅ enforced by projector
Surfaces contradictions ❌ May overwrite ❌ ✅ first-class Conflict
Verifiable output ❌ ❌ ❌ ✅ proof + deterministic verifier
Where guarantees live — Storage The prompt The runtime

Corollary is not a memory plugin and not a prompting strategy. The belief base is the agent; memory, context, tool results and inter-agent messages are all projections of it. See Related work for the prior art it builds on.

Installation

Corollary requires Python 3.10+. The core has no runtime dependencies; model SDKs are optional extras.

pip install corollary                 # the kernel and the agent runtime
pip install "corollary[anthropic]"    # + the Claude adapter (official Anthropic SDK)
pip install "corollary[openai]"       # + the OpenAI-compatible adapter

To develop:

git clone https://github.com/gabe-santana/corollary.git
cd corollary
pip install -e ".[dev]"
pytest

Then follow Getting started, or run an example. Neither of the first two needs an API key:

python examples/self_repairing_report.py   # 20 conclusions repair themselves, zero model calls
python examples/agent_offline.py           # a full agent run with a scripted model
python examples/agent_claude.py            # the same with Claude (needs ANTHROPIC_API_KEY)

Documentation

Getting started Install, a self-repairing belief base, an offline agent, Claude
Core concepts Beliefs, revisions, justifications, labels, confidence, conflicts, proofs
Guides Belief base · Agents · Contract · Verification · Conflicts · Documents · Time · Models · Persistence
Architecture and guarantees How the kernel works, what is guaranteed, and the limits
API reference Every public class and method
FAQ Quick answers

Roadmap

v0.1: the kernel ✅

  • Belief, Justification, BeliefBase with JTMS labeling and retraction
  • Claim contract (structured output) with model-agnostic adapters
  • Runtime-executed tools that emit premises
  • Projector: context built from IN beliefs only
  • Flagship demo: a 20-conclusion research report that repairs itself when one input is corrected

v0.2: proof and verification

  • Deterministic verifiers: arithmetic re-execution, citation spans, temporal consistency
  • Proof export (JSON, Mermaid, Graphviz)
  • Interactive proof graph viewer
  • Ablation-based dependency narrowing
  • Conflict detection and pluggable resolvers

v0.3: time and memory

  • Validity windows and automatic re-verification of stale beliefs
  • Persistent belief snapshots (JSON)
  • Persistent belief stores (SQLite, Postgres)
  • Cross-session inheritance policies for verified beliefs

v0.4: many agents

  • Belief exchange protocol: accept / reject / demand proof
  • Assumption-based (ATMS) mode for exploring alternative hypotheses in parallel
  • MCP server so any agent can use a Corollary belief base

Open problems we want help with

These are the hard parts. Each one is a research contribution waiting to happen.

  • Confabulated dependencies. How cheaply can we detect a declared dependency the model didn't actually use?
  • Granularity. Too fine and the graph explodes; too coarse and cascades become blunt. What is the right atomic claim?
  • Equivalence. When are two differently worded claims the same belief?
  • Developer experience. It has to feel as simple as appending to a message list, or nobody will switch.
  • Benchmarks. There is no standard evaluation for how well an agent recovers from a corrected input. We want to build one.

Contributing

Corollary is in the design phase, which is the best time to shape it. Good first contributions:

  • Comment on the belief schema in Discussions
  • Propose a demo scenario from your domain (finance, law, science, operations)
  • Implement a verifier
  • Write a model adapter for your favorite provider

See CONTRIBUTING.md to get started. Everyone taking part is expected to follow the Code of Conduct. To report a vulnerability, see SECURITY.md.

Background

Corollary stands on decades of work in knowledge representation:

  • Jon Doyle, A Truth Maintenance System (1979)
  • Johan de Kleer, An Assumption-based TMS (1986)
  • Alchourrón, Gärdenfors & Makinson, the AGM theory of belief revision (1985)

What's new is the pairing: classical reason maintenance as the kernel, LLMs as the engine that proposes beliefs and justifications at scale.

License

MIT. See LICENSE. You can use, modify and distribute Corollary freely, including in commercial and closed-source projects.


Agents shouldn't just conclude. Their conclusions should follow.

Metadata

Release files for corollary 0.1.0a2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for corollary 0.1.0a2
File Size Uploaded
corollary-0.1.0a2.tar.gz 1.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for corollary 0.1.0a2
File Interpreter ABI Platform
corollary-0.1.0a2-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / corollary-0.1.0a2.tar.gz

Download URL corollary-0.1.0a2.tar.gz
Size 1.0 MB
Tags Source
SHA-256 checksum
How to use checksums
11964b498f4ae1e8cc4c51f7d6bce2dde6dbc49546a49539b5918f775fecf857
BLAKE2b-256 checksum
How to use checksums
e333bf8166de148af3bbda20c3639ee213005f5eafe99ded095c75afc8b282c6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / corollary-0.1.0a2-py3-none-any.whl

Download URL corollary-0.1.0a2-py3-none-any.whl
Size 90.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a2d749991fa5d28e5192a4022aa3f2e7aa3713d905a51cf33a9a95cf86fa0ec0
BLAKE2b-256 checksum
How to use checksums
98348a065a5f91001df9c0bba779144d6d4181cd9df54e1047ad3dd370b2f5f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0a2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page