Skip to main content

Agenda Intelligence MD

PyPI version CI License: MIT

Agenda Intelligence MD is a deterministic evidence-packet linter and compliance orchestration engine for claim-backed AI output. It provides verifiable trust boundaries, guardrail enforcement, and evidence-readiness triage across A2A (Agent-to-Agent), MCP (Model Context Protocol), CLI / Python API, and Serverless Edge Workers (Cloudflare).


Core concepts

Agenda Intelligence MD is a deterministic evidence-packet linter for claim-backed AI output.

Give it claims, the source IDs each claim relies on, optional quotations, and the supplied source text. It returns broken references, quote mismatches, lexical-support gaps, unmatched numbers, claims that negate the source they cite, and the next reviewer actions.

It reports packet completeness, not whether a claim is true:

  • not a factuality verifier;
  • no autonomous live source retrieval;
  • no authorization, approval, or compliance decision;
  • human review is required for every result.

First run

Run the canonical synthetic packet from a source checkout:

git clone https://github.com/vassiliylakhonin/agenda-intelligence-md
cd agenda-intelligence-md
python -m venv .venv
.venv/bin/python -m pip install -e .
.venv/bin/agenda-intelligence check examples/evidence-packet/request.json

Expected shape:

packet_status=packet_complete claims=2 sources=1 factuality=not_assessed
  c1: packet_complete (lexical_support=supported, coverage=1.0)
  c2: packet_complete (lexical_support=supported, coverage=1.0)

Use JSON for an agent loop or CI pipeline:

.venv/bin/agenda-intelligence check examples/evidence-packet/request.json --format json
.venv/bin/agenda-intelligence check examples/evidence-packet/request.json --strict

--strict exits non-zero unless every claim is packet_complete.

Find where a claim could be supported, before deciding what it cites:

.venv/bin/agenda-intelligence discover examples/evidence-review/manifest.json

discover derives literal patterns from each claim — figures and quoted spans first, then content terms, rarest first — and matches every one against every source, reporting the line that matched. Nothing is sampled and no model is called, so it behaves the same on 40 sources and on 4,000. It names the sources a claim's own figures reach but it does not cite, and the ones it cites where not one pattern occurs. Candidates are places to look: nothing here verifies a claim, and a source that supports one in different words does not appear at all.

Review local source files without copying their full text into JSON:

.venv/bin/agenda-intelligence review examples/evidence-review/manifest.json \
  --out evidence-review.md --strict

The manifest keeps claims explicit and points to local UTF-8, Markdown, DOCX, or PDF sources. Paths are resolved inside the manifest directory. DOCX support uses the Python standard library; PDF extraction requires pip install -e ".[documents]". The command makes no network or model call and does not include source text in its JSON or Markdown result. See docs/evidence-review.md.

Install the pinned release without cloning the source and check your own packet:

pip install "agenda-intelligence-md==1.11.0"
agenda-intelligence check /path/to/evidence-packet.json --strict

Generate an interactive standalone HTML reviewer report from local documents:

.venv/bin/agenda-intelligence review examples/evidence-review/manifest.json --format html

The evidence-packet contract

The request has two required collections:

  • claims: a claim ID, claim text, declared source_ids, and optional verbatim quotes;
  • sources: a source ID and the text supplied by the caller.

Request schema: schemas/v1/evidence-packet-request.schema.json

Response schema: schemas/v1/evidence-packet-response.schema.json

Runnable example: examples/evidence-packet/request.json

The response has three packet statuses:

Status Meaning
packet_complete References resolve and the named source text has strong lexical overlap with the claim.
source_review_required References resolve, but lexical support is weak, a numeric value is not present, or the claim and its closest source sentence disagree on negation.
packet_incomplete A source is missing, a quote is absent, or the claim has no source reference.

factuality_status is always not_assessed. A complete packet can still rely on a wrong, stale, biased, or irrelevant source.

Numeric support is format-aware but deliberately conservative. Equivalent scaled values, percentages, and common date forms are compared canonically ($10M ↔ 10,000,000 USD, 62% ↔ 62 percent, and 12 May 2024 ↔ 2024-05-12). Currency is part of the comparison: 10M USD does not support 10M EUR, and the linter performs no currency conversion or approximate-value inference.

Quote presence remains strict after Unicode, typography, whitespace, ellipsis, soft-hyphen, and PDF line-break hyphenation normalization. When an otherwise absent quote has a typo-level candidate at 95% similarity or higher, the quote check may include a bounded near_miss diff for the reviewer. It still reports status: absent and keeps the packet incomplete. Candidates whose numeric facts or negation cues differ are not presented as harmless near misses.

What weighted term overlap can and cannot see

Lexical support is an IDF-weighted share of a claim's content terms that appear in the source it names. Terms that occur throughout the supplied corpus carry less weight than rare entities, while a single-document packet preserves the original plain-overlap scale. Corpus text, sentences, numeric facts, and term sets are indexed once per check run and reused across claims.

Negation is checked. not and no are stopwords and never reach the ratio, so "the board approved it" and "the board did not approve it" score the same against the same source. Where a claim and its closest sentence in the cited source disagree on negation or denial, the claim is downgraded to weak and carries lexical_support_polarity_mismatch. Polarity is read at sentence scope: a negation elsewhere in the same document does not flag an unrelated claim.

Reversed roles are not checked, and are not claimed to be. "A approved a facility for B" and "B approved a facility for A" contain the same terms and both score supported. Deciding who did what to whom is not something term overlap can do, and no heuristic here pretends otherwise. A reviewer still has to read the sentence. The limit is pinned by a test (test_polarity_check_does_not_claim_to_catch_reversed_roles) so it stays visible.

Unicode text is tokenized, but language understanding is not claimed. Cyrillic and Arabic words are no longer discarded, common Russian and Arabic function words are excluded from lexical coverage, and common English, Russian, and Arabic negation cues are checked. A conservative deterministic fold covers common English plurals/verb suffixes and Russian noun/adjective inflections. It is not a full morphological analyzer and does not resolve translation, cross-language support, paraphrases, or semantic roles. Those remain model or reviewer tasks.


Agent Guardrail & Self-Correction Loop

Validate packets and automatically run agent self-correction feedback loops in LangChain, LlamaIndex, CrewAI, DSPy, or vanilla LLM loops:

from agenda_intelligence.integrations import EvidenceClaim, EvidencePacket, EvidencePacketGuardrail, EvidenceSource

guardrail = EvidencePacketGuardrail(strict=True, max_repair_attempts=2)

# Optional zero-dependency typed input; plain dictionaries remain supported.
packet = EvidencePacket(
    claims=(EvidenceClaim("c1", "The board approved the budget.", ("s1",)),),
    sources=(EvidenceSource("s1", "The board approved the budget after review."),),
)

# Direct check
result = guardrail.check(packet)
if not guardrail.is_complete(result):
    repair_prompt = guardrail.get_repair_prompt(packet_json, result)
    # Provide repair_prompt back to LLM to revise output

# Automated retry loop with custom LLM generation function
final_packet, success, repair_history = guardrail.validate_or_repair(
    packet_json,
    llm_repair_fn=lambda prompt: my_llm_chain.invoke({"prompt": prompt}),
)

# Event-loop pipelines can await check_async(...) or validate_or_repair_async(...).
# LangGraph can use the dependency-free async node returned by:
node = guardrail.as_langgraph_node(packet_key="evidence_packet", result_key="evidence_check")

Concurrency & A2A Demos

The repository includes runnable end-to-end demonstrations of the agent-first architecture:

  • Bounded concurrency example (examples/infinite-swarm-batch.py): Sends 250 synthetic requests and reports transport latency and actual task states. It is a load demonstration, not a capacity benchmark or comparison with staff.
  • A2A step-up simulation (examples/agent-to-agent-negotiation.py): Demonstrates a synthetic request being stopped until operator-authorization evidence is supplied. No real transaction is authorized.
  • Profile scaffolder (scripts/agent-factory.py): Creates starter files for a proposed vertical profile. Generated files are inactive until schemas, implementation, tests, and review are added.

GitHub Action CI Integration

Add deterministic evidence linting to your repository CI workflow (.github/workflows/evidence-lint.yml):

name: Evidence Lint
on: [push, pull_request]

permissions:
  contents: read
  security-events: write

jobs:
  lint-evidence:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Validate evidence packet
        uses: vassiliylakhonin/agenda-intelligence-md@main
        with:
          path: 'evidence/packet.json'
          command: 'check'
          format: 'sarif'
          strict: 'true'

With format: sarif, findings are uploaded to GitHub code scanning and point to the corresponding claim_id line in the packet JSON. text and json output remain available.


Python API

import json
from pathlib import Path

from agenda_intelligence.services import check_evidence_packet, build_repair_prompt

packet = json.loads(Path("examples/evidence-packet/request.json").read_text())
result = check_evidence_packet(packet)
print(result["response"]["packet_status"])

# Generate actionable markdown repair instructions for an agent
if result["response"]["packet_status"] != "packet_complete":
    prompt = build_repair_prompt(packet, result["response"])
    print(prompt)

The service layer is stateless. It does not persist packet contents or fetch missing sources.


What this is

  • A small JSON contract for claim-backed AI output.
  • A deterministic preflight before human review.
  • A CLI and Python service suitable for local and CI use.
  • A local-file review adapter that returns a reviewer-facing Markdown or JSON result.
  • An inspectable base for domain-specific compatibility profiles.

What this is not

  • A general LLM evaluation platform.
  • A GRC, vendor-management, or document-storage system.
  • An agent authorization or policy-enforcement layer.
  • Legal, compliance, sanctions, financial, investment, insurance, or trading advice.
  • Proof that a source or claim is factually correct.

Why a repo full of markdown?

The repository predates the evidence-packet focus and also packages agent reasoning instructions. Files under skills/ are executable instructions for compatible agent runtimes, not ordinary prose documentation. They remain available for compatibility, but they are not the primary product interface.


MCP

The packaged MCP server exposes the same evidence-packet preflight to agent clients:

{
  "mcpServers": {
    "agenda-intelligence": {
      "command": "uvx",
      "args": ["--from", "agenda-intelligence-md", "agenda-intelligence-mcp"]
    }
  }
}

Run a focused stdio example against an editable install:

.venv/bin/python examples/evidence-packet/mcp_client.py \
  --command ".venv/bin/agenda-intelligence-mcp"

The example initializes the MCP server, calls check_evidence_packet with the synthetic packet, and prints only the review summary. See examples/evidence-packet/mcp_client.py and MCP.md.

Before using the result for an irreversible or high-stakes action, record the goal, supplied evidence, suspected unreliable evidence, assumptions, intended action, and stop/escalation conditions. The tool checks packet structure, not whether a claim is true or an action is authorized.

Existing MCP tools such as audit_claims, verify_quotes, grounded_check, and verify_claims remain compatible; no tool was removed or renamed.

pre_action_check adds a stateless action boundary on top of the existing claim audit. It returns continue, request_evidence, require_approval, or stop from caller-supplied evidence, risk, policy-check results, and an optional external approval reference. The caller still authenticates the actor, stores approvals, enforces the result, and performs the action. The request and response contracts are pre-action-check-request.schema.json and pre-action-check-response.schema.json. Twenty illustrative replay cases are in examples/pre-action-check/replay-cases.json.

Two authoring tools, create_brief and append_evidence, let an agent assemble a brief or an evidence pack step by step inside the contract instead of hand-building JSON and validating it afterwards. Both are deterministic and stateless: they validate on every call and return the document to the caller. They do not write files, retrieve sources, draft prose, or assess factual truth, and append_evidence never infers a supported claim status on its own.

Claude Code plugin installation also remains available:

/plugin marketplace add vassiliylakhonin/agenda-intelligence-md
/plugin install agenda-intelligence@agenda-intelligence

Compatibility profiles and adapters

The strategic-intelligence shell, HTTP API, A2A adapter, Cloudflare Workers, and five domain profiles remain in the repository. They demonstrate how the same service layer can be wrapped for different transports and domains. They represent active prototypes and technical wedges for vertical domains.

Compatibility surface Reference
Strategic agenda analysis Agenda-Intelligence.md
HTTP API docs/deployment/http-api.md
A2A adapter docs/deployment/a2a-adapter.md
Middle Corridor example docs/use-cases/kazakhstan-middle-corridor.md
CIS secondary-sanctions example docs/use-cases/cis-secondary-sanctions.md
Agentic interaction example docs/use-cases/agentic-interaction-trust.md
Gulf maritime example docs/use-cases/gulf-maritime-exposure.md
Kazakhstan market-entry example docs/use-cases/kazakhstan-market-entry-readiness.md
Live A2A demo pack docs/agenstry/demo-pack.md

The compatibility profiles are evidence-routing examples only. They do not provide legal, compliance, sanctions, financial, investment, insurance, or trading advice. Human review is required before any commercial action.


Verification Contract

The repository keeps three checks separate:

  1. check reports packet completeness and lexical-support diagnostics.
  2. grounded-check performs the older claim-to-corpus lexical diagnostic.
  3. verify-claims applies declared freshness, authority, independence, jurisdiction, and identifier rules to caller-supplied evidence.

None discovers the right sources for the caller. verified in the bounded Claim Verdict contract means the supplied evidence meets that declared contract; it is not absolute truth.


Schemas

Canonical schemas live under schemas/v1/. Packaged copies under src/agenda_intelligence/data/schemas/v1/ must remain byte-equivalent; CI checks this invariant.

Start with:

The full registry is in agent-manifest.json.


Before / after and benchmarks

The older agenda-analysis evaluation surface remains available for regression and compatibility work:

These are evaluation fixtures, not customer evidence or production benchmarks.


AnalysisBank

analysis-bank/ contains compatibility fixtures for reasoning-memory retrieval and failure-pattern regression. It is not part of the primary evidence-packet workflow.


Status

Surface Status
Evidence-packet request/response schemas Implemented
check_evidence_packet Python service Implemented
agenda-intelligence check packet auto-detection Implemented
agenda-intelligence review local-file workflow Implemented for UTF-8, Markdown, DOCX, and optional PDF input
agenda-intelligence review --format html Implemented (Generative UI)
check_evidence_packet MCP tool Implemented
AI Fleet (Vertical Workers) Active (12 profiles deployed on Cloudflare Edge)
Agent Financial Guard Implemented (Pre-sign transaction firewall for AI agents)
M2M Escrow Arbiter & Base Contract Implemented (Autonomous B2B dispute resolution on Base)
Live Source Retrieval Optional per profile; currently unconfigured in the hosted fleet

Current classification: Ecosystem Expansion & R&D.


Documentation

Topic File
Pitch Deck (12 Slides) docs/pitch/PITCH_DECK.md
Case Studies docs/pitch/CASE_STUDIES.md
Unit Economics docs/pitch/UNIT_ECONOMICS.md
Adoption ADOPTION.md
Quickstart docs/quickstart.md
Evidence audit docs/evidence-audit.md
Local evidence review docs/evidence-review.md
Factuality boundary docs/factual-verification.md
Evaluation docs/evaluation.md
Source policy SOURCE_POLICY.md
Security SECURITY.md
Threat model docs/threat-model.md
Roadmap ROADMAP.md

Repository layout

schemas/v1/                    public JSON contracts
src/agenda_intelligence/       Python service and transport adapters
examples/evidence-packet/      canonical packet example
tests/                         contract and regression tests
skills/                        compatibility agent instructions
deploy/cloudflare-worker/      compatibility Worker implementation
docs/                          reference and compatibility documentation

Development

pip install -e ".[dev]"
make ci
make verification-report

make verify-local also runs the compatibility Cloudflare Worker tests. make verification-report runs both verification surfaces and writes .verification/results.json: a deterministic, machine-readable record of the checks and hashed contracts. It uses no paid APIs and deliberately makes no claim about factual truth, live deployment health, adoption, or market value.


Roadmap

The current phase focuses on Product-Led Growth & Ecosystem Expansion. We are rapidly iterating on Generative UI for interactive evidence dashboards, deploying new vertical AI workers for adjacent domains (e.g., ESG, supply chain), and registering capabilities with agent catalogs (Agenstry).

See ROADMAP.md for the active expansion initiatives.


License

MIT

Release files for agenda-intelligence-md 1.11.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agenda-intelligence-md 1.11.0
File Size Uploaded
agenda_intelligence_md-1.11.0.tar.gz 767.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agenda-intelligence-md 1.11.0
File Interpreter ABI Platform
agenda_intelligence_md-1.11.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.2 MB

Release files / agenda_intelligence_md-1.11.0.tar.gz

Download URL agenda_intelligence_md-1.11.0.tar.gz
Size 767.0 kB
Tags Source
SHA-256 checksum
How to use checksums
0f400ecd8307d933bb477827b912a516c3eea09dc22e5196a9b5d78cc7b551ee
BLAKE2b-256 checksum
How to use checksums
dabc4d0234d50c7e81ed6bc3c16317946d74c36c55834c60e4755157f2f43d07
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / agenda_intelligence_md-1.11.0-py3-none-any.whl

Download URL agenda_intelligence_md-1.11.0-py3-none-any.whl
Size 440.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
65c6f430458d26126b824b90a57fbd37308216dbde0879a2d64d27b9f2ff5697
BLAKE2b-256 checksum
How to use checksums
eff46bf80f16be0f552a21df36b28b947872fa99bc4102d20af2d4cf6744c924
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

1.12.1

2 release files

1.12.0

2 release files

1.11.1

2 release files

This release

1.11.0 This release

2 release files

1.10.0

2 release files

1.9.0

2 release files

1.8.0

2 release files

1.7.1

2 release files

1.7.0

2 release files

1.6.0

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.8

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page