Skip to main content

srdcheck

Deterministic rails for game-running agents — so intelligence is spent only where intelligence is the only thing that works.

Machine verdicts over the rules of the System Reference Document 5.2.1: cited, reproducible, delivered in milliseconds with zero tokens, and honest enough to refuse questions that aren't the rules' to answer. The rules lawyer for agents.

Status: v0.1 — young but real, building in the open. The kill tests that shaped the product — including the one that killed half our original idea — are in eval/RESULTS-phase0.md; the with/without-rails demo is in demo/mage-hand/; the truth scorecard below is generated by CI.

Why this exists

A model running a game is a brilliant improviser with a finite attention budget. Every mechanical micro-check it handles in-context — is this legal, is that slot spent, does the reaction refresh this round — spends tokens and attention that belong to the only work that needs a mind: the story, the improvisation, the table. And the checks a model can answer, it cannot prove, cannot reproduce, and — as our own benchmark showed — will not refuse when the question is outside the rules' jurisdiction.

We tested this before building. Frontier models answered our SRD rules questions nearly perfectly — and confidently ruled on house rules, GM discretion, and content that doesn't exist in the SRD, where the only correct answer is "not my call." Small local models got 19–30% wrong with zero refusals. So srdcheck does not compete with what models know. It is a rail: state in, verdict out, citations attached, deterministically, every time.

It matters most when the table isn't one model but several. A game increasingly runs as multiple agents sharing one evolving state — a GM, player agents, a rules referee, tool-callers. There the problem stops being any single agent's attention budget and becomes consistency: five capable models will each read the SRD and adjudicate grapple, concentration, or a death save a little differently, and the shared game state quietly desyncs. No amount of individual intelligence guarantees they agree with each other, and every re-derivation is another chance to drift. A deterministic rail they all call — the same atoms, the same citations, the same verdict every time — is the thing that keeps them in lockstep and gives every ruling a single provenance any of them (or a human) can audit. That is where srdcheck stops being a convenience and becomes load-bearing.

What it is

srdcheck answers one kind of question: is this legal under the rules? — and one better one: what is legal right now?

  • Verdicts, not vibes. Exit code 0 = legal, 1 = illegal, 2 = cannot adjudicate. Every verdict carries its chain of SRD 5.2.1 citations. A rule we cannot cite is a rule we do not have.

  • Judge, never simulate. No dice, no narration, no owned game state. State comes in with the query; a verdict goes out. The kernel is a stateless pure function — embeddable in anyone's DM product, VTT, or agent.

  • Deterministic and fast. No LLM call anywhere in the verdict path. Runs local and offline.

  • For agents first. MCP + CLI, --pipe, --schema, tool.json at the repo root. Humans get a plain-English why in the same payload.

  • Rulesets are adapters. The kernel knows no game; all rule content loads from adapter packages, each carrying its own provenance manifest — source document, hash, license, attribution — that every verdict cites through. The SRD 5.2.1 adapter ships in this repo as the reference implementation. Anyone can build an adapter for another ruleset — a community, a private table, or a publisher shipping a first-party adapter for their own IP — and their content never passes through this project. The adapter catalog points; it never hosts.

See docs/product-truths.md for the invariants this project holds itself to, and docs/anatomy-of-a-turn.md for where srdcheck sits in a game-running agent's pipeline — a combat turn, a stealth infiltration, and the Mage Hand test, worked end to end.

Show your work. A big context window can read the SRD; srdcheck's value is the layer above — a consistent, cited set of rule atoms and a shared schema that several agents can rely on without each re-deriving the rules into context and drifting apart. docs/atom-concordance.md is the audit trail for that layer: every rule the engine applies, mapped to the exact SRD text it's grounded in (forward and inverse), generated and CI-checked-fresh, with a verbatim-citation oracle proving each quote really is on its cited page. A machine-readable atom-concordance.json is the slim index an agent can load to know what the rail covers and where it came from.

What srdcheck does not check yet

Honesty is the product (truth T8), so the boundaries are stated, not implied. As of v0.1 the SRD adapter covers combat-turn fundamentals; it does not yet check:

  • Feature prerequisites. turn.plan judges the turn's action economy — that you spent at most one action, one bonus action, one reaction, one spell slot. It does not verify that a feature actually grants a given action. A lone two-weapon-fighting offhand attack is "action-economy legal" even though the 2024 rules require the Attack action first; that prerequisite lives in the character's features, which this version doesn't model. The success message says so explicitly.
  • Per-spell effects. Damage/HP, death saves, and saving throws are now modeled (the reducer folds damage/heal/death-save; save.check and concentration.check resolve caller-rolled saves), but what a given spell does — Fireball's 8d6, Hold Person's paralysis on a failed save — is the long-tail swamp and stays refused. srdcheck computes the save; it does not roll the dice or apply the spell's bespoke effect.
  • Conditions — all 15 are fully modeled across every surface srdcheck exposes: attack rolls/legality, action economy and Speed (incl. Exhaustion's graduated reduction), condition-aware saving throws (auto-fail Str/Dex, Restrained's Dex Disadvantage) and ability checks (Blinded/Deafened auto-fail sight/hearing, Poisoned/Frightened Disadvantage), damage typing (Petrified's Resistance to all), and condition immunity (Petrified vs Poisoned). A completeness oracle enforces in CI that every codified clause is cited-and-modeled — the only clauses left deferred are the three srdcheck never owns by doctrine (T6): positional geometry (Frightened's "can't approach"), the initiative order, and stealth/perception contests. Those return exit 2 by principle, not for lack of building.
  • Content outside the SRD 5.2.1 — subclasses, feats, spells, and monsters not in the SRD return exit 2 from jurisdiction.

Nonsensical inputs (negative Speed, a 99th-level spell, exhaustion past 6) return exit 2 rather than a confident-looking answer. When in doubt, srdcheck refuses — a wrong verdict is the only unforgivable bug.

Try it now

$ pip install git+https://github.com/chaoz23/srdcheck
$ python -m srdcheck jurisdiction "Fireball"          # exit 0 — known content
$ python -m srdcheck jurisdiction "Hexblade"          # exit 2 — not in the SRD, honestly refused
$ python -m srdcheck query mage-hand.use '{"kind": "attack"}'
{
  "verdict": "illegal",
  "exit_code": 1,
  "why": "The hand can't attack.",
  "citations": [{"section": "SRD 5.2.1 p.145 'Spells > Mage Hand'", "page": 145,
                 "quote": "The hand can't attack"}],
  "rule_ids": ["mage-hand.cant-attack"],
  "adapter": "srd-5.2.1@0.1.4"
}
$ python -m srdcheck --schema                          # I/O contract for agents

Deterministic, offline, no tokens, sub-millisecond. The query surface is young and growing slice by slice — the architecture (kernel + adapters, spec at v0.9 RC) is the point.

For agents (MCP)

srdcheck is an MCP server with zero dependencies — stdlib only. After pip install, the command is srdcheck-mcp; from a clone it's:

{
  "mcpServers": {
    "srdcheck": {
      "command": "python3",
      "args": ["-m", "srdcheck.mcp"],
      "cwd": "/path/to/srdcheck"
    }
  }
}

Twenty tools: jurisdiction, turn_plan, turn_options, reaction_available, roll_compose, attack_modifiers (conditions, distance, and the ranged-in-close-combat Disadvantage), mage_hand_use, event_apply (the state reducer — folds a declared event, including damage/heal/death-save with damage typing, into a verdict plus a hash-stamped next state), save_check/check_make/concentration_check (resolve a caller-rolled d20 against a DC — condition-aware saving throws and ability checks, and the concentration-save DC from damage taken — srdcheck folds, never rolls), opportunity_attack_provoked (the provoke rule and its exceptions), grapple_initiate (Grapple/Shove DC + size legality), passive_perception (10 + modifier, and it refuses the non-SRD ±5), help_assist (the Help action's proficiency gate), creature_valid/creature_stats (cited SRD 5.2.1 creature validity + CR/XP), encounter_xp_budget (the p.202 party-level XP budget), and the toy adapter's ttt_move/ttt_options (which exist to prove the adapter spec). Every call returns the same verdict object as the CLI (verdict, exit_code, why, citations with source quotes) as structured content. An illegal verdict is a result, not an error; cannot-adjudicate is an honest refusal, not a failure. Tool descriptions and schemas come from the loaded adapters, so new adapters extend the tool list without kernel changes. See also tool.json for the CLI surface.

As a Python library

Consume a ruleset's content through a stable, versioned interface — no coupling to internal file paths:

from srdcheck import load_adapter, available_adapters

available_adapters()               # ['srd-5.1', 'srd-5.2.1', 'toy-tictactoe']
a = load_adapter("srd-5.2.1")      # a versioned identifier
a.categories()                     # the content categories this adapter carries
a.names("creature")                # every creature name (326, complete from the stat blocks)
a.record("creature", "Ghast")      # {"name": "Ghast", "cr": "2", "xp": 450, "citation": "SRD 5.2.1 p.287"}
a.query("encounter.xp-budget", {"level": 3, "difficulty": "moderate"})  # a verdict dict

Cross-version edition-trap detection — catch a name that was valid in an older edition but renamed/removed in the current one (the classic LLM failure, e.g. "Goblin" in 2014 → "Goblin Warrior" in 2024):

from srdcheck import edition_check
edition_check("Goblin", "creature")          # exit 1 — a srd-5.1 name, not in 5.2.1;
                                             #   data.candidates_in_current suggests Goblin Warrior/Minion
edition_check("Aboleth", "creature")         # exit 0 — valid in both
# or from the CLI: srdcheck edition-check "Goblin" --category creature

Adapter identifiers carry their version, so load_adapter("srd-5.2.1") and load_adapter("srd-5.1") coexist without a breaking change; edition_check is caller-parameterized (current / priors), so it generalizes to any versions and any category. The handle is content-neutral: it knows about categories and records, never about a specific ruleset's vocabulary.

The benchmark

bench/ is the rules-fidelity referee: versioned question sets with SRD-cited gold verdicts, a harness that scores any model or agent (gemini:, ollama:, or cmd:your-agent on stdin/stdout), and a generated scorecard that reports wrong-rate, refusal-rate, and false-confidence separately, per category, with no aggregate number — ever. A generated leaderboard ranks subjects by wrong-count alone (the one unforgivable failure), and any third party can submit a tamper-checked result — you can't grade your own homework, because the golds are the set's. Published findings: frontier models ace codified rules and fail by false confidence exactly where the rules end — and, on a per-horizon drift lane whose gold verdicts are derived by the engine itself, they don't drift even at 30-round horizons, while a local 8B model's errors are mistakes of rule-knowledge, not memory. Benchmark your own DM product with one command.

Truth scorecard

Every tagged release publishes a scorecard against the product truths — generated by CI, never hand-edited, no aggregate score.

Generated by scripts/truth_scorecard.py — regenerated and diff-checked in CI, never hand-edited. Statuses are honest: structural and held in review mean exactly that.

truth claim status evidence
T1 wrong verdicts enforced in CI 219 tests including gold suites ported from the Phase 0 eval; any wrong verdict fails the build
T2 no citation, no rule enforced in CI 88/88 rule atoms carry verbatim source quotes; every-verdict-cites tests on all adjudicated paths
T3 advise, never overrule structural the API has no blocking or veto interface to wire; verdicts are advisory by construction
T4 one payload, two audiences enforced in CI every verdict carries machine fields plus a templated plain-English why; schema-tested
T5 enumeration is the product proven in CI consistency sweeps (50 turn states + toy boards) verify enumerate/validate agreement in both directions on every push
T6 judge, never simulate enforced in CI determinism test plus a purity lint: no randomness anywhere in the kernel, no network or subprocess in the verdict path
T7 mechanism never knows the game enforced in CI kernel lint scans every kernel module for game vocabulary (it caught a real violation during development)
T8 honest boundaries enforced in CI refusal goldens: unknown content, unmodeled conditions, and genuinely ambiguous rules text all return exit 2 with citations
T9 never a single number enforced in CI bench scorecard freshness test; per-category tables, no aggregate score exists anywhere in this repository — even the leaderboard ranks by wrong-count alone (T1) and shows the other failure modes unblended
T10 stranger-agent bootstrap enforced in CI cold-start conformance test reaches a first verdict from tool.json/--schema/MCP alone (20 tools); live probe: a frontier model given only tool.json produced a correct first verdict in 1 attempt(s), 3.9s (2026-07-16)
T11 table speed enforced in CI p95 latency budget test: 100 verdicts must stay under 100 ms at p95 (typically sub-millisecond)
T12 never sell what the model has held in review a strategy invariant: features pitched on knowledge parity are cut in review — enforced by humans and admitted as such
T13 the benchmark is a product shipped bench/ publishes 5 sets across 5 subjects with cited gold verdicts, incl. a per-horizon drift lane whose golds are engine-derived; a generated LEADERBOARD ranks any subject; cmd: driver + validate command let a third party submit a tamper-checked result
T14 every state has a lineage enforced in CI event.apply reducer stamps every transition (predecessor hash, causing event, rule ids, rule-vs-ruling kind); tests cover replay verification, tamper detection, the schema minimality ratchet, and reducer/validator agreement; demo replays 15 rounds hash-for-hash

Licensing

  • Code: MIT.
  • data/: includes material derived from the System Reference Document 5.2.1 under CC-BY-4.0 — see srdcheck/adapters/srd-5.2.1/sources/README.md for provenance and the required attribution.
  • srdcheck is unofficial and is not affiliated with or endorsed by Wizards of the Coast.

vs alternatives

  • Just asking the LLM — frontier models know these rules nearly perfectly (we measured it and killed half our own idea). What they can't do by construction: prove a verdict, reproduce it, or refuse when the question is outside the rules' jurisdiction — they claimed rules authority in the discretion zone in 5 of 6 demo runs. srdcheck sells proof, refusal, determinism, and economy — never knowledge.
  • Lookup APIs and reference MCP servers (Open5e, 5e-srd-api and their wrappers) — they answer "what does the book say," not "is this legal given this state." Complementary, not competing; srdcheck cites the same text they serve.
  • Closed engines inside AI-DM products — the strongest products in this vertical run deterministic rules layers, each rebuilt privately, uncited, unbenchmarked, and locked to their platform. srdcheck is that layer open, cited, benchmarked, and embeddable — including in theirs.
  • Simulator libraries (combat engines, character libraries) — they run games and own state. srdcheck judges and owns nothing (state travels with the query, stamped with lineage), which is exactly what makes it embeddable anywhere.

Prior art

srdcheck stands on lessons from Temple of Elemental Evil / Temple+ (dispatcher architecture), PCGen (prerequisite predicates), the FoundryVTT PF2e system (rules as data), Datasworn (official rules-as-JSON precedent), and FIREBALL (structured play state). Patterns were studied; no code was taken from any of them.

mcp-name: io.github.chaoz23/srdcheck

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

srdcheck-0.1.6.tar.gz (108.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

srdcheck-0.1.6-py3-none-any.whl (82.5 kB view details)

Uploaded Python 3

File details

Details for the file srdcheck-0.1.6.tar.gz.

File metadata

  • Download URL: srdcheck-0.1.6.tar.gz
  • Upload date:
  • Size: 108.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for srdcheck-0.1.6.tar.gz
Algorithm Hash digest
SHA256 60ddf9b98b28f15e392c2e57882dcb72e747c021a9da7a061d4b330f5d2c412a
MD5 180b1f0f273f5e0f947538dea179a651
BLAKE2b-256 5976575aa09928dff335b2f995dc5edd3d0f8d6a5469919824e0e4fcb37d0663

See more details on using hashes here.

File details

Details for the file srdcheck-0.1.6-py3-none-any.whl.

File metadata

  • Download URL: srdcheck-0.1.6-py3-none-any.whl
  • Upload date:
  • Size: 82.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for srdcheck-0.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 26f5980216bf5155b8b881e2af2e7bc23fb93daa2a3709577672d1aef3d11500
MD5 59d6bf6ed82cfcc9b91f722e4ee11cd7
BLAKE2b-256 2272bab7eba7acfc866044d097e69816a811d8121e7f4530b58220ea5977895f

See more details on using hashes here.

Release history Release notifications | RSS feed

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

This release

0.1.6 This release

2 files

0.1.5

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page