InferRouter: loop-aware router and tracing for coding agents (Claude Code, Codex). Routes workflow steps and records cost/cache/verification/outcome receipts.
Reason this release was yanked:
Superseded by 0.1.1. 0.1.0 was an accidental build with the wrong internal module name (import inferrouter failed) and a missing PyYAML dependency.
Project description
Touchdown InferGuard InferRoute MVP
Status: offline-demoable MVP with a live-capable GMI serverless replay path
Primary metric: cost per accepted patch
Primary wedge: Fable 5 / Claude Code / Codex / Cursor routing, cache, backend, and trace optimization
InferGuard plugs into a coding-agent workflow, records model/tool/context/cache/patch events, uses InferRoute for each workflow step, and outputs investor-demo artifacts:
coding-agent workflow
→ trace
→ phase classifier
→ InferRoute static policy
→ InferRoute tiny-router shadow hook
→ cache-aware routing gate
→ backend allocator
→ GMI/GLM-5.2 serverless candidate replay
→ backend promotion plan
→ cost per accepted patch report
→ trace eligibility summary
→ post-training and endpoint handoff artifacts
This is not a generic proxy and not a model-hosting wrapper. It starts API/serverless-first and promotes traffic to rented/dedicated infrastructure only when traces prove utilization and quality.
Quick demo (no API keys required)
cd inferguard
python -m td_inferguard.cli demo --out .td/demo
python -m checks.smoke_check
Outputs include:
report.md 14-section investor report
api_response.json routes + analysis + eval gates
SKILL.md generated coding-agent skill
routing.policy.yaml per-step route policy
backend_ladder.md API → serverless → rented → owned → WorkloadPod
backend_promotion_plan.yaml real gates computed from the trace (plan_only)
gmi_glm52_replay.json replay plan (plan-only in the demo, always)
serverless_to_rented_break_even.md
tiny_router_shadow.md
oracle_ceiling.md
distillation_eligibility.md
tinker_job.yaml
prime_intellect_train_manifest.yaml
gmi_sidekick_endpoint_plan.md
terra_backend_plan.md
CLI
python -m td_inferguard.cli init --target claude-code --provider gmi --out .td
python -m td_inferguard.cli skill write --target claude-code --out SKILL.md
python -m td_inferguard.cli demo --out .td/demo
python -m td_inferguard.cli route examples/sample_claude_code_run.jsonl --out .td/policies/routing.policy.yaml
python -m td_inferguard.cli report examples/sample_claude_code_run.jsonl --out .td/reports/inferguard.md
python -m td_inferguard.cli backend report --run examples/sample_claude_code_run.jsonl --out .td/reports/backend_ladder.md
python -m td_inferguard.cli backend promotion-plan --run examples/sample_claude_code_run.jsonl --out .td/backends/promotion_plan.yaml
python -m td_inferguard.cli replay gmi --run examples/sample_claude_code_run.jsonl --model zai-org/GLM-5.2-FP8 --out .td/replay/gmi_glm52_replay.json
python -m td_inferguard.cli live-check --out .td/live/live_check.json
python -m td_inferguard.cli serve --host 127.0.0.1 --port 8765
td-inferguard and td-api console scripts are available after pip install -e .; every command also works uninstalled via python -m td_inferguard.cli.
API endpoints
GET /health
GET /v1/version
GET /usage/summary?by=dispatch_path usage ledger summary
GET /usage/events?limit=20 recent UsageEvent rows
GET /usage/budget Fable 5 daily budget view
GET /usage/limits Claude Code official or estimated limit view
POST /v1/route {"event": {...}} or a bare event object
POST /v1/optimize/claude-code {"events": [...]} → analysis + routes + backend_summary + evals
POST /v1/backends/report {"events": [...]} → tier costs/counts + promotion plan
POST /v1/backends/promotion-plan {"events": [...]} → gate status from the trace
POST /v1/replay/gmi {"events": [...], "model": "zai-org/GLM-5.2-FP8"}
POST /v1/live-check {"frontier_model"?: "...", "sidekick_model"?: "..."} → real endpoint pings
GET /v1/providers/check dry configuration check, no requests
POST /v1/providers/check {"live": true} → real per-provider pings
POST /v1/loop/audit {"events": [...]} → loop audit markdown
POST /v1/loop/blindspots {"task": "..."} → blind-spot interview markdown
POST /v1/loop/merge-quiz {"events": [...]} → merge quiz markdown
POST /v1/loop/ir {"events": [...]} → LoopIR snapshot
POST /v1/dataset/export {"events": [...], "out"?: "dir"} → governed dataset export
POST /v1/workload/signature {"events": [...]} → workload signature
POST /v1/kernels/opportunity-map {"events": [...]} → kernel opportunity map (no speedup claims)
POST /v1/ingest/weka {"dataset"|"source", "limit"} → Weka traces → TraceEvents
POST /v1/ingest/isb1 {"dataset"|"source", "limit"} → ISB-1 traces → TraceEvents
POST /v1/cache/simulate {"events": [...]} → 64-token block cache report
POST /v1/benchmark/agentic-replay {"events": [...], "profile": "..."} → replay report
POST /v1/benchmark/compare-real-vs-controlled {"real_events": [...], "controlled_events": [...]}
POST /v1/reports/generate {"events": [...], "out"?: "dir"} → 9 operator reports
Usage and cost tracing
Cross-vendor coding-agent usage is recorded as append-only UsageEvent rows in
inferguard/out/traces/usage-ledger.jsonl or TD_USAGE_LEDGER. The full
runbook is docs/inferguard/33-2026-07-07-cross-vendor-usage-cost-tracing.md.
From the repo root, the two most useful commands are:
PYTHONPATH=inferguard python3 -m td_inferguard.cli usage run --dispatch-path codex --task-id TASK --trace-id TRACE -- codex exec --json "reply OK"
PYTHONPATH=inferguard python3 -m td_inferguard.cli usage import-claude
PYTHONPATH=inferguard python3 -m td_inferguard.cli usage report --by dispatch_path --json
Pricing follows the registry discipline: runtime-reported cost first, sourced registry prices second, optional local overrides third, unknown otherwise. Subscription rows are API-equivalent estimates, not billed dollars, and reports keep them separate from actual API cost.
Signal Plane (InferRoute semantic layer)
Every route builds a ContextPack (phase/risk/cache/history/privacy; prompt
hash + 200-char summary, never the raw prompt) and a Signal Plane readout
(heuristic/policy/cache/modality/semantic signals). Routing picks a role
first, then resolves a model alias from tdroute.yaml (model_aliases:).
td-inferguard eval routes --config tdroute.yaml --out out/reports/route_eval_report.md # 9 deterministic policy cases, no live calls
td-inferguard mcp start --port 8790 # JSON-RPC tool server: inferguard.route_event, explain_route, build_context_pack, ...
td-inferguard route-receipt --run out/traces/proxy-events.jsonl --out out/route-receipt.json
Hook contracts (td_inferguard/hooks.py) give in-process agents pre/post
model-call, tool, test, patch, compaction, and human-decision integration.
All routing answers are projected; dollar figures in receipts are projected
from fixed cost bands until measured provider responses exist.
Agentic benchmarking (real vs controlled)
td-inferguard ingest hf-weka --dataset semianalysisai/cc-traces-weka-062126 --split train --limit 10 --out out/traces/weka.jsonl
td-inferguard ingest isb1 --dataset wchen22/isb1-cc-traces --limit 10 --out out/traces/isb1.jsonl
td-inferguard cache simulate --run out/traces/weka.jsonl --out out/cache/weka.json
td-inferguard benchmark agentic-replay --run out/traces/weka.jsonl --profile with_subagents --out out/reports/agentic-replay.md --json out/reports/agentic-replay.json
td-inferguard benchmark compare-real-vs-controlled --real out/traces/weka.jsonl --controlled out/traces/isb1.jsonl --out out/reports/real-vs-controlled.md
td-inferguard receipts generate --run out/traces/weka.jsonl --out receipts/
Weka (SemiAnalysis) = real Claude Code production distribution; ISB-1
(wchen22/isb1-cc-traces) = controlled KV-cache stress. The source registry
(td_inferguard/datasets/registry.py, generated source_registry.yaml)
enforces lane separation: Weka never becomes a correctness eval, ISB-1 never
poses as real distribution, SWE-bench stays eval-only, execution-free
trajectories are never test-passing evidence. Mechanism quality gates
(mechanisms.py) refuse supported for lossy optimizations without a
completed quality eval. Replay-only traces prove runtime realism, not
accepted-patch quality.
Evidence Spine
td_inferguard/evidence/ is the append-only JSONL trace layer (validated event
types, claim_status, trace_eligibility, redaction_status). td_inferguard/datasets/
exports governed training data: routing/eval rows by default, SFT/RL rows only
when the trace carries explicit sft_allowed/rl_allowed rights; secrets are
redacted and every row carries its source_bucket so measured, synthetic, and
plan-only data never mix silently. td_inferguard/workloads.py turns traces
into workload signatures and a kernel opportunity map (measured bottlenecks vs
projected opportunities vs plan-only NVIDIA/AMD futures, never a speedup claim).
td-inferguard route-event --input event.json --out route.json
td-inferguard providers check [--live] [--out providers.json]
td-inferguard replay gmi [--live] --run <trace> --out <file> # live is opt-in
td-inferguard replay openai-compatible [--provider openai|openrouter] [--live] --run <trace> --out <file>
td-inferguard dataset export --run <trace> --out out/dataset
td-inferguard workload signature --run <trace> --out sig.json
td-inferguard kernels opportunity-map --run <trace> --out kmap.json
td-inferguard loop ir --run <trace> --out loop_ir.json
td-inferguard reports --run <trace> --out out/reports
Provider env vars: ANTHROPIC_API_KEY/ANTHROPIC_BASE_URL/TD_FRONTIER_MODEL,
GMI_API_KEY/GMI_BASE_URL/GMI_MODEL (or TD_SIDEKICK_MODEL),
OPENAI_API_KEY/OPENAI_BASE_URL/OPENAI_MODEL,
OPENROUTER_API_KEY/OPENROUTER_BASE_URL/OPENROUTER_MODEL. Replay never
spends money without --live; --live without a key writes a blocked
artifact instead of pretending.
Live endpoint check
export ANTHROPIC_API_KEY="sk-ant-..." # frontier lane (Fable 5 via api.anthropic.com)
export GMI_API_KEY="..." # sidekick lane (GLM-5.2 via api.gmi-serving.com)
python -m td_inferguard.cli live-check --out .td/live/live_check.json
The CLI also reads ignored local .env files from the current directory and the InferGuard package directory. Exported environment variables win over .env.
live-check sends one real, tiny request (max_tokens=16) per configured provider and writes a measured artifact with latency and token usage straight off the wire. Nothing is simulated: an unconfigured lane is recorded not_configured (no request made), a failed request is recorded error with the provider's actual response. Model defaults come from TD_FRONTIER_MODEL (default claude-fable-5) and TD_SIDEKICK_MODEL (default zai-org/GLM-5.2-FP8); nothing is hardcoded beyond those documented, overridable defaults.
GMI serverless replay
export GMI_API_KEY="your key" # required for live replay
export GMI_BASE_URL="https://api.gmi-serving.com/v1" # optional default
python -m td_inferguard.cli replay gmi --run <trace> --model zai-org/GLM-5.2-FP8 --out <file>
For local runs, GMI_API_KEY, GMI_BASE_URL, and TD_SIDEKICK_MODEL can also live in ignored .env.
Without the key the command writes a plan-only artifact and makes no requests. With the key it replays only sidekick-eligible phases (repo search, file/test-log summaries, docs formatting, low-risk first-fail repair) and labels results candidate, never adopt, until quality evals pass. High-risk patching and final review never replay by default. GLM-5.2 is a sidekick candidate; it does not replace Fable.
The provider layer (td_inferguard/providers/) is stdlib-only: an OpenAI-compatible client, a GMI subclass, and an Anthropic Messages API client for the frontier lane (minimal payload so the same request shape is valid on Fable 5, which rejects sampling and thinking params). Provider errors come back as structured objects, never exceptions.
Tests
python -m pip install pytest # or: pip install -e ".[dev]"
python -m pytest
Covers the CLI, demo outputs, phase classifier, InferRoute static policy, cache-switch gate, backend allocator, promotion gates, cost-per-accepted-patch analysis, the 14-section report, the eval gates, plan-only replay, and every API endpoint.
Docs
docs/production-runbook.md: offline demo, API server, GMI configuration, replay, real-trace import, report reading, and what not to claim yet.../docs/inferguard/32-2026-07-06-inferroute-architecture-and-status.md: canonical InferRoute architecture, policy gates, backend tiers, realtime sidecar, receipts, and honest status.
Honesty labels
Artifacts carry measured, estimated, projected, shadow_only, candidate, plan_only, adopt, reject, or inconclusive labels. The demo does not claim live Fable savings, live Terra integration, live GMI deployment, a trained GLM-5.2/Qwen sidekick, or executed training jobs. InferRoute's tiny-router shadow is shadow-only; the static policy is what routes.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file inferrouter-0.1.0.tar.gz.
File metadata
- Download URL: inferrouter-0.1.0.tar.gz
- Upload date:
- Size: 340.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4cc7ca0ae5f4889cfbf6f4f123649f071e018dc84081d0a0747bf6885f1f0c33
|
|
| MD5 |
89e1fb71a8a02c666269b2ebfb295b22
|
|
| BLAKE2b-256 |
f2c67e7f113ca0dbc7677953f88fd7c713d3c61a2c50a58cf8c0796772633181
|
Provenance
The following attestation bundles were made for inferrouter-0.1.0.tar.gz:
Publisher:
inferguard-package-publish.yml on OCWC22/Touchdown-Labs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
inferrouter-0.1.0.tar.gz -
Subject digest:
4cc7ca0ae5f4889cfbf6f4f123649f071e018dc84081d0a0747bf6885f1f0c33 - Sigstore transparency entry: 2117612799
- Sigstore integration time:
-
Permalink:
OCWC22/Touchdown-Labs@97a97fae378ef3a2656c3cd62b97470bb868166a -
Branch / Tag:
refs/heads/main - Owner: https://github.com/OCWC22
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
inferguard-package-publish.yml@97a97fae378ef3a2656c3cd62b97470bb868166a -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file inferrouter-0.1.0-py3-none-any.whl.
File metadata
- Download URL: inferrouter-0.1.0-py3-none-any.whl
- Upload date:
- Size: 348.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
76491f90b61c31153a8e1bcc8465bc652d9a35304034ca4ee0f089b2d6e330b6
|
|
| MD5 |
f2caf269eacf6e00a14e326dbc339366
|
|
| BLAKE2b-256 |
e40b015e9d177002933b2da8531116fcefb37e8e879c0152f280c9925a212295
|
Provenance
The following attestation bundles were made for inferrouter-0.1.0-py3-none-any.whl:
Publisher:
inferguard-package-publish.yml on OCWC22/Touchdown-Labs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
inferrouter-0.1.0-py3-none-any.whl -
Subject digest:
76491f90b61c31153a8e1bcc8465bc652d9a35304034ca4ee0f089b2d6e330b6 - Sigstore transparency entry: 2117612866
- Sigstore integration time:
-
Permalink:
OCWC22/Touchdown-Labs@97a97fae378ef3a2656c3cd62b97470bb868166a -
Branch / Tag:
refs/heads/main - Owner: https://github.com/OCWC22
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
inferguard-package-publish.yml@97a97fae378ef3a2656c3cd62b97470bb868166a -
Trigger Event:
workflow_dispatch
-
Statement type: