LLM Workflow Router
Deterministic workflow topology enforcement for LLM-powered systems.
LLM Workflow Router is a stateless middleware engine designed to enforce explicit execution topology in AI systems that rely on large language models. It evaluates structured interaction metadata against strictly declared workflow rules and returns a terminal state.
It controls structure — not content.
Overview
Modern LLM-driven systems frequently suffer from:
- Recursive tool invocation loops
- Circular container transitions
- Cross-container contamination
- Implicit fallback behavior
- Unbounded workflow escalation
- Inconsistent refusal logic
Most mitigation strategies limit volume (timeouts, max tool calls, retries).
LLM Workflow Router enforces topology explicitly.
The engine evaluates structured interaction metadata and returns one of three terminal states:
- PROCEED
- REFUSE
- PAUSE
No content inspection.
No moderation.
No orchestration.
No mutation of input.
Only structural enforcement.
Core Design Principles
- Deterministic evaluation
- Stateless per evaluation
- Metadata-only inspection
- Explicit transitions only
- Static configuration (v1)
- Strict validation at load time
- No silent fallback
- Host application retains execution control
Same input → same output.
Architecture
Application
↓
WorkflowEngine.evaluate(metadata)
↓
[ PROCEED | REFUSE | PAUSE ]
↓
Application decides next action
The router does not:
- Execute tools
- Retry calls
- Modify prompts
- Orchestrate sessions
It enforces topology and returns a decision.
Configuration
Workflow rules are defined using static YAML configuration.
entrypoints:
- entry
containers:
entry:
allow_transitions:
- support
- REFUSE
allow_reentry: false
max_invocations: 2
support:
allow_transitions:
- faq
- REFUSE
allow_reentry: false
max_invocations: 3
faq:
allow_transitions:
- REFUSE
allow_reentry: true
max_invocations: 5
Each container defines:
- Explicit allowed transitions
- Whether re-entry is permitted
- Maximum invocation depth
Implicit transitions are not allowed.
Entrypoints
The optional top-level entrypoints list declares the authoritative root
containers. A workflow may have several roots — a support flow, a sales flow,
and an FAQ flow can each start in their own container — so entrypoints is a
list. When declared, any indegree-0 container that is not listed is treated
as an orphan and rejected. If entrypoints is omitted, roots are inferred from
graph shape (indegree 0) for backward compatibility, and an ambiguity warning
is raised when more than one root is inferred. See
examples/multi_entry_config.yaml.
Transition budget
max_total_transitions caps a whole run, not just one container:
max_total_transitions: 6
Once transition_history holds that many containers, any further
container-to-container transition is refused with
MAX_TOTAL_TRANSITIONS_EXCEEDED. Terminal targets (PROCEED, REFUSE,
PAUSE) are always allowed, so a run that spent its budget can still end
cleanly. Without it, a cycle of re-entrant containers is bounded only by each
container's own max_invocations.
Approval gates
Mark a container requires_approval: true and entering it needs a human (or
any approver you choose):
refunds:
requires_approval: true
allow_transitions: [PROCEED, REFUSE]
A structurally valid transition into refunds returns PAUSE with reason
APPROVAL_REQUIRED until the host re-evaluates the same metadata with
approval_granted: true. The engine stays stateless: approval is an input, not
something it remembers. See examples/approval_config.yaml.
Strict parsing and editor support
Unknown keys are rejected at load time, so a typo such as max_invocation: 5
fails loudly instead of silently falling back to the default. For
autocompletion and inline errors while editing, export the JSON Schema and
point your editor at it:
llm-router schema > workflow.schema.json
# yaml-language-server: $schema=./workflow.schema.json
Topology Validation
At configuration load time, the router performs strict validation:
- Invalid transition targets
- Unknown containers
- Unknown declared entrypoints
- Dead-end containers
- Missing entry points
- Orphan containers (indegree 0, not a declared entrypoint)
- Unreachable containers
- Self-transition contradictions
- Cycle detection
- Re-entry safety enforcement
- Containers beyond the transition budget
- Approval gates on entrypoints (where they have no effect)
Configuration errors raise exceptions immediately.
Fail loudly at load time.
Never fail silently at runtime.
Runtime Evaluation
The engine evaluates an immutable metadata structure:
InteractionMetadata:
container: str
previous_state: InteractionState
transition_history: List[str]
invocation_depth: Dict[str, int]
requested_action: str
trace_id: Optional[str]
approval_granted: bool = False
Returns:
EvaluationResult:
state: InteractionState
container: str
reason: Optional[ReasonCode]
trace_id: Optional[str]
allowed_transitions: Tuple[str, ...]
No exceptions during normal evaluation.
Only terminal states are returned.
Every REFUSE says what would have been accepted:
allowed_transitions lists the targets the same metadata could have requested
(empty when the container itself is exhausted). An agent that took a wrong turn
can recover without guessing, and engine.allowed_targets(metadata) gives the
same list before anything is attempted.
Load a config from Python with the same strict checks the CLI uses:
from router import WorkflowEngine, load_config
engine = WorkflowEngine(load_config("config.yaml"))
CLI Usage
Install:
pip install llm-workflow-router
Validate configuration:
llm-router validate config.yaml
Analyze topology:
llm-router analyze config.yaml
Evaluate metadata:
llm-router run --config config.yaml --metadata metadata.json
Print the JSON Schema for configs:
llm-router schema
Review a config change. Every change is tagged WIDENS, NARROWS or
NEUTRAL, and newly reachable containers are listed. Classification is
conservative, so --fail-on-widen works as a CI gate (exit code 1):
llm-router diff old.yaml new.yaml --fail-on-widen
Replay recorded traffic against a new config and see every decision that would
now come out differently. A trace is the JSON Lines log that
run --verbose writes to stderr (or EvaluationLogEvent.from_result(...) from
your own code). Add --secure to replay through the Security Layer too:
llm-router run --config config.yaml --metadata md.json --verbose 2>> trace.jsonl
llm-router replay trace.jsonl --config new.yaml --fail-on-change
Exit codes: 0 success, 1 a diff or replay gate tripped, 2 the input
could not be read or validated.
Render the topology as a diagram:
llm-router graph config.yaml # Mermaid (paste into any Mermaid renderer)
llm-router graph config.yaml --format dot # Graphviz DOT
Like analyze, graph renders broken topologies too — unknown transition
targets are drawn and flagged, which is exactly what you want while debugging
a config.
Session Layer (optional)
The engine is stateless by design: every evaluate() receives a complete
metadata snapshot. If you'd rather not do that bookkeeping yourself,
WorkflowSession does it for you — and only advances on PROCEED:
from router import WorkflowEngine, WorkflowSession
engine = WorkflowEngine(cfg)
session = WorkflowSession(engine, entry="entry", trace_id="req-42")
result = session.request("support") # entry -> support
result = session.request("faq") # support -> faq
result = session.request("REFUSE") # terminal; session closes
session.history # ("entry", "support")
session.invocation_depth # {"entry": 1, "support": 1, "faq": 1}
A REFUSE closes the session. A PAUSE suspends it until resume().
session.snapshot(target) exposes the exact metadata the next request would
evaluate, so the session is fully auditable and you can drop down to raw
engine.evaluate() at any time. The engine itself remains pure and stateless.
At an approval gate the session pauses and holds the pending target:
result = session.request("refunds") # PAUSE, reason APPROVAL_REQUIRED
session.pending_approval # "refunds"
result = session.approve() # re-evaluates with approval; PROCEED
# or session.resume() to decline and stay where you are
The gate itself does not consume an invocation; the approved retry does.
OpenAI Agents SDK Integration
Agents are containers. Handoffs are transitions. Attach a TopologyGuard to
a run and every handoff is structurally evaluated before the next agent
executes — a refused handoff raises TopologyViolation and aborts the run
loudly instead of letting the agent graph wander:
pip install "llm-workflow-router[openai-agents]"
from agents import Agent, Runner
from router import WorkflowEngine
from router.integrations.openai_agents import TopologyGuard, TopologyViolation
guard = TopologyGuard(engine, entry="triage", trace_id="run-001")
try:
result = await Runner.run(triage_agent, "I was double-charged.", hooks=guard)
except TopologyViolation as violation:
print("Refused:", violation.result.reason)
# Full structural audit trail of the run:
print(guard.session.history, guard.session.invocation_depth)
By default an agent's name is its container name; pass container_for= to
map differently. For handoffs into requires_approval containers, pass
approver=, a sync or async (source, target) -> bool. Without one, or when it
returns False, the handoff raises TopologyViolation. The guard is content-blind — it never reads prompts,
messages, or tool arguments. See
examples/openai_agents_example.py for a
complete runnable triage → billing → refunds system.
Tool Gating: Claude Agent SDK and MCP
For single-agent systems the risky structure is usually the sequence of tool calls, not handoffs. Here tools are containers and each tool call is a transition from the tool called before it:
entrypoints: [start]
containers:
start: { allow_transitions: [search], max_invocations: 3 }
search: { allow_transitions: [search, summarize], allow_reentry: true, max_invocations: 5 }
summarize: { allow_transitions: [send_email], max_invocations: 2 }
send_email: { allow_transitions: [PROCEED], requires_approval: true }
"Never send an email before summarizing, and never without a human" is now a
contract rather than a prompt instruction. A refused call does not crash
the run: the model is told which tools it may call next and can recover, and
repeated bad calls exhaust max_invocations, so recovery is bounded. Every
tool must be declared or listed in passthrough, and a tool whose name maps to
PROCEED, REFUSE or PAUSE is refused, since those are workflow states;
nothing is waved through silently. Calls in one conversation are decided one at
a time, so a call waiting on an approver holds the calls behind it until the
answer is in. The gate reads tool names only, never arguments or results.
Claude Agent SDK (pip install "llm-workflow-router[claude-agent-sdk]"):
from claude_agent_sdk import ClaudeAgentOptions, query
from router.integrations.claude_agent_sdk import TopologyHooks
topology = TopologyHooks(engine, entry="start", passthrough={"Read", "Glob"})
options = ClaudeAgentOptions(hooks=topology.hooks())
Refused calls are denied through PreToolUse with a reason naming the
allowed tools. Permitted calls return no decision, so your normal permission
rules still apply. Each sub-agent gets its own topology session (keyed by
agent_id, with entry_for= to start sub-agent types elsewhere), so parallel
sub-agents never tangle. At an approval gate, pass approver= to decide
in-process, or leave it out and the gate defers to the SDK's own permission
prompt.
MCP (pip install "llm-workflow-router[mcp]"):
from router.integrations.mcp import GatedClientSession
session = GatedClientSession(client_session, engine, entry="start")
result = await session.call_tool("search", {"q": "..."})
A refused call never reaches the server. It comes back as a normal
CallToolResult with isError set, which is how MCP reports tool failures to
a model, so your agent loop needs no special handling. Everything besides
call_tool is passed through to the wrapped session.
Both adapters are thin layers over router.integrations.tool_gate.ToolGate,
which you can call directly from any other framework. The gate keeps a session
per conversation key until told otherwise, so a long-running host should call
gate.forget(key) when a conversation ends (each adapter exposes its gate as
.gate).
LangGraph Integration
Nodes are containers; running a node is a transition. LangGraph's edges
say where a graph may go. The guard enforces your declared topology
independently, so a routing function or Command(goto=...) that strays
outside the contract raises TopologyViolation instead of running the node:
pip install "llm-workflow-router[langgraph]"
from router.integrations.langgraph import TopologyGuard
guard = TopologyGuard(engine, entry="triage")
builder.add_node("triage", guard.node("triage", triage))
builder.add_node("refunds", guard.node("refunds", refunds))
Approval gates use LangGraph's own human-in-the-loop: the guard calls
interrupt() with {"type": "wfrouter.approval_required", "source": ..., "target": ...}, and you resume with Command(resume=True) to approve or
Command(resume=False) to decline (a checkpointer is required, as for any
interrupt). A gated node can still call interrupt() itself, and its own
questions receive their own answers.
Sessions are kept per thread_id, in the guard's memory, and one session
covers one run of the graph. LangGraph starts every new invocation with fresh
input at START, so on a thread that has run before, call start_run first.
Resuming with Command(resume=...) continues the current run and needs no call:
config = {"configurable": {"thread_id": "chat-1"}}
guard.start_run("chat-1") # new input on a thread that has run before
graph.invoke(new_input, config)
Without it, the entry node is refused as a transition from wherever the last
run ended, and the error names the call to make. Call guard.forget(thread_id)
when a conversation ends, so a long-running host does not keep every thread in
memory. The guard follows one node at a time, so graphs that fan out to
parallel branches in a single step are outside its model.
Why not just LangGraph (or my orchestrator's built-in graph)?
Orchestrators describe structure. This engine enforces it — as a separate, framework-agnostic layer with properties orchestrators don't give you:
- Independent enforcement. The topology lives outside your agent framework, so a prompt-induced detour, a buggy handoff, or a framework upgrade can't silently widen what's reachable. The declared graph is a contract, and violations fail loudly with structured reason codes.
- Framework-agnostic. The same YAML config governs an OpenAI Agents SDK app today and whatever you migrate to next year. Adapters are thin; the contract is portable.
- Auditable determinism. Same metadata + same config → same decision,
every time. Combined with the
wfrouter.*OpenTelemetry conventions, you get compliance-grade answers to "why was this transition refused?" — a reason code, not a vibe. - Load-time topology analysis. Cycles, orphans, dead ends, and unreachable states are caught before anything runs — the kind of static validation industrial control systems have had for decades and agent frameworks mostly don't.
If you're happy inside one orchestrator and don't need independent structural guarantees, its built-in graph may be enough. This tool exists for when "probably follows the graph" isn't good enough.
Intended Audience
- AI SaaS developers
- Internal LLM tooling teams
- Agent orchestration builders
- Platform engineering teams
- Infrastructure-focused AI developers
Not intended for content moderation or prompt filtering.
Observability
Optional OpenTelemetry instrumentation is provided under a dedicated,
versioned namespace (wfrouter.*) that this project owns — it is deliberately
independent of the upstream gen_ai.* conventions, which assume a model at the
center of every span and do not fit a content-blind topology engine.
Install the extra:
pip install "llm-workflow-router[otel]"
Wrap evaluation:
from router.observability.otel import traced_evaluate
result = traced_evaluate(engine, metadata)
If opentelemetry-api is not installed, instrumentation degrades to a no-op
and the engine behaves identically. A PROCEED, REFUSE, or PAUSE outcome
is a successful decision (span status OK); only genuine failures are errors.
See OBSERVABILITY.md for the full attribute and span conventions.
Security Layer (router.security) — commercial
An optional layer that secures the structure of a workflow: it enforces which
trust levels may reach which side-effecting capabilities, emits
wfrouter.security.* telemetry, and keeps a tamper-evident audit trail. Same
content-blind, deterministic principles as the core.
Declare trust posture per container, then check it:
containers:
intake: { trust: UNTRUSTED, allow_transitions: [guard] }
guard: { sanitizer: true, allow_transitions: [tool] }
tool: { capability: TOOL_EXEC, allow_transitions: [PROCEED] }
llm-router secure config.yaml # reports trust-boundary issues
from router.engine import WorkflowEngine
from router.security import SecurityEngine, AuditLog
secure = SecurityEngine(WorkflowEngine(cfg), audit_log=AuditLog())
result = secure.evaluate(metadata) # REFUSE if a boundary is crossed
assert secure.audit_log.verify()
The core property: untrusted input must pass a sanitizer before it can reach a side-effecting capability. See SECURITY.md for the full model, findings, and telemetry conventions.
Licensing: the whole project is PolyForm Noncommercial 1.0.0 — free to inspect, run, and use for any noncommercial purpose, commercial use requires a paid license. See LICENSING.md and PRICING.md.
License
Source-available under PolyForm Noncommercial 1.0.0. The whole project — core engine and Security Layer alike — is free to inspect, run, modify, and use for any noncommercial purpose: personal projects, research, education, non-profits, and evaluation.
Commercial use requires a paid license. That means using any part of it in a product or service you sell or host, inside a for-profit company's production or internal systems, or in paid client work. See LICENSING.md, COMMERCIAL-LICENSE.md, and PRICING.md.
If you deploy this in production, a note about your use case is always appreciated (and helps prioritize the roadmap) — but never required.
Version
Current version: 1.1.1
- Static configuration model
- Explicit multi-entrypoint declaration (with inferred fallback)
- Whole-run transition budget and human approval gates
- Refusals that name the allowed next steps
- Optional stateful
WorkflowSessionconvenience layer - Adapters: OpenAI Agents SDK, Claude Agent SDK, LangGraph, MCP
- Mermaid / DOT topology export (
llm-router graph) - Change review:
llm-router diffandllm-router replay - JSON Schema for configs (
llm-router schema) - Security Layer: trust-boundary enforcement,
wfrouter.security.*telemetry, tamper-evident audit trail (commercial) - Licensed: PolyForm Noncommercial 1.0.0 (noncommercial free · commercial licensed)
No inheritance. No dynamic rule composition.
Future versions may extend topology modeling capabilities.
Author
Doby Baxter
Software systems developer focused on deterministic infrastructure and human-centered tooling.
Metadata
Release files for llm-workflow-router 1.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_workflow_router-1.1.1.tar.gz | 91.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_workflow_router-1.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 167.2 kB
Release files / llm_workflow_router-1.1.1.tar.gz
| Download URL | llm_workflow_router-1.1.1.tar.gz |
|---|---|
| Size | 91.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f6739e2f78a81136a736c284b1d4ed91940ee3907e3d8a31ac6e328c15f4bd2e
|
|
BLAKE2b-256 checksum How to use checksums |
a5f62a49893b2a9921060ae83c4a1f6352c90d79ddbda357182980d207379b77
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.8
|
Release files / llm_workflow_router-1.1.1-py3-none-any.whl
| Download URL | llm_workflow_router-1.1.1-py3-none-any.whl |
|---|---|
| Size | 76.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
fee0214a22df4f166a91cd3238c8b9bbf657b80ad338cb50e9992a00dcfec334
|
|
BLAKE2b-256 checksum How to use checksums |
83c9368eb30fbb7d7eddf70f068651127885823b108e66495f2ad345119e5765
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.8
|