Skip to main content

LLM Workflow Router

Deterministic workflow topology enforcement for LLM-powered systems.

LLM Workflow Router is a stateless middleware engine designed to enforce explicit execution topology in AI systems that rely on large language models. It evaluates structured interaction metadata against strictly declared workflow rules and returns a terminal state.

It controls structure — not content.


Overview

Modern LLM-driven systems frequently suffer from:

  • Recursive tool invocation loops
  • Circular container transitions
  • Cross-container contamination
  • Implicit fallback behavior
  • Unbounded workflow escalation
  • Inconsistent refusal logic

Most mitigation strategies limit volume (timeouts, max tool calls, retries).
LLM Workflow Router enforces topology explicitly.

The engine evaluates structured interaction metadata and returns one of three terminal states:

  • PROCEED
  • REFUSE
  • PAUSE

No content inspection.
No moderation.
No orchestration.
No mutation of input.

Only structural enforcement.


Core Design Principles

  • Deterministic evaluation
  • Stateless per evaluation
  • Metadata-only inspection
  • Explicit transitions only
  • Static configuration (v1)
  • Strict validation at load time
  • No silent fallback
  • Host application retains execution control

Same input → same output.


Architecture

Application
↓
WorkflowEngine.evaluate(metadata)
↓
[ PROCEED | REFUSE | PAUSE ]
↓
Application decides next action

The router does not:

  • Execute tools
  • Retry calls
  • Modify prompts
  • Orchestrate sessions

It enforces topology and returns a decision.


Configuration

Workflow rules are defined using static YAML configuration.

entrypoints:
  - entry

containers:
  entry:
    allow_transitions:
      - support
      - REFUSE
    allow_reentry: false
    max_invocations: 2

  support:
    allow_transitions:
      - faq
      - REFUSE
    allow_reentry: false
    max_invocations: 3

  faq:
    allow_transitions:
      - REFUSE
    allow_reentry: true
    max_invocations: 5

Each container defines:

  • Explicit allowed transitions
  • Whether re-entry is permitted
  • Maximum invocation depth

Implicit transitions are not allowed.

Entrypoints

The optional top-level entrypoints list declares the authoritative root containers. A workflow may have several roots — a support flow, a sales flow, and an FAQ flow can each start in their own container — so entrypoints is a list. When declared, any indegree-0 container that is not listed is treated as an orphan and rejected. If entrypoints is omitted, roots are inferred from graph shape (indegree 0) for backward compatibility, and an ambiguity warning is raised when more than one root is inferred. See examples/multi_entry_config.yaml.

Transition budget

max_total_transitions caps a whole run, not just one container:

max_total_transitions: 6

Once transition_history holds that many containers, any further container-to-container transition is refused with MAX_TOTAL_TRANSITIONS_EXCEEDED. Terminal targets (PROCEED, REFUSE, PAUSE) are always allowed, so a run that spent its budget can still end cleanly. Without it, a cycle of re-entrant containers is bounded only by each container's own max_invocations.

Approval gates

Mark a container requires_approval: true and entering it needs a human (or any approver you choose):

  refunds:
    requires_approval: true
    allow_transitions: [PROCEED, REFUSE]

A structurally valid transition into refunds returns PAUSE with reason APPROVAL_REQUIRED until the host re-evaluates the same metadata with approval_granted: true. The engine stays stateless: approval is an input, not something it remembers. See examples/approval_config.yaml.

Strict parsing and editor support

Unknown keys are rejected at load time, so a typo such as max_invocation: 5 fails loudly instead of silently falling back to the default. For autocompletion and inline errors while editing, export the JSON Schema and point your editor at it:

llm-router schema > workflow.schema.json
# yaml-language-server: $schema=./workflow.schema.json

Topology Validation

At configuration load time, the router performs strict validation:

  • Invalid transition targets
  • Unknown containers
  • Unknown declared entrypoints
  • Dead-end containers
  • Missing entry points
  • Orphan containers (indegree 0, not a declared entrypoint)
  • Unreachable containers
  • Self-transition contradictions
  • Cycle detection
  • Re-entry safety enforcement
  • Containers beyond the transition budget
  • Approval gates on entrypoints (where they have no effect)

Configuration errors raise exceptions immediately.

Fail loudly at load time.
Never fail silently at runtime.


Runtime Evaluation

The engine evaluates an immutable metadata structure:

InteractionMetadata:
    container: str
    previous_state: InteractionState
    transition_history: List[str]
    invocation_depth: Dict[str, int]
    requested_action: str
    trace_id: Optional[str]
    approval_granted: bool = False

Returns:

EvaluationResult:
    state: InteractionState
    container: str
    reason: Optional[ReasonCode]
    trace_id: Optional[str]
    allowed_transitions: Tuple[str, ...]

No exceptions during normal evaluation.
Only terminal states are returned.

Every REFUSE says what would have been accepted: allowed_transitions lists the targets the same metadata could have requested (empty when the container itself is exhausted). An agent that took a wrong turn can recover without guessing, and engine.allowed_targets(metadata) gives the same list before anything is attempted.

Load a config from Python with the same strict checks the CLI uses:

from router import WorkflowEngine, load_config

engine = WorkflowEngine(load_config("config.yaml"))

CLI Usage

Install:

pip install llm-workflow-router

Validate configuration:

llm-router validate config.yaml

Analyze topology:

llm-router analyze config.yaml

Evaluate metadata:

llm-router run --config config.yaml --metadata metadata.json

Print the JSON Schema for configs:

llm-router schema

Review a config change. Every change is tagged WIDENS, NARROWS or NEUTRAL, and newly reachable containers are listed. Classification is conservative, so --fail-on-widen works as a CI gate (exit code 1):

llm-router diff old.yaml new.yaml --fail-on-widen

Replay recorded traffic against a new config and see every decision that would now come out differently. A trace is the JSON Lines log that run --verbose writes to stderr (or EvaluationLogEvent.from_result(...) from your own code). Add --secure to replay through the Security Layer too:

llm-router run --config config.yaml --metadata md.json --verbose 2>> trace.jsonl
llm-router replay trace.jsonl --config new.yaml --fail-on-change

Exit codes: 0 success, 1 a diff or replay gate tripped, 2 the input could not be read or validated.

Render the topology as a diagram:

llm-router graph config.yaml                # Mermaid (paste into any Mermaid renderer)
llm-router graph config.yaml --format dot   # Graphviz DOT

Like analyze, graph renders broken topologies too — unknown transition targets are drawn and flagged, which is exactly what you want while debugging a config.


Session Layer (optional)

The engine is stateless by design: every evaluate() receives a complete metadata snapshot. If you'd rather not do that bookkeeping yourself, WorkflowSession does it for you — and only advances on PROCEED:

from router import WorkflowEngine, WorkflowSession

engine = WorkflowEngine(cfg)
session = WorkflowSession(engine, entry="entry", trace_id="req-42")

result = session.request("support")   # entry -> support
result = session.request("faq")       # support -> faq
result = session.request("REFUSE")    # terminal; session closes

session.history           # ("entry", "support")
session.invocation_depth  # {"entry": 1, "support": 1, "faq": 1}

A REFUSE closes the session. A PAUSE suspends it until resume(). session.snapshot(target) exposes the exact metadata the next request would evaluate, so the session is fully auditable and you can drop down to raw engine.evaluate() at any time. The engine itself remains pure and stateless.

At an approval gate the session pauses and holds the pending target:

result = session.request("refunds")   # PAUSE, reason APPROVAL_REQUIRED
session.pending_approval               # "refunds"
result = session.approve()             # re-evaluates with approval; PROCEED
# or session.resume() to decline and stay where you are

The gate itself does not consume an invocation; the approved retry does.


OpenAI Agents SDK Integration

Agents are containers. Handoffs are transitions. Attach a TopologyGuard to a run and every handoff is structurally evaluated before the next agent executes — a refused handoff raises TopologyViolation and aborts the run loudly instead of letting the agent graph wander:

pip install "llm-workflow-router[openai-agents]"
from agents import Agent, Runner
from router import WorkflowEngine
from router.integrations.openai_agents import TopologyGuard, TopologyViolation

guard = TopologyGuard(engine, entry="triage", trace_id="run-001")

try:
    result = await Runner.run(triage_agent, "I was double-charged.", hooks=guard)
except TopologyViolation as violation:
    print("Refused:", violation.result.reason)

# Full structural audit trail of the run:
print(guard.session.history, guard.session.invocation_depth)

By default an agent's name is its container name; pass container_for= to map differently. For handoffs into requires_approval containers, pass approver=, a sync or async (source, target) -> bool. Without one, or when it returns False, the handoff raises TopologyViolation. The guard is content-blind — it never reads prompts, messages, or tool arguments. See examples/openai_agents_example.py for a complete runnable triage → billing → refunds system.


Tool Gating: Claude Agent SDK and MCP

For single-agent systems the risky structure is usually the sequence of tool calls, not handoffs. Here tools are containers and each tool call is a transition from the tool called before it:

entrypoints: [start]
containers:
  start:      { allow_transitions: [search], max_invocations: 3 }
  search:     { allow_transitions: [search, summarize], allow_reentry: true, max_invocations: 5 }
  summarize:  { allow_transitions: [send_email], max_invocations: 2 }
  send_email: { allow_transitions: [PROCEED], requires_approval: true }

"Never send an email before summarizing, and never without a human" is now a contract rather than a prompt instruction. A refused call does not crash the run: the model is told which tools it may call next and can recover, and repeated bad calls exhaust max_invocations, so recovery is bounded. Every tool must be declared or listed in passthrough, and a tool whose name maps to PROCEED, REFUSE or PAUSE is refused, since those are workflow states; nothing is waved through silently. Calls in one conversation are decided one at a time, so a call waiting on an approver holds the calls behind it until the answer is in. The gate reads tool names only, never arguments or results.

Claude Agent SDK (pip install "llm-workflow-router[claude-agent-sdk]"):

from claude_agent_sdk import ClaudeAgentOptions, query
from router.integrations.claude_agent_sdk import TopologyHooks

topology = TopologyHooks(engine, entry="start", passthrough={"Read", "Glob"})
options = ClaudeAgentOptions(hooks=topology.hooks())

Refused calls are denied through PreToolUse with a reason naming the allowed tools. Permitted calls return no decision, so your normal permission rules still apply. Each sub-agent gets its own topology session (keyed by agent_id, with entry_for= to start sub-agent types elsewhere), so parallel sub-agents never tangle. At an approval gate, pass approver= to decide in-process, or leave it out and the gate defers to the SDK's own permission prompt.

MCP (pip install "llm-workflow-router[mcp]"):

from router.integrations.mcp import GatedClientSession

session = GatedClientSession(client_session, engine, entry="start")
result = await session.call_tool("search", {"q": "..."})

A refused call never reaches the server. It comes back as a normal CallToolResult with isError set, which is how MCP reports tool failures to a model, so your agent loop needs no special handling. Everything besides call_tool is passed through to the wrapped session.

Both adapters are thin layers over router.integrations.tool_gate.ToolGate, which you can call directly from any other framework. The gate keeps a session per conversation key until told otherwise, so a long-running host should call gate.forget(key) when a conversation ends (each adapter exposes its gate as .gate).


LangGraph Integration

Nodes are containers; running a node is a transition. LangGraph's edges say where a graph may go. The guard enforces your declared topology independently, so a routing function or Command(goto=...) that strays outside the contract raises TopologyViolation instead of running the node:

pip install "llm-workflow-router[langgraph]"
from router.integrations.langgraph import TopologyGuard

guard = TopologyGuard(engine, entry="triage")
builder.add_node("triage", guard.node("triage", triage))
builder.add_node("refunds", guard.node("refunds", refunds))

Approval gates use LangGraph's own human-in-the-loop: the guard calls interrupt() with {"type": "wfrouter.approval_required", "source": ..., "target": ...}, and you resume with Command(resume=True) to approve or Command(resume=False) to decline (a checkpointer is required, as for any interrupt). A gated node can still call interrupt() itself, and its own questions receive their own answers.

Sessions are kept per thread_id, in the guard's memory, and one session covers one run of the graph. LangGraph starts every new invocation with fresh input at START, so on a thread that has run before, call start_run first. Resuming with Command(resume=...) continues the current run and needs no call:

config = {"configurable": {"thread_id": "chat-1"}}
guard.start_run("chat-1")  # new input on a thread that has run before
graph.invoke(new_input, config)

Without it, the entry node is refused as a transition from wherever the last run ended, and the error names the call to make. Call guard.forget(thread_id) when a conversation ends, so a long-running host does not keep every thread in memory. The guard follows one node at a time, so graphs that fan out to parallel branches in a single step are outside its model.


Why not just LangGraph (or my orchestrator's built-in graph)?

Orchestrators describe structure. This engine enforces it — as a separate, framework-agnostic layer with properties orchestrators don't give you:

  • Independent enforcement. The topology lives outside your agent framework, so a prompt-induced detour, a buggy handoff, or a framework upgrade can't silently widen what's reachable. The declared graph is a contract, and violations fail loudly with structured reason codes.
  • Framework-agnostic. The same YAML config governs an OpenAI Agents SDK app today and whatever you migrate to next year. Adapters are thin; the contract is portable.
  • Auditable determinism. Same metadata + same config → same decision, every time. Combined with the wfrouter.* OpenTelemetry conventions, you get compliance-grade answers to "why was this transition refused?" — a reason code, not a vibe.
  • Load-time topology analysis. Cycles, orphans, dead ends, and unreachable states are caught before anything runs — the kind of static validation industrial control systems have had for decades and agent frameworks mostly don't.

If you're happy inside one orchestrator and don't need independent structural guarantees, its built-in graph may be enough. This tool exists for when "probably follows the graph" isn't good enough.


Intended Audience

  • AI SaaS developers
  • Internal LLM tooling teams
  • Agent orchestration builders
  • Platform engineering teams
  • Infrastructure-focused AI developers

Not intended for content moderation or prompt filtering.


Observability

Optional OpenTelemetry instrumentation is provided under a dedicated, versioned namespace (wfrouter.*) that this project owns — it is deliberately independent of the upstream gen_ai.* conventions, which assume a model at the center of every span and do not fit a content-blind topology engine.

Install the extra:

pip install "llm-workflow-router[otel]"

Wrap evaluation:

from router.observability.otel import traced_evaluate
result = traced_evaluate(engine, metadata)

If opentelemetry-api is not installed, instrumentation degrades to a no-op and the engine behaves identically. A PROCEED, REFUSE, or PAUSE outcome is a successful decision (span status OK); only genuine failures are errors. See OBSERVABILITY.md for the full attribute and span conventions.


Security Layer (router.security) — commercial

An optional layer that secures the structure of a workflow: it enforces which trust levels may reach which side-effecting capabilities, emits wfrouter.security.* telemetry, and keeps a tamper-evident audit trail. Same content-blind, deterministic principles as the core.

Declare trust posture per container, then check it:

containers:
  intake:      { trust: UNTRUSTED, allow_transitions: [guard] }
  guard:       { sanitizer: true,  allow_transitions: [tool] }
  tool:        { capability: TOOL_EXEC, allow_transitions: [PROCEED] }
llm-router secure config.yaml        # reports trust-boundary issues
from router.engine import WorkflowEngine
from router.security import SecurityEngine, AuditLog

secure = SecurityEngine(WorkflowEngine(cfg), audit_log=AuditLog())
result = secure.evaluate(metadata)   # REFUSE if a boundary is crossed
assert secure.audit_log.verify()

The core property: untrusted input must pass a sanitizer before it can reach a side-effecting capability. See SECURITY.md for the full model, findings, and telemetry conventions.

Licensing: the whole project is PolyForm Noncommercial 1.0.0 — free to inspect, run, and use for any noncommercial purpose, commercial use requires a paid license. See LICENSING.md and PRICING.md.


License

Source-available under PolyForm Noncommercial 1.0.0. The whole project — core engine and Security Layer alike — is free to inspect, run, modify, and use for any noncommercial purpose: personal projects, research, education, non-profits, and evaluation.

Commercial use requires a paid license. That means using any part of it in a product or service you sell or host, inside a for-profit company's production or internal systems, or in paid client work. See LICENSING.md, COMMERCIAL-LICENSE.md, and PRICING.md.

If you deploy this in production, a note about your use case is always appreciated (and helps prioritize the roadmap) — but never required.


Version

Current version: 1.1.1

  • Static configuration model
  • Explicit multi-entrypoint declaration (with inferred fallback)
  • Whole-run transition budget and human approval gates
  • Refusals that name the allowed next steps
  • Optional stateful WorkflowSession convenience layer
  • Adapters: OpenAI Agents SDK, Claude Agent SDK, LangGraph, MCP
  • Mermaid / DOT topology export (llm-router graph)
  • Change review: llm-router diff and llm-router replay
  • JSON Schema for configs (llm-router schema)
  • Security Layer: trust-boundary enforcement, wfrouter.security.* telemetry, tamper-evident audit trail (commercial)
  • Licensed: PolyForm Noncommercial 1.0.0 (noncommercial free · commercial licensed)

No inheritance. No dynamic rule composition.

Future versions may extend topology modeling capabilities.


Author

Doby Baxter
Software systems developer focused on deterministic infrastructure and human-centered tooling.

Metadata

Release files for llm-workflow-router 1.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-workflow-router 1.1.1
File Size Uploaded
llm_workflow_router-1.1.1.tar.gz 91.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-workflow-router 1.1.1
File Interpreter ABI Platform
llm_workflow_router-1.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 167.2 kB

Release files / llm_workflow_router-1.1.1.tar.gz

Download URL llm_workflow_router-1.1.1.tar.gz
Size 91.0 kB
Tags Source
SHA-256 checksum
How to use checksums
f6739e2f78a81136a736c284b1d4ed91940ee3907e3d8a31ac6e328c15f4bd2e
BLAKE2b-256 checksum
How to use checksums
a5f62a49893b2a9921060ae83c4a1f6352c90d79ddbda357182980d207379b77
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release files / llm_workflow_router-1.1.1-py3-none-any.whl

Download URL llm_workflow_router-1.1.1-py3-none-any.whl
Size 76.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fee0214a22df4f166a91cd3238c8b9bbf657b80ad338cb50e9992a00dcfec334
BLAKE2b-256 checksum
How to use checksums
83c9368eb30fbb7d7eddf70f068651127885823b108e66495f2ad345119e5765
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release history Release notifications | RSS feed

This release

1.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page