Skip to main content

simulhausen

PyPI Python CI License

Simulation testing for agno agents that runs entirely on your side. An LLM plays the user and holds a multi-turn dialogue with your agent, team or workflow. Then an LLM judge grades the result against your criteria, one verdict and one reason per criterion.

The API follows LangWatch Scenario, minus the cloud: no LangWatch account, no tracing backend, no litellm. The user simulator and the judge are plain agno agents on whatever model you configure, and the reports land in your terminal, in a JSON file and in a markdown comment for your pull request.

import simulhausen as sim

result = await sim.run(
    name="refund: duplicated charge",
    description="The user was charged twice for order 1042 and wants the extra charge refunded.",
    agents=[
        sim.agno_adapter(lambda: support_agent),
        sim.UserSimulatorAgent(),
        sim.JudgeAgent(criteria=[
            "The agent looks the order up before promising anything",
            "The agent does not invent a refund amount",
        ]),
    ],
)
assert result.success, result.failure_summary()

Install

uv add --dev simulhausen

or pip install simulhausen. Python 3.12+ and agno 3.x. The pytest plugin registers itself on install.

Quick start

Configure the model for the simulator and the judge once, in conftest.py. Keep model construction inside the fixture, so collecting tests never needs credentials:

# conftest.py
import pytest

import simulhausen as sim


@pytest.fixture(scope="session", autouse=True)
def configure_simulations() -> None:
    from agno.models.openai import OpenAIChat

    sim.configure(model=OpenAIChat(id="gpt-4.1-mini"), max_turns=8)

Then write scenarios as ordinary async tests:

# test_support.py
import pytest

import simulhausen as sim


def _support_agent():
    from myapp.agents import support_agent  # lazy: imported only when the test runs

    return support_agent


def _no_lookup_yet(state: sim.ScenarioState) -> None:
    assert not state.has_tool_call("lookup_order"), state.tool_call_names()


@pytest.mark.simulation
async def test_refund_asks_for_order_number() -> None:
    result = await sim.run(
        name="refund: missing order number",
        description="The user wants a refund but does not say which order it is about.",
        agents=[
            sim.agno_adapter(_support_agent, user_id="user-42"),
            sim.UserSimulatorAgent(),
            sim.JudgeAgent(criteria=["The agent asks for the order number before anything else"]),
        ],
        script=[
            sim.user("i want my money back"),
            sim.agent(),
            _no_lookup_yet,
            sim.judge(),
        ],
    )
    assert result.success, result.failure_summary()

agno_adapter takes the agent itself or a factory, and forwards extra keyword arguments to arun (session_state=..., user_id=...). The whole scenario runs in one agno session. A runnable version lives in examples/.

Scripts

Without script, a run is [proceed()] followed by the judge's verdict. With a script you decide what happens and in which order:

Step What it does
user("text") user message with fixed text
user() user message written by the simulator
agent() call the agent under test
agent("text") inject an agent reply without calling the agent
message(role, "text") append any message as is, to seed earlier context
proceed(turns=None, on_turn=..., on_step=...) simulator and agent talk on their own (see below)
judge() final verdict on the judge's criteria; later steps still run
judge(criteria=[...], additional_context=...) checkpoint: fails the scenario on the spot if a criterion fails, goes on otherwise
succeed(reasoning) / fail(reasoning) end the scenario right here
any callable gets the ScenarioState; sync or async; raise AssertionError to fail

proceed() stops at the first of these: turns turns played, max_turns reached, the simulator reporting the goal as achieved, or the judge concluding it has seen enough. The judge is not consulted before min_turns agent turns. If it concludes while some criteria are still inconclusive and turns remain, the dialogue goes on to collect the missing evidence.

When a script has a judge with criteria but no judge() step, the judge runs after the last step. It also runs again there if the dialogue went on after its last verdict, so the final verdict always covers the whole dialogue; a new final verdict replaces the previous one, while checkpoint verdicts accumulate. additional_context gives the judge evidence the dialogue cannot show, such as a database row the agent was supposed to write.

Verdicts and statuses

The judge answers per criterion: passed, failed or inconclusive, each with a restated requirement and its reasoning. A criterion phrased as a prohibition ("the agent must not X") passes when X did not happen. inconclusive means the evidence was not there, for example because the dialogue ended before the moment the criterion depends on.

Every result gets one of three statuses:

  • PASS: every criterion passed and no assertion failed.
  • FAIL: a criterion failed, a script assertion failed, the judge score was below threshold, or fail() ran.
  • WARN: the scenario could not be verified. Either the harness broke (SimulationHarnessError: the judge or the simulator returned no structured output, the model backend kept failing), or some criteria came back inconclusive and none failed.

WARN is kept out of the pass rate on purpose. A broken judge says nothing about your agent, and counting it as a failure turns the pass rate into a measure of helper-LLM reliability. pass_rate = PASS / (PASS + FAIL), and the WARN count is reported next to it. Harness errors are still raised, so a plain pytest run fails on them.

ScenarioResult carries verdicts, passed_criteria, failed_criteria, inconclusive_criteria, the judge's reasoning, messages, tool_calls, turns, total_time and agent_time.

For an overall quality gate on top of the criteria, JudgeAgent(criteria, score_threshold=7) also grades the whole dialogue 1-10 with agno's JudgeScorer. Calibrate the threshold on known good and bad runs first.

Checking tools

agno_adapter collects every ToolExecution of the run, including those of team members, so assertions work for teams too:

state.has_tool_call("lookup_order")         # ran cleanly at least once
state.tool_call_names()                     # names in call order, with repeats
state.has_tool_calls_in_order(["search", "lookup_order"])  # subsequence; other calls may sit between
state.last_tool_call("lookup_order")        # the ToolExecution, or None

Only clean executions count: a call that errored, was rejected or is paused waiting for confirmation does not satisfy has_tool_call. That matches agno.scorer.ToolCallScorer.

The judge sees the same evidence: the transcript, the RAG references the agent retrieved (with source URLs) and a capped summary of tool executions, errors included. Each block is fenced with a random per-call tag and the judge is told never to follow instructions inside them.

Writing scenarios that hold up

One scenario checks one behaviour from one starting state. When it fails, the name alone should tell you what broke.

  • name is the outcome in a few words, readable in a PR comment: refund: asks for the order number.
  • description is for the simulator: starting conditions, the user's goal, how they behave. It is not a criterion.
  • Each criterion is a single observable requirement. "The answer is good" is not one.

Check in code whatever code can check: tool calls and their order, exact identifiers, an expected error, a call that must not happen. Leave semantics to the judge: did the agent make something up, explain a limitation, stick to what the tools returned.

Spell out the user messages that matter for the regression with user("..."). Use user() or proceed() only where variety in the user's wording is part of what you test. Most regressions need one or two user/agent pairs; every extra turn costs money and adds variance.

Do not paper over a reproducible failure with reruns. Look at failure_summary(), the tool evidence and the transcript in the JSON report first: they tell an agent bug from a broken backend, a wrong assertion or a vague criterion.

Running

Simulation tests call real models, so keep them out of the default run:

[tool.pytest.ini_options]
addopts = "-m 'not simulation'"
asyncio_mode = "auto"
uv run pytest -m simulation                    # failures fail the run
uv run pytest -m simulation --sim-benchmark \
  --sim-report reports/simulations.json \
  --sim-report-md reports/simulations.md       # always exits 0, writes reports
uv run pytest -m simulation -k refund -s --sim-debug  # you type the user's messages

Benchmark mode is for CI jobs that should stay green and post the outcome instead. After any run with simulations, a rich table with every scenario, its verdicts, tools and dialogue is printed. With pytest-rerunfailures only the last attempt is reported.

Parallel runs work with pytest-xdist (-n 4): each worker is its own process with its own event loop, and the controller merges every worker's results into one report. Prefer that to asyncio.gather inside a test, because agno agents and HTTP clients are usually process-wide singletons. When an agent's async clients bind to the first event loop they see, run simulations on one session-scoped loop:

# conftest.py
def pytest_collection_modifyitems(items):
    for item in items:
        if item.get_closest_marker("simulation"):
            item.add_marker(pytest.mark.asyncio(loop_scope="session"), append=False)

Reports in pull requests

The markdown report starts with a hidden <!-- simulhausen-report --> marker and a headline such as 🟡 88% ▓▓▓▓▓▓▓▓▓░ · PASS 154 / FAIL 21. The emoji follows the pass rate (🟢 from 90% with no WARN, 🟡 from 75%, 🔴 below), since one failure out of 175 is noise, not a regression. The bar shows passed, not verified and failed scenarios out of all of them. On GitHub Actions:

- run: uv run pytest -m simulation --sim-benchmark --sim-report-md reports/simulations.md
- if: github.event_name == 'pull_request'
  run: gh pr comment ${{ github.event.pull_request.number }} --body-file reports/simulations.md
  env:
    GH_TOKEN: ${{ github.token }}

The report links the CI job when CI_JOB_URL (GitLab) or the GITHUB_* run variables are set.

Configuration

sim.configure(
    model=...,             # agno Model for the simulator and the judge (required)
    parser_model=...,      # optional agno Model that turns their answers into the schema
    max_turns=10,          # ceiling for proceed() and unscripted runs
    min_turns=1,           # the judge may not end the dialogue earlier
    verbose=True,          # log every message and verdict (logger "simulhausen")
    debug=False,           # ask for the user's messages on stdin
    user_name="User",      # labels in reports
    agent_name="Agent",
)

run() accepts max_turns, min_turns, verbose and debug per scenario. UserSimulatorAgent and JudgeAgent accept their own model and parser_model, plus system_prompt to replace the built-in prompt (simulhausen.prompts). The simulator also takes a persona.

Set parser_model when your backend breaks structured output in thinking mode. Some vLLM setups with a reasoning parser return empty or broken content when response_format meets reasoning; with a parser model, agno sends response_format only to the parser call. The helper agents also retry an off-schema answer twice, with a stricter instruction each time.

Custom adapters

agno_adapter covers agno agents, teams and workflows. It retries backend failures (a run that ended in RunStatus.error, a proxy's HTML error page) and raises SimulationBackendError when they persist, while a guardrail refusal (InputCheckError/OutputCheckError) counts as the agent's answer. For agents with output_schema, pass content_fn to render the structured reply as text.

Anything else is a function or an AgentAdapter subclass:

async def call_my_api(input: sim.AgentInput) -> str:
    return await my_client.chat(input.thread_id, input.last_new_user_message_str())

adapter = sim.function_adapter(call_my_api)


class MyAdapter(sim.AgentAdapter):
    async def call(self, input: sim.AgentInput) -> sim.AgentTurn:
        reply = await my_agent.respond(input.messages)
        return sim.AgentTurn(content=reply.text, tools=reply.tool_executions)

AgentInput has the whole transcript (messages), what changed since the agent last spoke (new_messages) and the live scenario_state.

Compared with LangWatch Scenario

Ported: the script steps (user, agent, message, proceed, judge with inline criteria and additional_context, succeed, fail), per-criterion passed/failed/inconclusive verdicts, the two-phase judge (an argument-free continue/conclude decision, then the verdict), min_turns, simulator persona and system_prompt, debug mode, total_time/agent_time.

Different: models are agno Model objects instead of litellm strings; the simulator reports goal_achieved so unscripted dialogues end without filler; harness failures and inconclusive scenarios are WARN and stay out of the pass rate; a script that ends without a judge verdict passes if its assertions passed; the judge also gets RAG references and tool executions.

Not ported: anything that needs LangWatch (tracing, events, evaluators, remote traces), voice and realtime agents, the red-team agents, the joblib cache and long-transcript discovery tools.

Development

uv sync
just ci          # ruff, ty, pytest with coverage
just precommit   # every prek hook

License

Apache-2.0, see LICENSE. simulhausen began as a local reimplementation of LangWatch Scenario (Apache-2.0, Copyright LangWatch); its API and parts of its prompts follow Scenario. See NOTICE.

Metadata

Release files for simulhausen 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for simulhausen 0.1.0
File Size Uploaded
simulhausen-0.1.0.tar.gz 45.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for simulhausen 0.1.0
File Interpreter ABI Platform
simulhausen-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 91.7 kB

Release files / simulhausen-0.1.0.tar.gz

Download URL simulhausen-0.1.0.tar.gz
Size 45.4 kB
Tags Source
SHA-256 checksum
How to use checksums
41705a16c30c27d5884506ddcd31ff988814b9975b28fd4cf69208d0dfe91319
BLAKE2b-256 checksum
How to use checksums
cbd563344394b87d794c5cda56ea0fae20c956777c94c2e648292269e866aa1f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / simulhausen-0.1.0-py3-none-any.whl

Download URL simulhausen-0.1.0-py3-none-any.whl
Size 46.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
34b622cd6d5d837fb5ba175348f9b413f49846f27fe5e79607f37af3e31aefc2
BLAKE2b-256 checksum
How to use checksums
46c2089f8534429613ed0fba1f618ee23aba4ff01729c5556cfc3e7fc9482c06
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page