Skip to main content

AgentRegress

PyPI

Regression testing for AI agents. Turn production failures into automated tests.

┌──────────────┐     ┌──────────────┐     ┌──────────────┐
│  Langfuse /  │     │              │     │   CI / CD    │
│  Langsmith / │────►│ AgentRegress │────►│   Pipeline   │
│  Arize / any │     │              │     │              │
│              │     │  Policies    │     │  Run tests   │
│  "What did   │     │  Tests       │     │  Block bad   │
│   it do?"    │     │  Generation  │     │   deploys    │
└──────────────┘     └──────────────┘     └──────────────┘
      OBSERVE              TEST               PREVENT
    (not our job)       (our job)          (our job)

AgentRegress is the missing layer between your observability tool and your deploy pipeline. It is not a tracing tool. It is not a dashboard. It does one thing: it tests your agent and blocks bad deploys.


Install

pip install agentregress
# or
uv add agentregress

For Langfuse integration:

pip install "agentregress[langfuse]"

5-Minute Quickstart

1. Write a policy

Policies live in your repo as YAML files. Anyone can write them.

# policies/support.yaml
policies:
  - id: no_repeat_info_ask
    severity: critical
    text: >
      If the customer has already provided a piece of information
      (name, account number, email), the agent must not ask for it again.

  - id: no_price_without_api
    severity: critical
    text: >
      The agent must never state a specific price without first calling
      the pricing API in the current conversation.

2. Score a trace

export ANTHROPIC_API_KEY=sk-ant-...

# Score a trace exported from Langfuse
agentregress score --trace ./exports/run-4821.json --policies ./policies/

# Output:
# ✗ run-4821: FAIL
#   Policy violated: no_repeat_info_ask (critical)
#   Turn 7: Agent asked for account number (previously given at turn 2)
#   Confidence: high

3. Write a test

# tests/no_repeat_account_ask_001.yaml
test:
  id: no_repeat_account_ask_001
  severity: critical

  persona: >
    A frustrated customer who wants to upgrade their plan.
    Has account number A-7291. Gets annoyed if asked twice.

  goal: >
    Upgrade from Basic to Pro plan.

  script:
    - user: "Hi, I want to upgrade my plan"
    - user: "My account number is A-7291"
    - persona_driven: true           # LLM plays the user from here
    - user: "Can you confirm the price?"

  checks:
    - policy: no_repeat_info_ask
    - policy: no_price_without_api

  runs: 10
  pass_threshold: 0.8

4. Run against your agent

# Against an HTTP endpoint
agentregress test --suite ./tests/ --agent-url http://localhost:8000/chat

# Against a subprocess (stdin/stdout JSON protocol)
agentregress test --suite ./tests/ --agent-cmd "python my_agent.py"

# Output:
# AgentRegress — 2 tests
#
# ✓ no_repeat_account_ask_001    9/10 passed  (threshold: 8/10)
# ✗ no_price_without_api_001     3/10 passed  (threshold: 8/10)  ← REGRESSION
#
# Result: 1 failure. Exit code 1.

5. Block bad deploys in CI

# .github/workflows/agent-eval.yml
name: Agent Tests
on: [pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: agentregress/agentregress@v1
        with:
          suite: ./tests
          agent-cmd: "python my_agent.py"
          severity: critical,functional
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}

Agent Protocol

HTTP (--agent-url)

AgentRegress sends a POST request with the conversation:

{
  "messages": [
    {"role": "user", "content": "Hi, I want to upgrade"},
    {"role": "assistant", "content": "Sure! What's your account number?"}
  ]
}

Expected response:

{"response": "Your account has been upgraded to Pro."}

OpenAI-compatible endpoints (/v1/chat/completions) are auto-detected.

Subprocess (--agent-cmd)

Your agent reads one JSON line from stdin and writes one JSON line to stdout:

# my_agent.py
import json, sys

data = json.load(sys.stdin)
messages = data["messages"]
new_message = data["new_message"]

response = call_your_agent(messages, new_message)
print(json.dumps({"response": response}))

Configuration

# agentregress.yaml
agent:
  command: "python agent/main.py"   # or url: http://localhost:8000/chat
  timeout: 60                        # seconds per agent turn

scoring:
  model: claude-haiku-4-5-20251001  # LLM judge model

testing:
  default_runs: 10
  default_pass_threshold: 0.8
  concurrency: 4

sources:
  langfuse:
    project: my-support-agent
    # Keys via: LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY

Supported Trace Sources

Source Flag / URI
JSON file (any format) --trace ./run.json
Directory of JSONs --traces ./exports/
Langfuse trace --source langfuse://trace/abc123
Langfuse project --source langfuse://project/my-agent

Supported JSON formats: plain messages array, OpenAI chat completions response, Anthropic messages response, Langfuse export.


Policy Severity

Severity Use when
critical Violations are unshippable (safety, legal, core UX)
functional Violations degrade the product significantly
informational Tracking only — never fails CI

Community Policy Templates

Ready-to-use policies in examples/policies/:

  • support.yaml — no-repeat-ask, retention-before-cancel, no-price-without-api, escalate-on-frustration
  • general.yaml — no-hallucination, no-infinite-loop, confirm-before-destructive-action, no-pii-in-responses, graceful-out-of-scope

What AgentRegress is NOT

Not this Use this instead
Tracing / observability Langfuse, Langsmith, Arize
LLM playground Langsmith, PromptLayer
Model comparison Braintrust, Humanloop
Agent framework LangGraph, CrewAI
Guardrails / safety layer NeMo Guardrails, Lakera

Release files for agentregress 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agentregress 0.1.1
File Size Uploaded
agentregress-0.1.1.tar.gz 207.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agentregress 0.1.1
File Interpreter ABI Platform
agentregress-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 262.8 kB

Release files / agentregress-0.1.1.tar.gz

Download URL agentregress-0.1.1.tar.gz
Size 207.1 kB
Tags Source
SHA-256 checksum
How to use checksums
78e4e8cc1e005667ab383359957cb608a4a8fd13c69a126ce45ab51fbf2aade1
BLAKE2b-256 checksum
How to use checksums
c724e59e51492f2d9ab7016d9b4fdb2206fcac9a2bde2050fb9eb2c006d844bc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.6 {"installer":{"name":"uv","version":"0.10.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / agentregress-0.1.1-py3-none-any.whl

Download URL agentregress-0.1.1-py3-none-any.whl
Size 55.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6d6a7b0086e59eaf37482066d566a00569478ea3fc5af2be9484506c02f5ced9
BLAKE2b-256 checksum
How to use checksums
158540573d24fb1abaa8aba59718024d1d1bdd1aeaf3bd2dbb3b5b9f9ede7d10
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.6 {"installer":{"name":"uv","version":"0.10.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page