Skip to main content

EvalView
Snapshot testing for AI agents.
Record what your agent does today. Get told when it silently changes.

PyPI version PyPI downloads Daily dogfood GitHub stars License


Your agent returns 200 and looks fine. But a model update, a provider change, or a one-line prompt edit just made it skip a clarification, call the wrong tool, or quietly drop output quality. Your tests still pass. Your users notice before you do.

EvalView snapshots your agent's behavior — the tools it calls, in what order, with what output — and tells you the moment that behavior changes. Like Jest snapshots, but for tool-calling, multi-turn agents.

demo.gif

↑ 30-second live demo — no API key needed

Quick Start

pip install evalview
evalview snapshot    # Record your agent's current behavior as the baseline
evalview check       # After any change, diff against the baseline

That's the whole loop. check returns one of:

  ✓ login-flow        PASSED          behavior matches baseline
  ⚠ refund-request    TOOLS_CHANGED   called a different tool, or in a different order
  ✗ billing-dispute   REGRESSION      score dropped — output quality fell

It diffs the whole trajectory — tool names, parameters, and order — not just the final string. The deterministic tool + sequence diff runs offline, with no API key. Add an LLM judge only when you want output-quality scoring.

No agent yet? See it work in 30 seconds:

evalview demo

Why snapshot testing (and not assertions)?

Most eval tools ask you to write down what "good" looks like — assertions, metrics, rubrics. That's a lot of upfront work, and you can only catch the failures you thought to assert.

EvalView inverts it: it records what your agent actually does now, and flags any drift from that. You catch regressions you never anticipated, with zero assertions written. When the new behavior is correct, evalview snapshot accepts it as the new baseline — same as updating a snapshot in Jest.

EvalView Assertion-based eval tools
Setup Record current behavior Write assertions/metrics first
Catches Any drift from baseline Only what you asserted
Non-determinism Multi-variant baselines (up to 5 valid paths) You handle it
Unit of comparison Full tool-call trajectory Usually final output

This makes EvalView a merge-time regression gate, which is a different job from observability (Langfuse, LangSmith) or metric scoring (promptfoo, DeepEval, Braintrust). Many teams run one of those for visibility and EvalView as the gate. Honest comparisons →

EvalView tests itself in public, every day

The badge at the top is live. Every day at 09:00 UTC, a GitHub Action runs EvalView against EvalView — including a regression check where the tool snapshots a live agent and diffs it with the same snapshot / check loop this README asks you to trust. It also runs the full test suite, type checks, evalview demo, the end-to-end flows, an evalview monitor smoke test, and chat-mode self-tests.

When something breaks, the run opens a single rolling 🐕 dogfood issue and keeps updating it until the tool is green again — so failures are public, not quietly patched.

Live dogfood runs → · How it works →

CI: block regressions in every PR

# .github/workflows/evalview.yml
name: EvalView
on: [pull_request]
jobs:
  agent-check:
    runs-on: ubuntu-latest
    permissions: { pull-requests: write }
    steps:
      - uses: actions/checkout@v4
      - uses: hidai25/eval-view@v0.8.1
        with:
          openai-api-key: ${{ secrets.OPENAI_API_KEY }}

You get a PR comment with the diff, cost/latency deltas, and a pass/fail gate. CI/CD guide →

Works with your stack

LangGraph · CrewAI · OpenAI · Claude · Mistral · Ollama · MCP · any HTTP API.

evalview check --agent http://localhost:8000/invoke

Framework details →

Use it as a library

from evalview import gate

result = gate(test_dir="tests/")
result.passed   # bool
result.diffs    # per-test scores and tool diffs

Python API →

More

EvalView also does multi-turn testing, statistical/pass@k runs, record/replay cassettes, model-drift canaries, production monitoring with Slack alerts, and auto-generated regression tests from incidents. These are power-user features — start with snapshot and check, reach for the rest when you need them.

Full feature reference · Getting Started · FAQ

Contributing

This is a young project built mostly by one developer. Issues, PRs, and "I tried it and X was confusing" feedback are all genuinely valuable.

License: Apache 2.0


Star History Chart

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

evalview-0.8.1.tar.gz (1.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

evalview-0.8.1-py3-none-any.whl (939.8 kB view details)

Uploaded Python 3

File details

Details for the file evalview-0.8.1.tar.gz.

File metadata

  • Download URL: evalview-0.8.1.tar.gz
  • Upload date:
  • Size: 1.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.13

File hashes

Hashes for evalview-0.8.1.tar.gz
Algorithm Hash digest
SHA256 0a55adcaac3fb633065be0e2ec94c3b0e9e62f3a128cafba756970559079f288
MD5 6bc5356a398374dd94dedae589d8d875
BLAKE2b-256 aaddad07e69692ad453cef3ff472f4315ab5c3c9201166e0ed671ae3585ce4a1

See more details on using hashes here.

File details

Details for the file evalview-0.8.1-py3-none-any.whl.

File metadata

  • Download URL: evalview-0.8.1-py3-none-any.whl
  • Upload date:
  • Size: 939.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.13

File hashes

Hashes for evalview-0.8.1-py3-none-any.whl
Algorithm Hash digest
SHA256 971db037a171b1889354412a39ff275afce1be1950736457da8a5f41f8bd7d72
MD5 7c768df5b0ccd0e56651f2f4b098e11c
BLAKE2b-256 f4095570816dd0ed9b43f0c3a8736e5715e1a883072de5fdb57eb5649615da5c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page