veval-sdk
Python SDK for Veval — trace, evaluate, and test AI agents.
Install
pip install veval-sdk
Quick start
Wrap your agent with run_async to send a trace to the Veval dashboard:
import asyncio
from veval import VevalSdk, VevalOptions, VevalExecutionContext
sdk = VevalSdk(VevalOptions(api_key="veval-..."))
async def my_agent(ctx: VevalExecutionContext) -> str:
result = await ctx.track_step_async(
name="call-llm",
input=ctx.input,
step=lambda: call_my_llm(ctx.input),
)
return result
output = asyncio.run(sdk.run_async("my-agent", my_agent, input="Hello"))
Every call to track_step_async records the step name, input, output, timing, and any metadata you attach.
Tracking steps
Use the StepHandle to attach token counts, cost, and model name:
async def agent(ctx: VevalExecutionContext) -> str:
async def llm_call(handle):
response = await call_my_llm(ctx.input)
handle.set_meta("tokens_in", response.usage.input_tokens)
handle.set_meta("tokens_out", response.usage.output_tokens)
handle.set_meta("cost_usd", response.usage.cost)
handle.set_meta("model", "claude-opus-4-6")
return response.text
return await ctx.track_step_async("call-llm", ctx.input, llm_call)
Assertions
Use TraceAssert to validate behaviour at the trace level:
from veval import TraceAssert
assertions = [
TraceAssert.no_errors(),
TraceAssert.max_steps(10),
TraceAssert.step_exists("call-llm"),
TraceAssert.max_cost(0.05),
TraceAssert.max_duration(30_000),
TraceAssert.output_contains("success"),
TraceAssert.tool_called("web-search"),
]
Scenarios
Run a named scenario against a list of test items and assertions:
from veval import ScenarioItem
items = [
ScenarioItem(name="basic greeting", input="Hello"),
ScenarioItem(name="edge case", input=""),
]
result = asyncio.run(
sdk.run_scenario_async("smoke-test", my_agent, assertions, items=items)
)
print(f"Passed: {result.pass_count}/{len(result.results)}")
for r in result.results:
if not r.passed:
print(f" FAIL {r.item.name}: {r.failures}")
You can also pull items from the Veval dashboard by omitting items:
result = asyncio.run(sdk.run_scenario_async("smoke-test", my_agent, assertions))
Replay and snapshot testing
Load a previously recorded trace and replay it with mocked LLM responses — no real calls, deterministic results:
trace = asyncio.run(sdk.get_trace_async("tr_abc123"))
replay = asyncio.run(sdk.replay_async(
trace,
my_agent,
ReplayOptions(mock_llm_responses=True, assertions=assertions),
))
print("Passed" if replay.passed else replay.failures)
Store a known-good run — every step with its input and output — and detect when a new run drifts from it: a changed prompt or tool argument, a step added, dropped, repeated, or reordered:
# Save a baseline from a recorded trace. The trace is pinned, so retention never deletes it.
asyncio.run(sdk.save_snapshot_async("my-baseline", "tr_abc123"))
# In tests: an assertion like any other. A missing baseline fails; it never passes silently.
replay = asyncio.run(sdk.replay_async(trace, my_agent, ReplayOptions(
mock_llm_responses=True,
assertions=[TraceAssert.matches_snapshot(sdk, "my-baseline")],
compare_with_recording=SnapshotOptions(), # also fail if steps/inputs drift from the trace
)))
# Or compare directly, and record the result in the dashboard.
baseline = asyncio.run(sdk.get_snapshot_async("my-baseline"))
diff = asyncio.run(sdk.compare_snapshot_async("my-baseline", baseline, ctx))
if diff.has_changes:
print(diff.summary()) # every change, with a line diff of changed inputs
Test SDK
VevalTestSdk is a drop-in replacement that blocks real LLM calls, making it safe to use in unit tests. It still reports to the dashboard.
from veval import VevalTestSdk
sdk = VevalTestSdk(options).with_replay(trace)
output = asyncio.run(sdk.run_async("my-agent", my_agent, input="test input"))
API reference
VevalOptions
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key |
str |
"" |
Your Veval API key — it also determines the workspace |
project_id |
str |
"" |
Deprecated and ignored; will be removed |
flush_interval_ms |
int |
5000 |
Batch flush interval |
flush_batch_size |
int |
50 |
Max traces per flush |
VevalSdk
| Method | Description |
|---|---|
run_async(name, callback, input) |
Run agent and send trace |
get_trace_async(trace_id) |
Fetch a recorded trace |
save_snapshot_async(name, trace_id_or_ctx) |
Store a named baseline (pins the trace) |
get_snapshot_async(name) |
Load the latest stored baseline |
load_snapshot_async(trace_id) |
Build a snapshot from a trace (not pinned) |
compare_snapshot_async(name, snapshot, ctx, options) |
Diff a run against a baseline and record it |
replay_async(trace, callback, options) |
Replay trace with mocked outputs |
run_scenario_async(name, agent, assertions, items) |
Run a test scenario |
TraceAssert built-ins
| Assertion | Description |
|---|---|
no_errors() |
All steps must succeed |
max_steps(n) |
Total step count must not exceed n |
step_exists(name) |
A step with this name must appear |
max_cost(usd) |
Total cost must not exceed usd |
max_duration(ms) |
Total step duration must not exceed ms |
output_contains(text) |
At least one step output must contain text |
tool_called(name) |
A tool step with this name must appear |
Custom assertions implement async ITraceAssertion.evaluate_async(ctx) -> Optional[str] — return None to pass or an error string to fail.
Requirements
- Python 3.10+
- No third-party dependencies
Metadata
Release files for veval-sdk 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| veval_sdk-1.1.0.tar.gz | 21.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| veval_sdk-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 44.6 kB
Release files / veval_sdk-1.1.0.tar.gz
| Download URL | veval_sdk-1.1.0.tar.gz |
|---|---|
| Size | 21.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d99872f696b30419cf901acf8b72091342701543c07d0e46bd3d26668bc9c4a4
|
|
BLAKE2b-256 checksum How to use checksums |
02ac01774e464862a1ca1a4a06930411031910e40d538c9e9cefb2935c51bc8f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / veval_sdk-1.1.0-py3-none-any.whl
| Download URL | veval_sdk-1.1.0-py3-none-any.whl |
|---|---|
| Size | 23.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b8373759e0f5b88959e0e9a8497ca4a994058cc24612012121e0db7775de040c
|
|
BLAKE2b-256 checksum How to use checksums |
7c9ccf1f833ac1df3462ce8e980d0462cec8235e460312b1529a76b64096b187
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log