Phthos Eval
Score an agent run, not a chat reply.
An agent can answer fluently and still call the wrong tool, loop, blow the budget, or break policy. Phthos Eval reads recorded traces (LLM steps + tool calls) and writes a diagnosis JSON you can gate CI on or hand to another system. It does not rewrite prompts, open PRs, or fine-tune.
Python 3.11+. Install from PyPI:
pip install phthos-eval
flowchart LR
subgraph stacks [Your agent]
LC[LangChain]
ADK[Google ADK]
OTH[CrewAI / custom]
end
stacks --> spans["spans: llm + tool"]
spans --> scorers[Phthos Eval scorers]
scorers --> dx[diagnosis.json]
dx --> ci[CI gate]
dx --> hint[change_class hint]
hint -.-> fix[You / other product applies the fix]
What you get
- Deterministic checks (no API key): expected tools, argument shape, cost/step budget, deny-list policy, repeated tool loops.
- A success profile, not one score:
task_success(pass^N),n_run_reliability(mean repeat pass-rate),pass_at_n(pass@N), plus cost-per-task, p95 latency, tool steps, policy hits. - A diagnosis file: scores, typed failures with span ids, and a
change_classhint (tool,policy,prompt, …).
Optional: an LLM judge using your key. Without a key, everything above still runs.
Sample scores (bundled fixture)
python -m phthos_eval run -d fixtures/dataset.json — 5 cases, 2 runs each. One case is clean; the others are seeded failures.
| Metric | Value | What that means on this suite |
|---|---|---|
task_success (pass^N) |
0.20 | 1 of 5 cases passed every repeat |
n_run_reliability |
0.20 | Mean per-case pass fraction (here the same: no mixed cases) |
pass_at_n (pass@N) |
0.20 | Share of cases with at least one passing repeat |
cost / cost_mean |
1.71 / 0.171 | Total USD / USD per trace |
latency_p50_ms / latency_p95_ms |
30 / 4100 | SLA percentiles of per-trace latency |
steps / steps_mean |
14 / 1.4 | Tool-call count (step efficiency) |
policy_hits |
2 | Deny-list hits (send_money) |
wrong_tool_hits / budget_hits / loop_hits |
4 / 2 / 2 | Other typed failures |
change_class |
policy |
Highest-priority hint (policy before tool/budget/loop) |
| Case | Passed | What fired |
|---|---|---|
pass-search |
yes | — |
fail-wrong-tool |
no | wrong_tool |
fail-budget |
no | budget |
fail-policy |
no | policy (and wrong_tool: denied tool is not in the allow-list) |
fail-loop |
no | loop |
Support-agent dogfood (examples/support_agent/dataset.json): task_success 0.50, cost_mean 0.002, latency_p95_ms 230, policy_hits 2, change_class policy. status-ok passes; refund-denied fails.
Any agent stack (LangChain, Google ADK, …)
Phthos Eval does not import LangChain, Google ADK, CrewAI, LlamaIndex, AutoGen, or similar. Those packages run the agent. We score what it did.
Seamless here means one shared trace shape, not a plugin inside each framework:
flowchart TB
LC[LangChain / LangGraph]
ADK[Google ADK]
CR[CrewAI / AutoGen / custom]
OT[OpenTelemetry / OpenInference]
LC --> MAP[TraceSink.wrap / OTel]
ADK --> MAP
CR --> MAP
OT --> MAP
MAP --> SP["spans: id, type llm or tool, name, args, cost"]
SP --> PE[phthos-eval]
PE --> DJ[diagnosis.json]
Today: wrap the agent you already have. TraceSink attaches collectors; the framework still runs the agent. We only score spans.
from phthos_eval import TraceSink
sink = TraceSink()
agent = sink.wrap(agent) # see TraceSink.frameworks
# run the agent as usual, then:
doc = sink.diagnose(expected_tools=["search"])
Unknown stack: sink.add_llm(...) / sink.add_tool(...), or export OpenTelemetry OpenInference to POST /v1/otel/traces. Examples: agent_integration_examples/.
You do not replace the agent runtime. Swap LangChain for ADK and the eval file stays valid as long as spans are still llm / tool JSON.
Live: POST /v1/traces with that span JSON, or OTLP/HTTP JSON (openinference.span.kind / gen_ai.tool.name) to a self-hosted engine. See Live engine.
Quick start
Save traces as JSON (see Dataset format), then:
python -m phthos_eval run -d eval/dataset.json -o diagnosis.json
python -m phthos_eval check diagnosis.json
Fail CI when anything is wrong:
python -m phthos_eval run -d eval/dataset.json -o diagnosis.json --fail-on-findings
Try the bundled examples after cloning this repo:
python -m phthos_eval run -d fixtures/dataset.json -o diagnosis.json
python -m phthos_eval run -d examples/support_agent/dataset.json -o diagnosis.json
# Wrap + score: python agent_integration_examples/google_adk/lib/agent.py
# Wrap + live: python agent_integration_examples/google_adk/live/agent.py
Live engine (self-host)
Same scorers as offline, on a sampled production stream. You run the process; data stays on your machine. This is not LangSmith. Optional hosted mode (--hosted) is the same binary with login and tenants — OSS self-host stays the default.
sequenceDiagram
participant Agent
participant Engine as Live engine
participant Worker as Async scorers
Agent->>Engine: POST /v1/traces
Engine-->>Agent: 202 accepted (sampled or not)
Note over Agent: Agent request already finished
Engine->>Worker: if sampled (~5%)
Worker->>Worker: same scorers as offline
Worker-->>Engine: diagnosis.json in SQLite
# default sample rate 5%, judge off
phthos-eval live -c examples/live/config.json
# or
docker compose up --build
UI at http://127.0.0.1:8765 — pass rate, cost, policy hits, open a run to see diagnosis JSON. No prompt editor.
Ingest does not wait for scoring:
from phthos_eval.live import LiveClient
client = LiveClient("http://127.0.0.1:8765")
client.ingest(spans=[...], agent_id="support", expected_tools=["search"])
GET /v1/scores · GET /v1/diagnoses · GET /v1/diagnoses/{id} · POST /v1/compare · POST /v1/diagnoses/{id}/export writes an offline dataset you can phthos-eval run. After you change the agent: phthos-eval compare --before a.json --after b.json. Contract: docs/CONSUMER.md. Example: examples/consumer/.
Demo (forces 100% sample — not for production):
PHTHOS_EVAL_SAMPLE_RATE=1 docker compose up --build
phthos-eval live-demo
Full Compose / OTel / cost knobs: examples/live/README.md.
Do not bankrupt yourself: default is 5% sample and no LLM judge. A judge key on 100% of live traffic is usually more expensive than the agent. Opt in with --live-judge plus OPENAI_API_KEY / PHTHOS_EVAL_API_KEY only if you accept that bill.
Hosted mode (same engine)
We operate this in cloud; you can also run it yourself. It is not a second product: PHTHOS_EVAL_HOSTED=1 or phthos-eval live --hosted.
- Sign-up / login, isolated tenants, dashboard (live, history, datasets), alerts when pass rate drops
- Default BYOK — traces are not sent to a model we own; live judge still off unless
--live-judgeand your key - Export diagnoses + datasets (
GET /v1/export) — no hostage data - Self-host path above stays complete without accounts
phthos-eval live --hosted --host 0.0.0.0
# or
docker compose -f docker-compose.hosted.yml up --build
from phthos_eval.live import LiveClient
client = LiveClient("https://your-eval-url", api_key="pk_…")
client.ingest(spans=[...], expected_tools=["search"])
CI can keep using phthos-eval run locally, or put_dataset / run_dataset on the hosted URL. Compare two runs: POST /v1/compare. Poll: GET /v1/diagnoses?since=. Details: docs/hosted.md, docs/CONSUMER.md. What is stored: docs/PRIVACY.md. GET /status for health.
Paid cloud extras (retention, SAML, hosted judges we meter, seats) are ops, not a paywall on scores. Catalog: docs/PLANS.md. Stripe / SAML UI is the private overlay, not this package.
Integrate in a project
1. Record traces
When your agent runs (tests or a small harness), write spans like:
{
"spans": [
{ "id": "s0", "type": "llm", "latency_ms": 120, "cost_usd": 0.002 },
{
"id": "s1",
"type": "tool",
"name": "lookup_order",
"args": { "order_id": "A-100" },
"latency_ms": 30,
"cost_usd": 0.0
}
]
}
You do not run the agent through Phthos Eval. You export what it did, then score the export.
2. Put cases in a dataset
One file per suite. Each case needs n_runs traces (default in examples: 2) so reliability is real.
{
"id": "my-agent",
"n_runs": 2,
"budget": { "max_cost_usd": 0.05, "max_steps": 8 },
"policy": { "deny_tools": ["issue_refund"] },
"tool_schemas": {
"lookup_order": { "required": ["order_id"] }
},
"cases": [
{
"id": "status-ok",
"expected_tools": ["lookup_order"],
"traces": [{ "spans": [] }, { "spans": [] }]
}
]
}
3. Call from Python (pytest)
import json
from pathlib import Path
from phthos_eval import run_dataset, validate_diagnosis
def test_agent_eval():
dataset = json.loads(Path("eval/dataset.json").read_text())
doc = run_dataset(dataset)
assert validate_diagnosis(doc) == []
assert doc["change_class"] == "none"
assert doc["scores"]["n_run_reliability"] == 1.0
Custom check (still deterministic — no LLM):
from phthos_eval import failure, run_dataset
class NoEmptyTrace:
def score(self, trace, *, case, dataset, case_id, trace_index):
if not trace.get("spans"):
return [failure("policy", "empty", case_id=case_id, trace_index=trace_index)]
return []
doc = run_dataset(dataset, scorers=[NoEmptyTrace()]) # replaces defaults; add default_scorers() to keep them
To keep built-in scorers and yours:
from phthos_eval import default_scorers, run_dataset
doc = run_dataset(dataset, scorers=[*default_scorers(), NoEmptyTrace()])
4. GitHub Actions
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install phthos-eval
- run: python -m phthos_eval run -d eval/dataset.json -o diagnosis.json --fail-on-findings
Full copy-paste: examples/github-eval.yml.
Dataset format
| Field | Role |
|---|---|
n_runs |
How many traces per case must exist and be scored |
budget.max_cost_usd / max_steps |
Fail the trace if cost or tool-call count is over the cap |
policy.deny_tools |
Fail if that tool was called |
tool_schemas |
Fail if a known tool is missing required args |
cases[].expected_tools |
Fail if a tool call is not in this allow-list |
cases[].traces |
Recorded runs; each span needs id, type (llm or tool) |
Tool spans: name, args, optional cost_usd, latency_ms.
Metrics (what they mean)
Suite-level diagnosis.json → scores. This is an AgentSLABench-style profile (success under cost/latency/policy), not a single 0–1 judge number.
Success (codegen / agent standard: pass^k vs pass@k):
| Metric | Range | Example | Why it matters |
|---|---|---|---|
task_success |
0–1 | 0.20 | pass^N: share of cases where every N-run passed. Job done, not a fluent last message. |
n_run_reliability |
0–1 | 0.20 | Mean over cases of (passing traces / N). A flaky case scores 0.5, not 0. Distinct from task_success. |
pass_at_n |
0–1 | 0.20 | pass@N: share of cases with at least one passing repeat. Lucky-once still counts here. |
On the bundled fixture every case is all-pass or all-fail, so the three numbers match. They diverge as soon as a case is mixed (see tests).
Cost, latency, steps (ops / SLA):
| Metric | Range | Example | Why it matters |
|---|---|---|---|
cost |
USD (sum) | 1.71 | Total spend on scored traces. |
cost_mean |
USD / trace | 0.171 | Cost-per-task — a correct 40-call run is still a failed product. |
latency_ms |
ms (sum) | 8593 | Sum of per-trace latency (audit). |
latency_mean_ms |
ms | 859.3 | Typical trace time. |
latency_p50_ms / latency_p95_ms |
ms | 30 / 4100 | SLA tail. p95 is what AgentSLABench / APM use, not the sum. |
steps / steps_mean |
count | 14 / 1.4 | Tool-call count (DeepEval-style step efficiency). |
tokens |
count or null | null | Sum of span tokens / input_tokens+output_tokens if you logged them. |
Safety / typed hits:
| Metric | Example | Why it matters |
|---|---|---|
policy_hits |
2 | Deny-list (or custom policy) fires. |
wrong_tool_hits |
4 | Allow-list / schema misses. |
budget_hits |
2 | Over max_cost_usd or max_steps. |
loop_hits |
2 | Same tool+args ≥ 3 times. |
judge.score (0–1) appears only if you set a judge key. Treat it as extra signal, not the verdict. We do not ship hallucination/fluency vanity scores — those are judge-only and not the contract.
Per case: cases[] has passed, pass_rate, cost, latency_ms, steps, and that case’s failures.
Failures and change_class
Each failure has a type, a span_id, and evidence (span / step / case / which of the N traces). Types:
| Type | Trigger | Typical change_class |
|---|---|---|
wrong_tool |
Tool not in expected_tools, or missing required args |
tool |
policy |
Tool on deny_tools |
policy |
budget |
Over max_cost_usd or max_steps |
model |
loop |
Same tool + args ≥ 3 times in one trace | prompt |
change_class is a hint for whatever improves the agent (you, CI, a later tool). This package does not apply the change. Another system can poll or take a webhook, ship a change, and we re-eval — docs/CONSUMER.md.
Values: prompt · tool · policy · model · finetune_data · none (clean run).
Optional LLM judge
Not required. Deterministic scorers always run.
| Variable | Use |
|---|---|
OPENAI_API_KEY or PHTHOS_EVAL_API_KEY |
Your judge key |
PHTHOS_EVAL_JUDGE_BASE_URL |
OpenAI-compatible URL (OpenAI, Ollama, a gateway, …) |
PHTHOS_EVAL_JUDGE_MODEL |
Model id (default gpt-4o-mini) |
PHTHOS_EVAL_LIVE_JUDGE |
Live engine only: set to 1 (or --live-judge) to run the judge on sampled traces. Off by default so a leftover key cannot bill every ingest. |
PHTHOS_EVAL_HOSTED |
1 or --hosted: require sign-up / API keys and isolate tenants. Omit for open self-host. |
PHTHOS_EVAL_RETENTION_DAYS |
Hosted auto-prune fallback (plan retention wins: free 30d, pro 365d). |
PHTHOS_EVAL_OPS_SECRET |
Operator header X-Phthos-Ops to set a workspace plan (cloud overlay / Stripe). |
PHTHOS_EVAL_SSO_SECRET |
HMAC secret for POST /v1/sso/consume (SAML overlay). |
PHTHOS_EVAL_HOSTED_JUDGE_API_KEY |
Our judge key for Pro hosted-judge (metered). Unset on OSS. |
Do not reuse the agent’s production keys as the judge unless you intend that. Agent keys run the system under test; judge keys only score. With no judge key, judge.skipped is true and reason is no_key.
What this is not
- Not LangSmith (no prompt playground). Hosted mode is login + scores, not a prompt IDE.
- Not an auto-fixer or fine-tuner. Export a failing live run; you (or another product) apply the change.
export-finetuneis labeled traces for their trainer. - Not a hosted LLM. Judge is BYOK and off by default on live. Traces are not required to go to a model we own.
Develop this repo
pip install -e ".[dev]"
python -m pytest
python -m ruff check src tests
License: MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file phthos_eval-0.6.9.tar.gz.
File metadata
- Download URL: phthos_eval-0.6.9.tar.gz
- Upload date:
- Size: 67.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ca2a4248c3a1d0212891c18170ec9c9a1303dc6d79f34ab0ca287f24388b5e7b
|
|
| MD5 |
87b86025de9d5d92636cdbab0cdd0571
|
|
| BLAKE2b-256 |
9589e72ea822d2368b706ff01e5853384ab5dcf3fc71d6f2c96741651753a85a
|
File details
Details for the file phthos_eval-0.6.9-py3-none-any.whl.
File metadata
- Download URL: phthos_eval-0.6.9-py3-none-any.whl
- Upload date:
- Size: 60.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
564997fb0b31528482b3fb32a5fc8c5bc7704bbff6c6a23f7e93ed7ebd839625
|
|
| MD5 |
590e21ee1f3f950046d23add6f9642b9
|
|
| BLAKE2b-256 |
133b546a1269ce980b6f571ef43445f935ed565df8d3807f27d53c2b1e297e32
|