Skip to main content

Phthos Eval

Score an agent run, not a chat reply.

An agent can answer fluently and still call the wrong tool, loop, blow the budget, or break policy. Phthos Eval reads recorded traces (LLM steps + tool calls) and writes a diagnosis JSON you can gate CI on or hand to another system. It does not rewrite prompts, open PRs, or fine-tune.

Python 3.11+. Install from PyPI:

pip install phthos-eval

PyPI · GitHub


What you get

  1. Deterministic checks (no API key): expected tools, argument shape, cost/step budget, deny-list policy, repeated tool loops.
  2. N-run reliability: the same case is scored more than once so a lucky pass does not look like a good agent.
  3. A diagnosis file: scores, typed failures with span ids, and a change_class hint (tool, policy, prompt, …).

Optional: an LLM judge using your key. Without a key, everything above still runs.


Any agent stack (LangChain, Google ADK, …)

Phthos Eval does not import LangChain, Google ADK, CrewAI, LlamaIndex, AutoGen, or similar. Those packages run the agent. We score what it did.

Seamless here means one shared trace shape, not a plugin inside each framework:

  LangChain / LangGraph
  Google ADK
  CrewAI, LlamaIndex, AutoGen, custom
           │
           │  callbacks / OTel / your logger
           ▼
  spans: { id, type: llm|tool, name, args, cost_usd, latency_ms }
           │
           ▼
  phthos-eval  →  diagnosis.json

Today: you map your framework’s events into that JSON (a small callback or post-run dump). Then phthos-eval run is the same for every stack.

Typical mappings

Stack Where traces already exist What you map
LangChain / LangGraph Callbacks, LangSmith export, or run tree Each LLM/tool event → one span
Google ADK Session / event log Tool calls → type: tool
CrewAI / AutoGen Step / message log Same
Anything on OpenTelemetry / OpenInference Span export Filter LLM + tool spans

You do not wrap the agent in a Phthos runtime. Swap LangChain for ADK and the eval file stays valid as long as spans still look like the table above.

Live: POST /v1/traces with that span JSON, or OTLP/HTTP JSON (openinference.span.kind / gen_ai.tool.name) to a self-hosted engine. See Live engine.


Quick start

Save traces as JSON (see Dataset format), then:

python -m phthos_eval run -d eval/dataset.json -o diagnosis.json
python -m phthos_eval check diagnosis.json

Fail CI when anything is wrong:

python -m phthos_eval run -d eval/dataset.json -o diagnosis.json --fail-on-findings

Try the bundled examples after cloning this repo:

python -m phthos_eval run -d fixtures/dataset.json -o diagnosis.json
python -m phthos_eval run -d examples/support_agent/dataset.json -o diagnosis.json

Live engine (self-host)

Same scorers as offline, on a sampled production stream. You run the process; data stays on your machine. This is not a hosted SaaS and not a LangSmith clone.

# default sample rate 5%, judge off
phthos-eval live -c examples/live/config.json

# or
docker compose up --build

UI at http://127.0.0.1:8765 — pass rate, cost, policy hits, open a run to see diagnosis JSON. No prompt editor.

Ingest does not wait for scoring:

from phthos_eval.live import LiveClient

client = LiveClient("http://127.0.0.1:8765")
client.ingest(spans=[...], agent_id="support", expected_tools=["search"])

GET /v1/scores · GET /v1/diagnoses/{id} · POST /v1/diagnoses/{id}/export writes an offline dataset you can phthos-eval run.

Demo (forces 100% sample — not for production):

PHTHOS_EVAL_SAMPLE_RATE=1 docker compose up --build
phthos-eval live-demo

Full Compose / OTel / cost knobs: examples/live/README.md.

Do not bankrupt yourself: default is 5% sample and no LLM judge. A judge key on 100% of live traffic is usually more expensive than the agent. Opt in with --live-judge plus OPENAI_API_KEY / PHTHOS_EVAL_API_KEY only if you accept that bill.


Integrate in a project

1. Record traces

When your agent runs (tests or a small harness), write spans like:

{
  "spans": [
    { "id": "s0", "type": "llm", "latency_ms": 120, "cost_usd": 0.002 },
    {
      "id": "s1",
      "type": "tool",
      "name": "lookup_order",
      "args": { "order_id": "A-100" },
      "latency_ms": 30,
      "cost_usd": 0.0
    }
  ]
}

You do not run the agent through Phthos Eval. You export what it did, then score the export.

2. Put cases in a dataset

One file per suite. Each case needs n_runs traces (default in examples: 2) so reliability is real.

{
  "id": "my-agent",
  "n_runs": 2,
  "budget": { "max_cost_usd": 0.05, "max_steps": 8 },
  "policy": { "deny_tools": ["issue_refund"] },
  "tool_schemas": {
    "lookup_order": { "required": ["order_id"] }
  },
  "cases": [
    {
      "id": "status-ok",
      "expected_tools": ["lookup_order"],
      "traces": [{ "spans": [] }, { "spans": [] }]
    }
  ]
}

3. Call from Python (pytest)

import json
from pathlib import Path
from phthos_eval import run_dataset, validate_diagnosis

def test_agent_eval():
    dataset = json.loads(Path("eval/dataset.json").read_text())
    doc = run_dataset(dataset)
    assert validate_diagnosis(doc) == []
    assert doc["change_class"] == "none"
    assert doc["scores"]["n_run_reliability"] == 1.0

Custom check (still deterministic — no LLM):

from phthos_eval import failure, run_dataset

class NoEmptyTrace:
    def score(self, trace, *, case, dataset, case_id, trace_index):
        if not trace.get("spans"):
            return [failure("policy", "empty", case_id=case_id, trace_index=trace_index)]
        return []

doc = run_dataset(dataset, scorers=[NoEmptyTrace()])  # replaces defaults; add default_scorers() to keep them

To keep built-in scorers and yours:

from phthos_eval import default_scorers, run_dataset

doc = run_dataset(dataset, scorers=[*default_scorers(), NoEmptyTrace()])

4. GitHub Actions

- uses: actions/setup-python@v5
  with:
    python-version: "3.12"
- run: pip install phthos-eval
- run: python -m phthos_eval run -d eval/dataset.json -o diagnosis.json --fail-on-findings

Full copy-paste: examples/github-eval.yml.


Dataset format

Field Role
n_runs How many traces per case must exist and be scored
budget.max_cost_usd / max_steps Fail the trace if cost or tool-call count is over the cap
policy.deny_tools Fail if that tool was called
tool_schemas Fail if a known tool is missing required args
cases[].expected_tools Fail if a tool call is not in this allow-list
cases[].traces Recorded runs; each span needs id, type (llm or tool)

Tool spans: name, args, optional cost_usd, latency_ms.


Metrics (what they mean)

These are on the whole suite in diagnosis.jsonscores:

Metric Range Why it matters
task_success 0–1 Share of cases where every N-run passed. A pretty last message does not count.
n_run_reliability 0–1 Same as task_success in this version: did the case pass on all repeats? Single-run “90%” is often luck.
cost USD (sum) Token/tool spend across scored traces. A correct 40-call run can still be a failed product.
latency_ms ms (sum) Time in the recorded spans. Useful for p95-style budgets later; here it is the total you logged.
policy_hits count How many deny-list (or custom policy) failures fired. Safety/compliance, not “quality vibe”.

judge.score (0–1) appears only if you set a judge key. Treat it as extra signal, not the verdict. Deterministic failures are the source of truth.

Per case: cases[] has passed and that case’s failures.


Failures and change_class

Each failure has a type, a span_id, and evidence (span / step / case / which of the N traces). Types:

Type Trigger Typical change_class
wrong_tool Tool not in expected_tools, or missing required args tool
policy Tool on deny_tools policy
budget Over max_cost_usd or max_steps model
loop Same tool + args ≥ 3 times in one trace prompt

change_class is a hint for whatever improves the agent (you, CI, a later tool). This package does not apply the change.

Values: prompt · tool · policy · model · finetune_data · none (clean run).


Optional LLM judge

Not required. Deterministic scorers always run.

Variable Use
OPENAI_API_KEY or PHTHOS_EVAL_API_KEY Your judge key
PHTHOS_EVAL_JUDGE_BASE_URL OpenAI-compatible URL (OpenAI, Ollama, a gateway, …)
PHTHOS_EVAL_JUDGE_MODEL Model id (default gpt-4o-mini)
PHTHOS_EVAL_LIVE_JUDGE Live engine only: set to 1 (or --live-judge) to run the judge on sampled traces. Off by default so a leftover key cannot bill every ingest.

Do not reuse the agent’s production keys as the judge unless you intend that. Agent keys run the system under test; judge keys only score. With no judge key, judge.skipped is true and reason is no_key.


What this is not

  • Not LangSmith (no prompt playground, no hosted cloud accounts).
  • Not an auto-fixer or fine-tuner. Export a failing live run; you (or another product) apply the change.
  • Not a hosted LLM. OSS / self-host uses your machine and, optionally, your judge key.

Develop this repo

pip install -e ".[dev]"
python -m pytest
python -m ruff check src tests

License: MIT.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

phthos_eval-0.2.0.tar.gz (31.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

phthos_eval-0.2.0-py3-none-any.whl (29.9 kB view details)

Uploaded Python 3

File details

Details for the file phthos_eval-0.2.0.tar.gz.

File metadata

  • Download URL: phthos_eval-0.2.0.tar.gz
  • Upload date:
  • Size: 31.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for phthos_eval-0.2.0.tar.gz
Algorithm Hash digest
SHA256 dd12c19e355451c4057eea5d5a8ac2613124954980c35e586b88645c411eecb2
MD5 689ce50b2ed41096d1ad0fb97670c294
BLAKE2b-256 088c9faf62e10729f08ea5fd2eb7b9db2712e5528ee45e2efa7b7fa77e0c91f1

See more details on using hashes here.

File details

Details for the file phthos_eval-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: phthos_eval-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 29.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for phthos_eval-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 2d07459b10907fa63aa40afcdefdfc1a57f620b278f9a1a011883433bde45810
MD5 5fa85640a9703e55644eed8fc9126074
BLAKE2b-256 15ee9b1bcc4736c3856fcf7a835189a516be076c6e75aa8f42abab11eb507613

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page