agentic-evals
A standalone, framework-agnostic evaluation and scoring engine for LLM and agent outputs.
Extracted from AgenticLens's
proven evaluation module — the same engine, usable on its own. It scores
whatever trace-shaped data you give it (see EvalTrace/EvalSpan below);
it has no dependency on any specific tracing/observability tool, and no
dependency on AgenticLens itself.
Status
Alpha. The engine (deterministic checks, LLM-as-judge and custom evaluators, a release gate, live Python/HTTP targets) is real, tested code lifted directly from AgenticLens's evaluation module. No PyPI release yet.
Why a separate package
AgenticLens's evaluation module was never actually AgenticLens-specific — scoring an LLM/agent output against expectations doesn't need AgenticLens's full trace schema, CLI, or dashboards. Pulling it out means:
agentic-sidecar,agentic-chaos, or any other project can score outputs without depending on all of AgenticLens.- Anyone with any trace-shaped data — not just AgenticLens users — can use
it, the same way Braintrust's
autoevalsdoesn't care what produced the string it's scoring.
Install
pip install agentic-evals
Core concepts
EvalTrace/EvalSpan— the minimal trace shape the engine inspects: a trace id, spans (each optionally naming atool_nameand carrying arbitraryattributes), total latency, estimated cost, and metadata. Deliberately not tied to any specific instrumentation format — build one from whatever you already have.Score— a single named judgment (0-1 value, pass/fail, explanation).Evaluator— anything with a.nameand an.evaluate(context) -> list[Score].CallableEvaluatoradapts a plain Python function;LLMJudgeEvaluatorandBusinessRuleEvaluatorare named convenience subclasses for readability/reporting.TestCase/TestSuite— declarative expectations (exact match, substring, JSON Schema, required fields, required/forbidden tool calls, required tool arguments, latency/cost/turn-count thresholds, or a named custom evaluator) plus the cases that make up a suite.evaluate_suite— runs a suite against suppliedEvaluationSamples and returns anEvaluationReport(per-case scores plus a pass-rate/cost/ latency summary).GateConfig/evaluate_gate— turn anEvaluationReportinto a pass/fail release decision on configurable thresholds.
Quickstart
from agentic_evals import (
EvalSpan,
EvalTrace,
EvaluationSample,
TestCase,
TestSuite,
evaluate_suite,
)
suite = TestSuite(
name="support-answers",
version="1",
cases=[
TestCase(
id="case-1",
name="Answer contains the right total",
expected_contains=["42"],
required_tools=["calculator"],
max_latency_ms=2000,
)
],
)
sample = EvaluationSample(
case_id="case-1",
output="The combined total is 42.",
trace=EvalTrace(
trace_id="trace-1",
total_latency_ms=350,
spans=[EvalSpan(tool_name="calculator")],
),
)
report = evaluate_suite(suite, [sample])
print(report.summary.pass_rate) # 1.0
LLM-as-judge
from agentic_evals import (
EvaluationContext,
EvaluatorConfig,
EvaluatorRegistry,
LLMJudgeEvaluator,
Score,
TestCase,
)
def judge(context: EvaluationContext) -> Score:
# Call whatever model/provider you like here.
correct = "42" in context.sample.output
return Score(
name="answer_quality",
value=0.95 if correct else 0.1,
passed=correct,
explanation="Judged against the rubric in context.config.config.",
)
registry = EvaluatorRegistry()
registry.register(LLMJudgeEvaluator("answer_quality_judge", judge))
case = TestCase(
id="case-1",
name="Answer quality",
evaluators=[EvaluatorConfig(name="answer_quality_judge", threshold=0.8)],
)
Release gates
from agentic_evals import GateConfig, evaluate_gate
decision = evaluate_gate(
report,
GateConfig(min_pass_rate=0.95, max_average_latency_ms=1500, max_total_cost_usd=0.25),
)
if not decision.passed:
raise SystemExit(f"Release gate failed: {decision.reasons}")
Never fabricates a value it can't back up: total_cost_usd on a summary or
gate decision stays None unless every case in scope has a known cost —
an incomplete cost picture is reported as unavailable, not $0.00.
Live targets
Point a suite at a real running system (a trusted Python callable, or an HTTP endpoint) instead of pre-recorded samples:
from agentic_evals import PythonTarget, run_live_suite
report = run_live_suite(suite, PythonTarget(callable_path="my_module:run_case"))
Live targets are intentionally powerful developer-facing integrations — Python targets execute local code and HTTP targets can reach arbitrary URLs. Only point them at trusted suite files and trusted target definitions.
Using it with AgenticLens's own traces
If you already have an AgenticLens Run (from its instrumentation API or
OTLP ingestion), AgenticLens itself provides the adapter —
agenticlens.evaluation.to_eval_trace(run) — so you don't have to hand-build
an EvalTrace. This package has no dependency in the other direction.
What's deliberately not here
Dataset versioning/splitting, judge calibration, and HTML report rendering
stay in AgenticLens for now — those are product features built on top of
this engine (the same way Braintrust's dataset/experiment platform is
separate from the autoevals library itself), not the engine.
License
MIT
Metadata
Release files for agentic-evals 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agentic_evals-0.1.0.tar.gz | 135.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agentic_evals-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 148.9 kB
Release files / agentic_evals-0.1.0.tar.gz
| Download URL | agentic_evals-0.1.0.tar.gz |
|---|---|
| Size | 135.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e1b2fb3a76be14ff86801c6b980672c52d8e3acf6ad3d4cdfb270dd54d3b156d
|
|
BLAKE2b-256 checksum How to use checksums |
80efdd77d5a6c95f5586319c2a219b4061478a3124da4fa6e9b7413c435bea0a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.
Transparency logRelease files / agentic_evals-0.1.0-py3-none-any.whl
| Download URL | agentic_evals-0.1.0-py3-none-any.whl |
|---|---|
| Size | 13.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
842b9862e495e20ad2ae32a25cb7a0e659da5099ea887d920b06d59f4ed3dc3a
|
|
BLAKE2b-256 checksum How to use checksums |
3fc932cd98e0ae01c6289da991db031fbf7027f9e01e5546af80961d2ddfbc48
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.
Transparency log