Skip to main content
Judgment Logo

The Continuous-Improvement Stack for Agents

Detect failures, triage root causes, and ship fixes backed by production data.

PyPI Docs

X LinkedIn

Overview

Judgeval is an open-source Python SDK for agent improvement. It provides tracing and agent-judge evaluation for LLM-powered applications — so you can detect failures, understand what went wrong, and validate fixes against real production cases before shipping.

To get started, dive into the docs.

Why Judgeval

OpenTelemetry-based tracing -- Instrument any function with @Tracer.observe(). Automatically captures inputs, outputs, and LLM token usage. Built on OpenTelemetry for full compatibility with existing observability stacks.

Agent judges -- Define prompt-based scorers to evaluate agent behaviors at scale. Judges produce structured behaviors — scored, labeled outputs that describe how your agent acted — which accumulate into a searchable record of agent behavior over time. Run judges against live production traffic or replay them on historical traces to validate fixes before shipping.

Online monitoring -- Automatically score live production traffic server-side with no latency impact. Detected behaviors surface as structured signals — configure Slack alerts so regressions and recurrences never go unnoticed.

Broad integrations -- Auto-instrumentation for OpenAI, Anthropic, Google GenAI, and Together AI. Framework support for LangGraph, OpenLit, and Claude Agent SDK.

Quickstart

Install the SDK:

pip install judgeval

Set your credentials:

export JUDGMENT_API_KEY=...
export JUDGMENT_ORG_ID=...

Add observability to your agent with two lines of setup:

from judgeval import Tracer, wrap
from openai import OpenAI

Tracer.init(project_name="my-project")
client = wrap(OpenAI())

@Tracer.observe(span_type="tool")
def search(query: str) -> str:
    results = vector_db.search(query)
    return results

@Tracer.observe(span_type="agent")
def run_agent(question: str) -> str:
    context = search(question)
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": f"{context}\n\n{question}"}],
    )
    return response.choices[0].message.content

run_agent("What is the capital of the United States?")

SQL

Use Judgeval.sql(sql_text) for read-only queries against Judgment's virtual schema, which abstracts the underlying storage. The server validates incoming queries, rejects writes, and enforces organization and project scope. The client uses its existing API key, organization membership, and resolved project. Viewer access and the public query rate limit apply.

from judgeval import Judgeval

client = Judgeval(project_name="my-project")
print(client.discover_schema())
result = client.sql("SELECT count() AS run_count FROM telemetry.traces")
print(result["rows"])

client.discover_schema() returns a Markdown string with the server's generated tables, column types and descriptions, row semantics, examples, and query limits, using the same reference as MCP discover_schema. It contains no project data and requires organization viewer access, but no resolved project or public query opt-in. The HTTP equivalent is GET /v1/sql/schema, which returns {"schema": "...Markdown reference..."}.

For query execution, use POST /v1/projects/{projectId}/sql with Authorization: Bearer <api-key>, X-Organization-Id: <organization-id>, and JSON body {"sql": "SELECT count() AS run_count FROM telemetry.traces"}. Organization and project scope are derived by the server. Use SQL predicates on supported catalog columns and LIMIT to narrow results. Physical tables, writes, multiple statements, and caller-specified execution limits are unsupported. DAL catalog allowlists, tenant isolation, and result limits of 1,000 rows and 5 MiB apply; over-limit results return an error. SQL text must contain a non-whitespace character and cannot exceed 50,000 characters.

The response contains catalog_version, columns (name, type, nullable), rows, row_count, and elapsed_ms. SQL integers outside JavaScript's safe range (-(2**53 - 1) to 2**53 - 1) arrive as exact decimal strings, including inside nested arrays and objects. For example, 9007199254740993 arrives as "9007199254740993"; use int(value) when you need a Python integer. Small integers and floating-point values remain numbers, and column types retain their original SQL types. Validation and execution errors use the SDK's existing exception mapping.

Integrations

Supports OpenAI, Anthropic, Google GenAI, Together AI, LangGraph, OpenLit, and Claude Agent SDK. See the full integrations docs.

CLI

Manage agents, traces, judges, behaviors, and evaluations from the terminal. Query trace history, deploy judges, inspect detected behaviors, and run evals against production data — all without leaving your shell. See the CLI repo and docs.

MCP Server

Connect Judgment to any MCP-compatible AI tool. Query agent traces, invoke judges, browse detected behaviors, and surface failures directly inside your AI assistant or IDE. See the docs.


Judgeval is created and maintained by Judgment Labs.

Metadata

Release files for judgeval 1.3.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for judgeval 1.3.3
File Size Uploaded
judgeval-1.3.3.tar.gz 118.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for judgeval 1.3.3
File Interpreter ABI Platform
judgeval-1.3.3-py3-none-any.whl Python 3 none any Details

Total release size: 323.6 kB

Release files / judgeval-1.3.3.tar.gz

Download URL judgeval-1.3.3.tar.gz
Size 118.9 kB
Tags Source
SHA-256 checksum
How to use checksums
a5ffa859fc93f13b38d441ab9aed5cfd5bc6505471f0a059f050211c61b53e46
BLAKE2b-256 checksum
How to use checksums
ec6967ddc986cb316f4ad2e2d6b3b87e694206a8399b45f8ab3bd2ab6bd167e4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / judgeval-1.3.3-py3-none-any.whl

Download URL judgeval-1.3.3-py3-none-any.whl
Size 204.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6be68d21809b5ff6e997458b8e2cb1130ffbab75284f9531e09975a68062a272
BLAKE2b-256 checksum
How to use checksums
b90438261fea3da090d934aa00b515e9c26fbbe4bfe848a81452357642cbe1e4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

This release

1.3.3 This release

2 release files

1.3.2

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.32.1

2 release files

0.32.0

2 release files

0.31.0

2 release files

0.30.0

2 release files

0.28.1

2 release files

0.28.0

2 release files

0.27.1

2 release files

0.26.0

2 release files

0.25.1

2 release files

0.25.0

2 release files

0.24.3

2 release files

0.24.2

2 release files

0.24.1

2 release files

0.24.0

2 release files

0.23.9

2 release files

0.23.8

2 release files

0.23.2

2 release files

0.23.1

2 release files

0.23.0

2 release files

0.22.8

2 release files

0.22.7

2 release files

0.22.6

2 release files

0.22.5

2 release files

0.22.4

2 release files

0.22.3

2 release files

0.20.1

2 release files

0.20.0

2 release files

0.19.0

2 release files

0.18.0

2 release files

0.17.0

2 release files

0.16.8

2 release files

0.16.7

2 release files

0.16.6

2 release files

0.16.5

2 release files

0.16.4

2 release files

0.14.1

2 release files

0.14.0

2 release files

0.13.1

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.4

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

0.0.55

2 release files

0.0.54

2 release files

0.0.53

2 release files

0.0.52

2 release files

0.0.51

2 release files

0.0.44

2 release files

0.0.43

2 release files

0.0.42

2 release files

0.0.40

2 release files

0.0.39

2 release files

0.0.38

2 release files

0.0.37

2 release files

0.0.35

2 release files

0.0.34

2 release files

0.0.33

2 release files

0.0.32

2 release files

0.0.31

2 release files

0.0.30

2 release files

0.0.29

2 release files

0.0.28

2 release files

0.0.25

2 release files

0.0.24

2 release files

0.0.23

2 release files

0.0.22

2 release files

0.0.21

2 release files

0.0.20

2 release files

0.0.19

2 release files

0.0.18

2 release files

0.0.13

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page