Skip to main content

agent-eval-rpc Python Client

agent-eval-rpc lets Python programs call the judging and ingestion APIs implemented by @tangle-network/agent-eval. The Python package validates requests and responses with Pydantic. The Node package owns rubric execution, model calls, and scoring.

Install

Python 3.10 or newer and Node.js 20 or newer are required. Install matching package versions:

pip install agent-eval-rpc
npm install --global @tangle-network/agent-eval

Configure an OpenAI-compatible model endpoint for judge calls:

export AGENT_EVAL_LLM_BASE_URL=https://api.openai.com/v1
export AGENT_EVAL_LLM_API_KEY="$YOUR_API_KEY"
export AGENT_EVAL_LLM_MODEL=gpt-4.1-mini

OPENAI_BASE_URL, OPENAI_API_KEY, and OPENAI_MODEL are also accepted. The endpoint receives the content, rubric, and context passed to client.judge().

Judge Content

from agent_eval_rpc import Client

client = Client()
result = client.judge(
    content="The retry budget is checked before each provider call.",
    rubric_name="anti-slop",
)

print(result.composite)
print(result.dimensions)
print(result.failure_modes)
print(result.rationale)

Client() first checks for an HTTP server at http://127.0.0.1:5005. If none is running, it invokes agent-eval rpc as a subprocess. Inspect client.transport to see which path was selected.

For repeated or concurrent calls, start the server once:

agent-eval serve --port 5005

Then force HTTP from Python when desired:

client = Client(transport="http", base_url="http://127.0.0.1:5005")

Define A Rubric

Use a built-in rubric by name or pass an inline rubric. Exactly one is required.

from agent_eval_rpc import Client, FailureMode, Rubric, RubricDimension

rubric = Rubric(
    name="commit-message",
    description="Checks whether a commit message explains why the change exists.",
    systemPrompt="Score the commit message using the supplied response schema.",
    dimensions=[
        RubricDimension(
            id="explains_why",
            description="The message states the reason for the change.",
            weight=1.0,
        ),
    ],
    failureModes=[
        FailureMode(
            id="what-only",
            description="The message states the edit without its reason.",
        ),
    ],
)

result = Client().judge(content="fix retry accounting", rubric=rubric)

List the built-in rubrics and their version hashes:

for rubric in Client().list_rubrics().rubrics:
    print(rubric.name, rubric.rubric_version)

Client Options

Client(
    base_url: str | None = None,
    cli_path: str | None = None,
    transport: "auto" | "http" | "subprocess" = "auto",
    timeout_s: float = 200.0,
)

client.judge() returns:

Field Meaning
composite Weighted score from 0 to 1
dimensions Score for each rubric dimension
failure_modes Detected negative-pattern IDs
wins Detected positive-pattern IDs
rationale Model explanation
rubric_version Stable rubric hash used for comparison
model Model reported by the provider
duration_ms Total call duration

Hosted Event Ingestion

HostedClient sends evaluation events and trace spans to a server that implements the hosted ingest format. This is separate from Client, which calls the local judging API.

from agent_eval_rpc import HostedClient

with HostedClient(
    endpoint="https://your-ingest.example",
    api_key="tenant-token",
    tenant_id="acme",
) as client:
    response = client.ingest_eval_run(event)
    assert response.accepted == 1

Review hosted.py for the typed event fields and retry behavior. The event payload can include run paths, scenario IDs, candidate values, scores, errors, costs, summaries, and trace attributes.

Official Optimizer Bridges

The TypeScript campaign API can run official GEPA and SkillOpt through this package. The Python client does not reimplement either algorithm.

GEPA

Install the client and published GEPA package for the standard engine:

python -m pip install agent-eval-rpc
python -m pip install "gepa[full]==0.1.4"

The published package runs direct recipes with the standard gepa engine. Sequential, adaptive, best-of, vote, Omni, AutoResearch, Meta Harness, and Best-of-N currently require this tested official source revision:

python -m pip install \
  "gepa[full] @ git+https://github.com/gepa-ai/gepa.git@f919db0a622e2e9f9204779b81fe00cc1b2d808f"

From an Agent Eval source checkout:

uv sync --frozen --group gepa-release
uv sync --frozen --group gepa-source

The bridge calls GEPA's official engine and composition functions. It supports direct engine, sequential, adaptive sequential, best-of, vote, and Omni recipes. The bridge forwards the run seed into every standard gepa engine configuration at engine.seed. Agent engines accept no seed parameter, so the output reports seedApplied: false for recipes that include one. A caller-supplied engineConfig.engine.seed is rejected because the run seed owns that field. GEPA receives only the serialized train and selection cases supplied by the caller. compareOptimizationMethods() keeps final cases in TypeScript and evaluates them only after GEPA exits. For a direct standard engine, the bridge writes a digest-addressed candidate population artifact. The artifact preserves GEPA's candidate indices, parent indices, selection scores, and discovery evaluation counts. It does not include final cases or large rollout outputs.

Every engine run requires an evaluation limit and an optimizer-model dollar limit. Agent Eval enforces callback counts before executing an agent or judge. For standard GEPA engines, the TypeScript optimizer option routes reflection through Agent Eval's local model proxy. The proxy enforces whole-run request and dollar limits, keeps the provider key out of Python, and records exact provider usage in the shared cost log. With optimizer.anthropicEndpoint: true, the AutoResearch and Meta Harness agent engines run fully metered too: the proxy serves an Anthropic Messages route, and every claude CLI call is admitted, receipted, and budget-enforced. The canonical description of that path is docs/campaign-proposers.md. An unproxied engine can still receive its native configuration; its external model spend remains incomplete unless that engine reports it.

SkillOpt

Install the client and the exact SkillOpt source revision tested by Agent Eval:

python -m pip install agent-eval-rpc
python -m pip install \
  "skillopt @ git+https://github.com/microsoft/SkillOpt.git@61735e3922efc2b90c6d6cab561e62e98452ca90"

From an Agent Eval source checkout, install the locked package with:

uv sync --frozen --group skillopt-source

The published skillopt==0.2.0 wheel omits the 21 prompt files required by ReflACTTrainer. The source revision contains those files and is checked before each release.

skillOptOptimizationMethod() runs SkillOpt's ReflACTTrainer with an Agent Eval environment adapter. The adapter sends candidate and case pairs back to the TypeScript process for execution and scoring. It disables SkillOpt's test split because final cases remain private to compareOptimizationMethods().

The TypeScript method requires:

  • an OpenAI-compatible endpoint and key in the TypeScript method's optimizer option,
  • exact input and output rates,
  • maximum model dollars, requests, request bytes, response bytes, and output tokens,
  • a maximum candidate evaluation count.

Agent Eval starts a local proxy, gives SkillOpt only the proxy credential, checks every request before forwarding it, and records provider token usage in the shared cost log. Before optimization starts, it records and hashes the installed optimizer source, bridge source, Python runtime, runner settings, endpoint settings, data, and evaluation ID into one run ID. The Python process checks that identity again before it restores any state. Missing provider usage fails the run instead of assuming zero cost.

DSPy

Install DSPy 3.2.1 and the Agent Eval adapters with:

python -m pip install "agent-eval-rpc[dspy]"
import dspy

from agent_eval_rpc import DspyJudgeMetric

dspy.configure_cache(restrict_pickle=True)
metric = DspyJudgeMetric(rubric_name="answer-quality")

gepa = dspy.GEPA(
    metric=metric.feedback,
    reflection_lm=dspy.LM("openai/gpt-4.1-mini"),
    max_metric_calls=100,
)
optimized = gepa.compile(program, trainset=train, valset=selection)

mipro = dspy.MIPROv2(metric=metric, auto="light")

Use metric.feedback for dspy.GEPA. It returns dspy.Prediction(score=..., feedback=...) with dimension scores, failure modes, wins, and rationale. Use the metric object directly for MIPROv2, SIMBA, bootstrap, and evaluation APIs that expect a number. Identical calls share one judge result, including concurrent calls. DspyJudgeMetric rejects DSPy's default unrestricted disk-cache pickle handling. Configure the official restricted cache as shown above, or call dspy.configure_cache(enable_disk_cache=False) before creating the metric.

DSPy programs should use DSPy's official optimizers directly. Agent Eval also runs the official dspy.RLM for recursive trace analysis. The TypeScript API starts the Python bridge, provides authenticated trace tools, calls a caller-owned model execution path, enforces model and trace-read limits, and validates cited findings. The Python bridge owns no trace storage or model credentials.

import {
  createDspyRlmTraceEngine,
  type DspyRlmTraceEngineOptions,
} from '@tangle-network/agent-eval/analyst'
import { analyzeTraces } from '@tangle-network/agent-eval/traces'

type ModelOwner = Pick<
  DspyRlmTraceEngineOptions,
  'call' | 'callRef' | 'recordExecution'
>

export async function analyzeRun(modelOwner: ModelOwner) {
  const engine = createDspyRlmTraceEngine({
    ...modelOwner,
    model: 'deepseek-v4-flash',
    pricing: {
      inputUsdPerMillion: 3,
      outputUsdPerMillion: 15,
    },
    runner: { command: '.venv/bin/python' },
  })

  return analyzeTraces(
    { question: 'What first caused this run to fail?' },
    { source: 'run.otlp.jsonl', engine, toolGroup: 'singleTrace' },
  )
}

See Trace Analysis for custom definitions, limits, result fields, and the public quality benchmark.

DSPy 3.2.1 pins GEPA 0.0.27. The general Optimize Anything bridge uses GEPA 0.1.4, so repository checks install them in separate environments:

uv sync --frozen --extra dev --group gepa-release
AGENT_EVAL_EXPECT_GEPA_RELEASE=1 \
  uv run --frozen --extra dev --group gepa-release \
  pytest tests/test_gepa_release_compatibility.py tests/test_gepa_bridge.py

uv sync --frozen --extra dev --group skillopt-source --group gepa-source
uv run --frozen pytest

uv sync --frozen --extra dev --extra dspy
uv run --frozen pytest tests/test_dspy_metric.py

The bridge records the installed upstream package version and source revision with each run. SkillOpt and a direct GEPA engine can restore official state only when the package revision, settings, starting candidate, described data, evaluation ID, and seed match. Direct GEPA resume also requires trustResumeState: true because its upstream checkpoint uses Python pickle. Enable it only for a checkpoint created locally in a directory you control. Composed GEPA recipes restart and never claim that upstream state was restored.

The official optimizer subprocess requires POSIX process-group cleanup. Use Linux or WSL rather than native Windows so a timeout can terminate the complete Python process tree.

Errors

Exception Meaning
ValidationError The request does not match the Python or server schema
RubricNotFoundError The named built-in rubric does not exist
TransportError The HTTP server or subprocess could not be reached
AgentEvalError Base class for client errors

Errors include .code and .details when the server returned structured error data.

Versions

The Python and npm packages are released with the same version. Use client.version() to check the running Node package and wire-format version:

version = Client().version()
print(version.version, version.wire_version)

Development

cd clients/python
pip install -e ".[dev]"
pytest

Run the cross-language tests after building the Node package:

cd ../..
pnpm build
cd clients/python
pytest

The runnable Python example is examples/judge_anti_slop.py.

Release files for agent-eval-rpc 0.150.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-eval-rpc 0.150.2
File Size Uploaded
agent_eval_rpc-0.150.2.tar.gz 457.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-eval-rpc 0.150.2
File Interpreter ABI Platform
agent_eval_rpc-0.150.2-py3-none-any.whl Python 3 none any Details

Total release size: 528.0 kB

Release files / agent_eval_rpc-0.150.2.tar.gz

Download URL agent_eval_rpc-0.150.2.tar.gz
Size 457.9 kB
Tags Source
SHA-256 checksum
How to use checksums
9bc5ff9eb487670b3901f121840346526e43534c8487dfda43f45cd215271226
BLAKE2b-256 checksum
How to use checksums
608b9f1f7a9abc113afa3f5739d2789d96e25135aee0d1ee7cfb21fb5a8d73fd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 20, 2026.

Transparency log

Release files / agent_eval_rpc-0.150.2-py3-none-any.whl

Download URL agent_eval_rpc-0.150.2-py3-none-any.whl
Size 70.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8ae85d2043f20d30ea9d74d40c30571d416b9a98d62d18bf240352608a75d554
BLAKE2b-256 checksum
How to use checksums
7308925b130df60034373e6b942d2c61bf626ea47c7fc4997001847a88ea6368
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 20, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.150.2 This release

2 release files

0.99.0

2 release files

0.98.0

2 release files

0.97.0

2 release files

0.96.5

2 release files

0.96.2

2 release files

0.96.1

2 release files

0.96.0

2 release files

0.95.1

2 release files

0.95.0

2 release files

0.94.0

2 release files

0.93.0

2 release files

0.92.0

2 release files

0.91.0

2 release files

0.90.1

2 release files

0.90.0

2 release files

0.89.0

2 release files

0.86.0

2 release files

0.59.1

2 release files

0.59.0

2 release files

0.58.2

2 release files

0.58.1

2 release files

0.58.0

2 release files

0.57.0

2 release files

0.56.0

2 release files

0.55.0

2 release files

0.54.0

2 release files

0.53.0

2 release files

0.52.0

2 release files

0.51.0

2 release files

0.50.2

2 release files

0.50.1

2 release files

0.50.0

2 release files

0.49.0

2 release files

0.48.0

2 release files

0.42.0

2 release files

0.41.0

2 release files

0.40.5

2 release files

0.40.4

2 release files

0.40.3

2 release files

0.40.2

2 release files

0.40.1

2 release files

0.34.1

2 release files

0.34.0

2 release files

0.32.0

2 release files

0.31.1

2 release files

0.31.0

2 release files

0.29.1

2 release files

0.29.0

2 release files

0.28.0

2 release files

0.27.2

2 release files

0.27.0

2 release files

0.25.0

2 release files

0.24.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page