Skip to main content

Evaluate, benchmark, and A/B-test AI coding agents (Claude Code, Codex, Gemini/Antigravity) with sandboxed, reproducible YAML task suites.

Project description

Coder Eval — evaluate & benchmark AI coding agents and Claude Code skills

PyPI Docs License: Apache 2.0 Python 3.13+ CI Code style: Ruff

Coder Eval (pip install coder-eval / uv tool install coder-eval) is an open-source framework for evaluating and benchmarking AI coding agents and their skills — built for CLI and skill builders — with sandboxing, reproducibility, and data-driven analysis. It runs a real agent (Claude Code, Codex, or Google Antigravity / Gemini) in a sandbox against declarative YAML tasks, then scores the files and commands it actually produced. Not an "agentic coding" benchmark: it measures how effective your CLI and skills are when used by coding agents.

Reach for it when you want to test whether a Claude Code skill triggers, A/B-test Claude Code vs. Codex vs. Gemini (or model vs. model, prompt vs. prompt), or gate CI on coding-agent quality. Unlike fixed datasets (SWE-bench, SkillsBench) that rank models on a shared leaderboard, Coder Eval evaluates the tasks, skills, and workflows you ship — with weighted 0.0–1.0 criteria, a skill_triggered activation check, an A/B experiment layer, and per-tool cost telemetry. See How it compares. 📚 Full docs: uipath.github.io/coder_eval.

Coder Eval running the hello_date task: a sandboxed agent writes and runs a script from a YAML task, then the scored result is browsed in evalboard

  • Declarative YAML tasks with pinned dependencies and clear success criteria
  • Sandboxed execution in isolated environments with resource limits
  • Weighted, continuous scoring (0.0–1.0) with fractional credit and thresholds
  • Many criterion types — from file checks to code similarity and LLM-graded rubrics
  • Agent abstraction — Claude Code, Codex, and Antigravity (Gemini) today, extensible via a plugin SPI
  • Experiment layer — A/B agent configs (models, tools, prompts) side-by-side
  • Full telemetry — every tool call, token counts, and cost, with real-time streaming

What you can do with it

  • Benchmark coding agents — score an agent across a suite of tasks with weighted scoring and pass/fail thresholds
  • Compare models & configs — A/B-test Claude vs. Codex vs. Gemini, model vs. model, tool-on vs. tool-off, prompt vs. prompt
  • Evaluate skills — verify an agent actually engages a target skill (skill_triggered) and score skill-driven suites (SkillsBench-style)
  • Keep skills up to date in CI — re-validate your skills on every change or on a schedule; catch silent regressions when models, prompts, or the skills themselves drift
  • Gate CI on agent quality — run the suite in GitHub Actions and fail the build on regressions
  • Bring your own dataset — fan one task out over many rows for larger benchmark suites

Keeping skills fresh? Run Coder Eval as a scheduled GitHub Actions job so your skills are continuously re-evaluated against the latest model — a skill that quietly stops triggering surfaces as a failing criterion before your users hit it. See Tutorial 02 — Running coder_eval in CI.

Quick Start

Prerequisites: Python 3.13+, uv 0.8+, and the Claude CLI (brew install claude). Developed on macOS; CI runs on Linux.

git clone https://github.com/UiPath/coder_eval.git
cd coder_eval

uv sync --extra dev          # install core + dev tools
cp .env.example .env         # then set ANTHROPIC_API_KEY — or skip that: an
                             # existing Claude Code login (`claude login`) is
                             # picked up automatically

uv run coder-eval plan tasks/hello_date.yaml   # validate (no tokens spent)
uv run coder-eval run  tasks/hello_date.yaml   # run your first evaluation
uv run coder-eval report runs/latest           # view the result

New here? Follow Tutorial 01 — Your First Evaluation.

The optional [uipath] extra (uv sync --extra dev --extra uipath) adds the in-host uipath SDK for local sandbox parity; it installs from public PyPI (no credentials required). Without it the framework runs end-to-end; uipath-dependent features fail at dispatch with a clear hint.

Using Coder Eval in CI or another project? Install the published package instead of cloning:

uv tool install coder-eval    # puts the `coder-eval` CLI on your PATH,
                              # in its own isolated environment

uv tool install "coder-eval[codex,antigravity]"   # same, with agent extras
coder-eval --version                              # verify the install

To add it as a project dependency instead: uv add coder-eval or pip install coder-eval. In a real CI gate, pin to a specific released version so a harness upgrade can't silently move your results. (The example tasks/ live in this repo — clone it or point the CLI at your own task files.) See Tutorial 02 — Running coder_eval in CI for the full setup.

Telemetry

📊 Usage telemetry is on by default. coder-eval sends anonymous usage telemetry (command names, outcomes, counts, durations, an anonymous install id, platform info) to help improve the tool. It never captures prompts, file contents, or repo paths, and prints a one-time notice on first run. To disable it, set TELEMETRY_ENABLED=false in your .env or environment. See Usage Telemetry for details and how to route it to your own resource.

Documentation

Guide What's in it
Tutorials Step-by-step walkthroughs — start here
User Guide Full CLI, configuration, output, and environment-variable reference
Task Definition Guide The task-file schema — all criterion types, scoring, templates
A/B Experiments Compare models / tools / prompts across the same tasks
Bring Your Own Dataset Fan a single task out over a dataset
Codex Agent Guide Running the Codex agent
Docker Isolation The container sandbox driver
CLAUDE.md Architecture, key patterns, and extension points
CONTRIBUTING.md Dev setup, quality bar, and how to contribute

How it compares

  • vs. fixed benchmarks (SWE-bench, SkillsBench) — they score a canonical dataset; Coder Eval scores your tasks with continuous 0.0–1.0 weighted criteria (and can still wrap a fixed dataset via Bring Your Own Dataset).
  • vs. large-scale / RL harnesses (Harbor) — Harbor targets scale and RL rollouts; Coder Eval targets weighted, skill-aware suites gated in CI.
  • vs. model-output eval tools (OpenAI Evals) — they grade model text; Coder Eval runs a full agent in a sandbox and scores the files and commands it produced.
  • vs. hand-rolled scripts — reproducible sandboxes, weighted criteria, cost/token telemetry, A/B experiments, and CI-ready pass/fail gates out of the box.

See the full comparison — with sources.

Task Definition

A task is a YAML file: a prompt, the agent config, a sandbox, and success criteria.

task_id: "hello_world"
description: "Create a Python script that prints Hello, World!"
initial_prompt: "Create hello.py that prints 'Hello, World!'"

agent:
  type: "claude-code"
  permission_mode: "acceptEdits"
  allowed_tools: ["Read", "Write", "Bash"]

sandbox:
  driver: "tempdir"
  python: {}

success_criteria:
  - type: "file_exists"
    path: "hello.py"
    description: "hello.py must be created"
  - type: "run_command"
    command: "python hello.py"
    timeout: 10
    description: "Script must execute successfully"

Tasks can omit the agent section entirely — defaults resolve from the experiment layer (experiments/default.yaml). For the full schema and every criterion type, see the Task Definition Guide.

Tip: In Claude Code, use /coder-eval-task-create to scaffold a task from a natural-language description, and /coder-eval-run-analysis runs/latest to get improvement suggestions from a completed run.

Development

make install    # package + dev + [uipath] deps + pre-commit hooks
make verify     # format + lint + typecheck + test + coverage (CI equivalent)

Run make verify before pushing — it mirrors CI (80% coverage threshold). See CONTRIBUTING.md for the full workflow, commit conventions, and extension points (new criteria, new agents).

Known limits & non-goals

  • Not a fixed benchmark or leaderboard — Coder Eval scores your tasks and ships example tasks, not a canonical scored dataset.
  • Tasks execute real code — run untrusted tasks only under the container driver (see Docker Isolation); the tempdir driver is not a security boundary.
  • Bring your own model credentials — Anthropic, Bedrock, or Gemini keys; Coder Eval does not proxy or supply model access.
  • Python 3.13+ only.

Support & security

License

© 2026 UiPath. Licensed under the Apache License, Version 2.0 — see LICENSE and NOTICE.

Acknowledgments

Built with the Claude Agent SDK, Pydantic, Typer, and Rich.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

coder_eval-0.8.7.tar.gz (7.5 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

coder_eval-0.8.7-py3-none-any.whl (496.0 kB view details)

Uploaded Python 3

File details

Details for the file coder_eval-0.8.7.tar.gz.

File metadata

  • Download URL: coder_eval-0.8.7.tar.gz
  • Upload date:
  • Size: 7.5 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for coder_eval-0.8.7.tar.gz
Algorithm Hash digest
SHA256 eeced467dfd330ce0ac5d654b5eceb9eae7e6016c050c1ccb505b2acc5d7e585
MD5 bf9ef76675cd094b031f1b067d2ea7d7
BLAKE2b-256 cedb542e52103dd2d1b83c136e227cd76bb4036348d4a52f93a01358c2dd5636

See more details on using hashes here.

Provenance

The following attestation bundles were made for coder_eval-0.8.7.tar.gz:

Publisher: release.yml on UiPath/coder_eval

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file coder_eval-0.8.7-py3-none-any.whl.

File metadata

  • Download URL: coder_eval-0.8.7-py3-none-any.whl
  • Upload date:
  • Size: 496.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for coder_eval-0.8.7-py3-none-any.whl
Algorithm Hash digest
SHA256 6172a246b89f439a7366c0f119a38f55bdadeadd151e2399fced200a4138fcd5
MD5 ba258968367e4b43159acd9b40687377
BLAKE2b-256 25f731a8acd108b5bcde8f96efb2330fac9fbde9abb13a88bdb1aaf29ef03a23

See more details on using hashes here.

Provenance

The following attestation bundles were made for coder_eval-0.8.7-py3-none-any.whl:

Publisher: release.yml on UiPath/coder_eval

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page