Skip to main content

CooperBench

arXiv Website Dataset PyPI License: MIT

Can AI agents work together as teammates? CooperBench is the first benchmark designed to measure how well AI agents can cooperate when handling individual tasks with potential conflicts.

We find that coordinating agents perform much worse than a single agent given the same total workload. This coordination deficit presents a fundamental barrier to deploying AI systems that can work alongside humans or other agents.

Installation

pip install cooperbench

For development:

git clone https://github.com/cooperbench/CooperBench.git
cd CooperBench
pip install -e ".[dev]"

Requirements

  • Python 3.12+
  • Execution Backend (choose one):
    • Modal (default, cloud-based)
    • GCP (Google Cloud Platform)
    • Docker (local execution)
  • Redis (for inter-agent communication in coop mode)

Setup

Option 1: Modal (Default)

  1. Modal: Sign up at modal.com and run modal setup
  2. Redis: Run locally with docker run -p 6379:6379 redis:7 or use a cloud provider
  3. LLM API keys: Set in .env file:
ANTHROPIC_API_KEY=your_key
OPENAI_API_KEY=your_key
GEMINI_API_KEY=your_key

Option 2: GCP (Recommended for Scale)

Prerequisites: Install gcloud CLI first

  • macOS: brew install google-cloud-sdk
  • Linux: curl https://sdk.cloud.google.com | bash

Setup:

# 1. Install GCP dependencies
pip install 'cooperbench[gcp]'

# 2. Run configuration wizard (handles authentication, project setup, validation)
cooperbench config gcp

# 3. You're ready to run experiments!
cooperbench run --backend gcp -s lite

Also needed: Redis and LLM API keys (same as Option 1)

See GCP Setup Guide for detailed instructions.

Dataset

Download the benchmark dataset from HuggingFace into ./dataset:

cooperbench prepare

Quick Start

CLI

Run agents on a task:

# Run cooperative agents (N peers, shared Redis messaging)
cooperbench run -n my-experiment -r llama_index_task -m gpt-4o

# Run solo agent (1 agent handling both features)
cooperbench run -n my-experiment -r llama_index_task -m gpt-4o --setting solo

# Run team mode (lead + members, shared task list + scratchpad)
cooperbench run -n my-experiment -r llama_index_task -m gpt-4o --setting team

# Evaluate results
cooperbench eval -n my-experiment

Python API

from cooperbench import run, evaluate

# Run agents
run(
    run_name="my-experiment",
    repo="llama_index_task",
    model_name="gpt-4o",
    setting="coop",  # or "solo", "team"
)

# Evaluate patches
evaluate(run_name="my-experiment")

Settings: solo / coop / team

CooperBench supports three settings, selected via --setting:

  • solo — one agent implements every feature in the task. The agent works alone in a single container; no Redis, no git server.
  • coop — N peer agents, each assigned one feature. Containers talk via Redis (coop-send / coop-recv shell commands inside the container, auto-injected as user messages in the Python-loop adapters). Optional --git enables a shared team git remote so agents can fetch/merge each other's branches.
  • team — N agents organized as one lead + N-1 members, with a Redis-backed shared task list (atomic claim via coop-task-claim), a shared scratchpad volume mounted at /workspace/shared in every container, and role-specific system prompts. The lead's prompt directs them to organize work via coop-task-create / coop-task-list before coding; members claim open tasks and report progress via coop-task-update.

Team-mode result.json includes a metrics block with coordination indicators computed from the task-list audit log: tasks_total, tasks_done, unowned_at_end, time_to_first_claim_seconds, claims_per_agent, updates_per_agent.

All three settings are evaluated the same way (cooperbench eval): per-feature tests against either the single agent's patch (solo) or each agent's individual patch (coop / team).

Running with Harbor

CooperBench is also available as a Harbor adapter, which provides parallelized cloud execution on Modal via Docker-in-Docker sandboxes, built-in oracle validation, and standardized result collection.

Requires Modal authentication (run modal setup) and an LLM API key.

Install and Prepare Tasks

# Install Harbor
uv tool install harbor

# Clone Harbor and prepare the adapter
git clone https://github.com/harbor-framework/harbor.git
cd harbor/adapters/cooperbench
uv sync

# Generate flash subset (50 pairs) with openhands-sdk harness
uv run python -m cooperbench.main \
  --subset flash \
  --agent-harness openhands-sdk \
  --output-dir ../../datasets/cooperbench

Run on Modal

cd ../..  # back to harbor root

# Set up .env with your API key
echo "GEMINI_API_KEY=your_key" > .env

# Run the flash subset (50 tasks, concurrency 10)
uv run harbor run -p datasets/cooperbench --agent nop -e modal \
  --env-file .env --n-concurrent 10 \
  --ae COOPERBENCH_MODEL=gemini/gemini-3-flash-preview

# Run oracle (validates infrastructure, expects 100% pass)
uv run harbor run -p datasets/cooperbench --agent oracle -e modal \
  --env-file .env --n-concurrent 28

See the Harbor CooperBench adapter for full documentation.

CLI Reference

cooperbench config

Configure execution backends (GCP, Modal, etc.).

# Configure GCP backend
cooperbench config gcp

# Skip validation tests for faster setup
cooperbench config gcp --skip-tests

See GCP Setup Guide for details.

cooperbench run

Run agents on benchmark tasks.

cooperbench run -n NAME [OPTIONS]
Option Description Default
-n, --name Experiment name (required) -
-r, --repo Filter by repository all
-t, --task Filter by task ID all
-f, --features Feature pair (e.g., 1,2) all pairs
-m, --model LLM model gemini/gemini-3-flash-preview
-a, --agent Agent framework mini_swe_agent
-c, --concurrency Parallel tasks 20
--setting coop or solo coop
--backend modal, docker, or gcp modal
--redis Redis URL redis://localhost:6379
--git Enable git collaboration disabled
--no-messaging Disable agent messaging enabled
--force Rerun existing results skip
--agent-config Path to agent config file none

Agent Configuration: Pass agent-specific parameters via a config file. CooperBench forwards the file path to your agent without parsing it.

cooperbench eval

Evaluate completed runs.

cooperbench eval -n NAME [OPTIONS]
Option Description Default
-n, --name Experiment name (required) -
-r, --repo Filter by repository all
-t, --task Filter by task ID all
-f, --features Feature pair (e.g., 1,2) all pairs
-c, --concurrency Parallel evaluations 10
--backend modal, docker, or gcp modal
--force Re-evaluate existing skip

Experiment Settings

Setting Agents Description
coop 2 Two agents with Redis messaging, each handles one feature
solo 1 Single agent handles both features sequentially

Dataset Structure

dataset/
  <repo_name>/
    task<id>/
      setup.sh          # Repository setup script
      run_tests.sh      # Test runner script
      feature1/
        feature.md      # Feature description
        feature.patch   # Golden implementation
        tests.patch     # Test cases
      feature2/
        ...

Output Structure

Results are saved to logs/:

logs/<run_name>/<repo>/task<id>/features_<i>_<j>/
  agent1/
    trajectory.json     # Full agent trajectory
    patch.diff          # Generated patch
  agent2/
    ...
  eval.json             # Evaluation results

Benchmark Statistics

Metric Value
Tasks 652
Repositories 12
Languages Python, TypeScript, Go, Rust

Key Findings

  1. Agents perform worse together than alone — GPT-5 and Claude Sonnet 4.5 achieve only 25% success with two-agent cooperation, roughly 50% lower than when a single agent handles both tasks.

  2. Communication reduces conflicts but not failures — Agents spend up to 20% of their budget on communication, reducing merge conflicts but not improving overall success.

  3. Three capability gaps underlie coordination failures:

    • Expectation failures (42%) — agents fail to integrate partner state information
    • Communication failures (26%) — questions go unanswered, breaking decision loops
    • Commitment failures (32%) — agents break promises or make unverifiable claims

Development

# Install dev dependencies
pip install -e ".[dev]"

# Run tests
pytest tests/ -v

# Run integration tests (requires Modal)
pytest tests/ -v --run-modal

# Lint
ruff check src/
ruff format src/

# Type check
mypy src/cooperbench/

Citation

@article{cooperbench2026,
  title={CooperBench: Why Coding Agents Cannot be Your Teammates Yet},
  author={Khatua*, Arpandeep and Zhu*, Hao and Tran†, Peter and Prabhudesai†, Arya
          and Sadrieh†, Frederic and Lieberwirth†, Johann K. and Yu, Xinkai
          and Fu, Yicheng and Ryan, Michael J. and Pei, Jiaxin and Yang, Diyi},
  journal={arXiv preprint},
  year={2026},
  url={https://arxiv.org/abs/2601.13295},
  note={*Equal contribution (Stanford) · †Equal contribution (SAP Labs)}
}

License

MIT

Metadata

Release files for cooperbench 0.0.29

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cooperbench 0.0.29
File Size Uploaded
cooperbench-0.0.29.tar.gz 979.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cooperbench 0.0.29
File Interpreter ABI Platform
cooperbench-0.0.29-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / cooperbench-0.0.29.tar.gz

Download URL cooperbench-0.0.29.tar.gz
Size 979.7 kB
Tags Source
SHA-256 checksum
How to use checksums
c31badb2bd01ac19331fe8c9808ae11f00e61233e11c12be047254995efb30ae
BLAKE2b-256 checksum
How to use checksums
f7a96f7854789fef64348d3c7307506377b19110d5ae8ffb71f4521fec9f5cd9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 15, 2026.

Transparency log

Release files / cooperbench-0.0.29-py3-none-any.whl

Download URL cooperbench-0.0.29-py3-none-any.whl
Size 828.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0f8f9501a89cfc271b71f9001aeb32bd78db68062d133e999eff34d1d8f63fc9
BLAKE2b-256 checksum
How to use checksums
cdcae83ab58b31bbe7f4d61ef4ea713474bb458064e60f761668947529ff719f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 15, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.0.29 This release

2 release files

0.0.28

2 release files

0.0.27

2 release files

0.0.26

2 release files

0.0.25

2 release files

0.0.19

2 release files

0.0.18

2 release files

0.0.17

2 release files

0.0.16

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.13

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page