Skip to main content

Plural

Plural is a Python SDK and CLI for building reproducible agent evaluations and keeping the evidence behind every score.

Evaluation in eight objects

  1. Environment: the world, its actions, private State, the Observation the Agent sees, resources, and Runtime.
  2. Verifier: deterministic, Agent, or Human scoring of a finished episode.
  3. Task: instructions bound to one Environment and its Verifiers.
  4. Benchmark: a versioned collection of Tasks and the rules for ranking results.
  5. Agent: a catalog model, instructions, and an optional Harness.
  6. Harness: the loop that connects an Agent to an Environment. An Agent with no Harness uses native, Plural's built-in tool loop.
  7. Job: one run of a Task or Benchmark with one or more Agents and attempts.
  8. Trial: one Agent on one Task for one attempt, with its trajectory, artifacts, and score.

A score is what a Verifier produces, and it is the only thing a Benchmark ranks on. Per-step rewards from an Environment are recorded for training and never contribute to a score. Only the Observation is ever shown to the model; scores, rewards, and private Verifier data never reach it.

Install

Plural requires Python 3.10 or newer.

pip install "plural>=0.15"

Try a finished project

The support queue project includes an Agent that follows keyword rules instead of calling a model, so it runs offline with no account or key. From a copy of examples/first-project:

plural benchmark validate support-triage
plural run --benchmark support-triage --agent scripted
plural job show <job-id>

The run prints the Job id and one line per Trial with its score.

Start your own project

A project is a directory with a project.yaml. Each resource is a directory named after it, and commands find the project by walking up from wherever you run them.

plural project init my-eval
cd my-eval
plural env init support-desk
plural verifier init resolved
plural task init refund --environment support-desk --verifier resolved

Each init writes a template. Fill in the places marked PLURAL-TODO, then validate. validate checks the resource and everything it depends on, and lists any template text you left unfinished.

plural task validate refund
plural run --task refund --model openai/gpt-5.6-luna --dry-run
plural run --task refund --model openai/gpt-5.6-luna

--dry-run shows the plan and the version of every input without running anything. --model without --harness uses native. Every model call goes through the Plural gateway, which bills it at the exact provider cost, so it needs a Plural API key: plural auth login --api-key-stdin or PLURAL_API_KEY. A browser login is not accepted for model calls. Every run is a new Job, recorded under .plural/jobs/<job-id>/ with every input pinned by version and content hash.

Push to Plural

Runs are local by default. To run on hosted infrastructure, register the project with a private hosted project, push what the run uses, and add --hosted:

plural auth login
plural project init my-eval --push
plural task push refund --with-deps
plural run --task refund --model openai/gpt-5.6-luna --hosted --follow

Credentials are stored in your user config directory or OS keyring, never in project files. Pushing never makes anything public; sharing is a separate, explicit action in the Plural web app. The local runtime runs code as a trusted subprocess on your machine; it is not a sandbox.

Python

The CLI and the SDK load the same resources. From inside the support queue project:

from plural import Job
from plural.project import Project, Workspace

workspace = Workspace(Project.find())
benchmark = workspace.get("benchmark", "support-triage")
agent = workspace.get("agent", "scripted")
result = Job(benchmark, agents=[agent]).run()

Read the documentation, starting with Getting started, then the support queue and Wordle tutorials. Coming from 0.14? Read Migrate to 0.15.

Development

uv sync --group dev --group docs
uv run ruff check .
uv run mypy --strict src/plural tests/typing/consumer.py
uv run pytest
uv run python scripts/check_docs.py
uv run mkdocs build --strict

License

Apache-2.0

Metadata

Release files for plural 0.19.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for plural 0.19.0
File Size Uploaded
plural-0.19.0.tar.gz 753.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for plural 0.19.0
File Interpreter ABI Platform
plural-0.19.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.2 MB

Release files / plural-0.19.0.tar.gz

Download URL plural-0.19.0.tar.gz
Size 753.1 kB
Tags Source
SHA-256 checksum
How to use checksums
a81a536808a9176405c21c5e6206d6ffcc7d1dc5814e93cd228d6bd0c33adf67
BLAKE2b-256 checksum
How to use checksums
0a05d342e33238c0e913116020b03f1a841f0136a6b4a3d6037122a5cd659aac
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / plural-0.19.0-py3-none-any.whl

Download URL plural-0.19.0-py3-none-any.whl
Size 435.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
468774f36d67734b9cfa8643ad19abc4c9d31cb22de3a65a0b5aa97f5446244a
BLAKE2b-256 checksum
How to use checksums
8ecc0a51aba8a0020608bd9303e8966ae5576069bc9d697115fe23576d7ecef8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.19.0 This release

2 release files

0.18.0

2 release files

0.17.2

2 release files

0.17.1

2 release files

0.17.0

2 release files

0.16.0

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.13.3

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.6.1

2 release files

0.6.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page