Skip to main content

Graded

Graded is a library to make computing rewards simple, defensive, and structured for agent evaluations, particularly within Harbor environments. It provides tools to declare structured grading criteria, execute LLM judges with automatic tracing, and manage evaluation artifacts.

Installation

pip install graded

Or using uv:

uv pip install graded

Quick Start

Create an evaluation script (e.g. verify.py) to grade a task workspace:

from pathlib import Path
from graded import Evaluator

# Initialize the evaluator
ev = Evaluator(
    workspace="/workspace",
    output_path="/logs/verifier/reward.json",
    auto_save_artifacts=True
)

# 1. Declare a standard criterion (workspace parameter is optional)
@ev.criterion(name="has_output_file", weight=1.0)
def check_output() -> bool:
    return ev.file_exists("output.txt")

# 2. Declare a fatal criterion (takes workspace Path parameter to inspect files directly)
@ev.criterion(name="no_syntax_errors", weight=2.0, fatal=True)
def check_syntax(workspace: Path) -> bool:
    # Use workspace parameter to inspect the files on disk
    return (workspace / "src").is_dir()

# 3. Declare a fractional scoring criterion
@ev.criterion(name="test_pass_rate", weight=3.0)
def check_tests() -> float:
    return 0.8  # Returns a score between 0.0 and 1.0

if __name__ == "__main__":
    ev.run()

Core Features

1. Criteria Declarations (@ev.criterion)

Define check functions using the @ev.criterion decorator.

Check functions can optionally accept the workspace directory as a pathlib.Path parameter if they need to perform custom filesystem operations. If a function does not accept any arguments, it will be executed without the workspace parameter.

  • name: Unique identifier for the criterion.
  • weight: Relative weight of the score in the final weighted average calculation.
  • fatal: If True, any score of 0.0 or False immediately short-circuits the final score to 0.0.
  • Return Value: Must return a bool, int, or float.

Programmatic Registration (Without Decorators)

If you prefer to define functions normally, you can register them programmatically without using decorator syntax:

def check_output(workspace: Path) -> bool:
    return (workspace / "output.txt").is_file()

# Register directly
ev.criterion("has_output_file", weight=1.0)(check_output)

2. LLM Judge with Automatic Tracing

Integrate with instructor to run structured, schema-validated LLM grading prompts. Prompt, parameters, response schema, and LLM responses are automatically logged to traces.json.

from pydantic import BaseModel, Field

class Rubric(BaseModel):
    score: float = Field(description="Score between 0.0 and 1.0 based on correctness.")
    reasoning: str = Field(description="Detailed reasoning for the score.")

# In your criterion:
result = ev.llm_judge(
    model="google/gemini-3.5-flash",
    response_model=Rubric,
    system="You are a strict code correctness evaluator.",
    prompt="Compare the student's solution in code.py with the requirements...",
)

# The return value is fully type-hinted as an instance of your Rubric class
print(result.score)
print(result.reasoning)

3. File & Artifact Management

Access files and copy evaluation artifacts to the logs directory safely:

  • ev.read_file(filename): Reads content as a string and auto-saves a copy to artifacts.
  • ev.load_json(filename): Parses JSON file content and auto-saves a copy to artifacts.
  • ev.save_file(filename, content): Saves arbitrary text to the artifacts directory.
  • ev.save_dir(dirname): Copies an entire directory from the workspace to the artifacts directory.
  • ev.load_trajectory(path): Loads and parses an agent's ATIF trajectory.json file.

Outputs

When ev.run() completes, the following files are written to the directory containing your configured output_path:

  1. reward.json: Flat JSON dictionary containing the final calculated reward and individual scores.
  2. reward.txt: Text file containing just the final reward float value.
  3. traces.json: List of structured LLM calls made via ev.llm_judge.
  4. metadata.json: Optional metadata.
  5. artifacts/: Subfolder containing copy-back files preserved during the evaluation run.

Agent Skills

You can install the graded-verifier skill to teach your AI coding agents (such as Cursor or Claude Code) how to write robust graded verifiers:

npx skills add <github-username>/eval-helpers/.agents/skills/graded-verifier

Release files for graded 1.0.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for graded 1.0.5
File Size Uploaded
graded-1.0.5.tar.gz 12.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for graded 1.0.5
File Interpreter ABI Platform
graded-1.0.5-py3-none-any.whl Python 3 none any Details

Total release size: 21.1 kB

Release files / graded-1.0.5.tar.gz

Download URL graded-1.0.5.tar.gz
Size 12.4 kB
Tags Source
SHA-256 checksum
How to use checksums
3167a2ba7a20e1444ea0607763d941a60dc936c976706a256cc4c4333f044c0f
BLAKE2b-256 checksum
How to use checksums
8a6df15f2f38e12abcceeac2b5f01acb5ea344cc57f6b4c912a24f0881ba57ac
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 12, 2026.

Transparency log

Release files / graded-1.0.5-py3-none-any.whl

Download URL graded-1.0.5-py3-none-any.whl
Size 8.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d5567a90fd6e9e54f00d9fa766e474abb24c625b4fe23240d2c62595f840473c
BLAKE2b-256 checksum
How to use checksums
6176ee95790d86109a23a7bf081fc0cd0b8217cd2176c0a4b9977975847901a7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 12, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.5 This release

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page