Eval Banana
Aspect-based evaluation framework - deterministic checks + harness judges. Score anything (agentic outputs, workflows, banana!) with simple YAML check definitions.
The name was inspired by this song (my kids love it)
What it does
Eval Banana discovers YAML check definitions from eval_checks/ directories, runs them, and produces a report. Every check scores 0 or 1 with equal weight.
Two check types:
| Type | Purpose | How it works |
|---|---|---|
deterministic |
Objective assertions (file existence, content, structure) | Runs a Python script via subprocess; exit 0 = pass |
harness_judge |
LLM-as-a-judge (coherence, accuracy, tone) | Invokes the configured AI agent; requires exit 0 and a {"score": 0|1} verdict |
The harness judge uses one of the following: codex, gemini, claude, openhands, opencode, pi
Writing checks
Create a directory called eval_checks/ anywhere in your project. Add YAML files -- one per check.
Deterministic check
schema_version: 1
id: output_file_exists
type: deterministic
description: Verify that output.json was generated.
script: |
import json, sys
from pathlib import Path
ctx = json.loads(Path(sys.argv[1]).read_text())
output = Path(ctx["project_root"]) / "output.json"
assert output.exists(), "output.json not found"
Harness judge check
schema_version: 1
id: summary_is_accurate
type: harness_judge
description: The generated summary accurately reflects source data.
instructions: |
Read summary.txt and source_data.json.
Compare the summary against the source data.
Score 1 if accurate, 0 if it contains fabricated claims.
Requires a configured harness agent. Set [harness] agent in config or pass --harness-agent.
Inspiration
Eval Banana's binary 0/1 scoring philosophy draws directly on two earlier bodies of work:
- Hamel Husain's Creating LLM-as-a-Judge that drives business results — argues that binary pass/fail judgments produce more reliable, actionable evals than Likert-style 1-5 scales.
- RAGAS's Aspect Critic metric — evaluates outputs against a natural-language aspect definition and returns a binary verdict.
The harness_judge check type is essentially an Aspect Critic: you describe what "good" looks like in plain language, and the judge returns {"score": 0|1}.
Skills
eval-banana ships agent skills in the skills/ directory of the repository, including:
eval-banana: guidance for writing and debugging eval-banana checksgemini_media_use: Gemini media upload and analysis helpersvoxtral-transcribe: Mistral Voxtral transcription patterns for generated audio evals
Install repo-local skills into your project with the npx skills CLI:
npx skills add https://github.com/writeitai/eval-banana
The CLI auto-detects installed agents and copies skills into their native directories (.claude/skills/, .codex/skills/, .agents/skills/, .gemini/skills/, etc.).
Quick start
# Install
uv sync
# Initialize project config
eb init
# Run all discovered checks
eb run
# List discovered checks without running
eb list
# Validate YAML definitions without running
eb validate
Installation
# Using uv (recommended)
uv add eval-banana
# Using pip
pip install eval-banana
# From source (development)
git clone https://github.com/writeitai/eval-banana.git
cd eval-banana
uv sync --extra dev
After installation the CLI is available as eb.
Harness configuration
harness_judge checks require a configured harness agent. Configure it via TOML or CLI flags.
TOML
# .eval-banana/config.toml
[harness]
agent = "codex"
# codex GPT-5.6 tiers: gpt-5.6-sol (flagship), gpt-5.6-terra, gpt-5.6-luna
model = "gpt-5.6-sol"
# reasoning_effort = "high"
# timeout_seconds = 300
Running in CI / cloud
The harness subprocess inherits the parent shell environment, so provide API keys the same way you would when running the agent locally:
| Agent | Environment variable |
|---|---|
claude |
ANTHROPIC_API_KEY |
codex |
OPENAI_API_KEY |
gemini |
GEMINI_API_KEY or GOOGLE_API_KEY (or Application Default Credentials) |
openhands |
depends on the configured LLM backend |
Example GitHub Actions step:
- name: Run evals
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: eb run
You can also inject extra env vars via [harness.env] in your config:
[harness.env]
MY_CUSTOM_VAR = "value"
Custom agent templates
Add [agents.<name>] sections to override built-in templates or define new ones:
[agents.myagent]
command = ["my-cli", "run"]
shared_flags = ["--headless"]
prompt_flag = "--prompt"
model_flag = "--model"
Configuration
Eval Banana uses a single project-level TOML config at .eval-banana/config.toml.
Create it with eb init.
Config precedence (highest to lowest)
- CLI arguments (
--output-dir,--harness-model, etc.) - Environment variables (
EVAL_BANANA_*) - Project config (
.eval-banana/config.toml) - Built-in defaults
Key settings
| Setting | Default | Env var |
|---|---|---|
output_dir |
.eval-banana/results |
EVAL_BANANA_OUTPUT_DIR |
pass_threshold |
1.0 |
EVAL_BANANA_PASS_THRESHOLD |
llm_max_input_chars |
0 |
EVAL_BANANA_LLM_MAX_INPUT_CHARS |
harness.agent |
unset | EVAL_BANANA_HARNESS_AGENT |
harness.model |
unset | EVAL_BANANA_HARNESS_MODEL |
CLI reference
eb init [--force] Create project config
eb run [OPTIONS] Run all discovered checks
eb list [OPTIONS] List discovered checks
eb validate [OPTIONS] Validate YAML without running
Options for run/list/validate:
--check-dir PATH Scan only this directory
--check-id TEXT Run only this check ID
--output-dir TEXT Override output directory
--flat-output Write directly to an explicit empty output dir
--pass-threshold FLOAT Minimum pass ratio (0.0-1.0)
--no-project-config Ignore .eval-banana/config.toml (run/validate)
--verbose Enable debug logging
--cwd TEXT Working directory
Harness options:
--harness-agent TEXT Agent used by harness_judge checks (run/validate)
--harness-model TEXT Model override for the agent
--harness-reasoning-effort TEXT Reasoning effort level
--harness-timeout-seconds INTEGER Per-agent wall-clock limit (default: 300)
Output
Each run creates a timestamped directory under the configured output_dir:
.eval-banana/results/<run_id>/
report.json # Machine-readable full report
report.md # Human-readable Markdown report
checks/
<safe_check_id_stem>.json # Per-check result
<safe_check_id_stem>.prompt.txt # Exact harness-judge input (harness only)
<safe_check_id_stem>.stdout.txt # Captured stdout (if any)
<safe_check_id_stem>.stderr.txt # Captured stderr (if any)
An orchestration system that already owns an attempt-unique directory can use an exact, non-nested layout:
eval-banana run \
--flat-output \
--output-dir /absolute/path/to/attempt/eval
--flat-output requires an explicit --output-dir. The exact path must be
absent or an empty, non-symlink directory; eval-banana refuses conflicting
contents instead of overwriting them. Without this flag, the timestamped
<output_dir>/<run_id>/ layout is unchanged.
For every harness process that is invoked, eval-banana writes the exact prompt
argument first as checks/<safe_check_id_stem>.prompt.txt. The prompt and its
result JSON use the same deterministic filesystem-safe stem: unsafe characters
become underscores and leading/trailing dots or underscores are trimmed. The
prompt remains available when the agent fails, times out, or returns an invalid
verdict, so an attempt-level trace naturally inventories both the judge input
and its outputs.
report.json declares "schema_version": 1; readers can therefore reject a
future incompatible report shape explicitly. Every check result in that report
and in checks/<safe_check_id_stem>.json includes check_definition_sha256,
formatted as sha256:<64 lowercase hex>. The digest binds every byte that defines
the check:
the exact YAML bytes and, for script_path deterministic checks, the referenced
script bytes frozen before execution. The runner executes that same frozen script
snapshot, so a concurrent file change cannot produce a verdict under the older
digest. Its private launcher preserves the original resolved script path as
__file__ and sys.argv[0], plus its parent as the first import path. Inline
scripts are already part of the YAML component.
The canonical digest input is versioned and length-framed. It starts with the
ASCII domain eval-banana/check-definition-sha256/v1 followed by a NUL byte.
Each component is u64be(name_length) || name || u64be(content_length) || content.
The first component is always definition.yaml; a readable referenced script adds
referenced-script. A missing or unreadable referenced script adds an empty
referenced-script-unavailable component and produces an error result. The
reported value is SHA-256 over the complete framed input.
Orchestrators must not substitute a raw YAML file hash for this canonical digest. Use the producer-owned API to validate a report without duplicating the framing algorithm:
from pathlib import Path
from eval_banana.runner import compute_check_definition_sha256
definition_sha256 = compute_check_definition_sha256(
source_path=Path("eval_checks/goal.yaml")
)
The function reads one exact YAML snapshot and applies the same referenced
script snapshot rules as the runner. It raises OSError when the YAML cannot
be read and ValueError when the bytes are not a valid check definition.
Harness-judge details also record the resolved agent_type, model, nullable
reasoning_effort, and positive timeout_seconds. Configure the per-agent
wall-clock limit with [harness] timeout_seconds,
EVAL_BANANA_HARNESS_TIMEOUT_SECONDS, or --harness-timeout-seconds; it
defaults to 300 seconds.
A judge result is accepted only when its process exits 0; a non-zero exit is
an error with score 0, even when stdout contains valid passing JSON.
If the selected template has no model flag/environment variable or no
reasoning-effort flag containing {effort}, an attempted override becomes an
error before launch and the corresponding effective detail remains null.
Development
uv sync --extra dev
make test # Run tests
make fix # Auto-fix lint + format
make pyright # Type check
make all-check # Lint + format + types + tests (matches CI)
Contributing
Issues and pull requests are welcome. Please run make all-check before opening a PR.
Changelog
See CHANGELOG.md for release notes.
License
Apache License 2.0 — see LICENSE for details.
Copyright 2026 WriteIt.ai s.r.o.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file eval_banana-0.3.3.tar.gz.
File metadata
- Download URL: eval_banana-0.3.3.tar.gz
- Upload date:
- Size: 2.8 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c73c4d88c735933b9242177c82dc545b1b304bfe38186742cc672ff3225ac92a
|
|
| MD5 |
38dee818ce2823a6c9fdc382aec8004d
|
|
| BLAKE2b-256 |
8315cb607403423333f77bce5e6e144bd57697afa3434ea09bad8f74c01dd13c
|
Provenance
The following attestation bundles were made for eval_banana-0.3.3.tar.gz:
Publisher:
release.yml on writeitai/eval-banana
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
eval_banana-0.3.3.tar.gz -
Subject digest:
c73c4d88c735933b9242177c82dc545b1b304bfe38186742cc672ff3225ac92a - Sigstore transparency entry: 2188168184
- Sigstore integration time:
-
Permalink:
writeitai/eval-banana@18233c7369f79b3d4ba5862f92ab13f15025f5e2 -
Branch / Tag:
refs/tags/v0.3.3 - Owner: https://github.com/writeitai
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@18233c7369f79b3d4ba5862f92ab13f15025f5e2 -
Trigger Event:
push
-
Statement type:
File details
Details for the file eval_banana-0.3.3-py3-none-any.whl.
File metadata
- Download URL: eval_banana-0.3.3-py3-none-any.whl
- Upload date:
- Size: 38.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f6a8dd6b53cd8ba69cc86f5df430f40733e2c58b3996aba87d0bfe1156812f95
|
|
| MD5 |
a79ad5a05dc739f38c22f240b90a7de6
|
|
| BLAKE2b-256 |
73c7354db207647cca16f366772c6ff6ef56065e205bcd50d6ca0faf9c3b2a07
|
Provenance
The following attestation bundles were made for eval_banana-0.3.3-py3-none-any.whl:
Publisher:
release.yml on writeitai/eval-banana
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
eval_banana-0.3.3-py3-none-any.whl -
Subject digest:
f6a8dd6b53cd8ba69cc86f5df430f40733e2c58b3996aba87d0bfe1156812f95 - Sigstore transparency entry: 2188168205
- Sigstore integration time:
-
Permalink:
writeitai/eval-banana@18233c7369f79b3d4ba5862f92ab13f15025f5e2 -
Branch / Tag:
refs/tags/v0.3.3 - Owner: https://github.com/writeitai
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@18233c7369f79b3d4ba5862f92ab13f15025f5e2 -
Trigger Event:
push
-
Statement type: