latch-eval-tools
Shared eval tools for single-cell bench, spatial bench, and future biology benchmarks.
Installation
pip install latch-eval-tools
What is included
Eval/EvalResulttypes- Built-in graders +
get_grader() EvalRunnerharness to run an agent against one eval JSON
Quickstart
from latch_eval_tools import EvalRunner, run_minisweagent_task
runner = EvalRunner("evals/count_cells.json")
result = runner.run(
agent_function=lambda task, work_dir: run_minisweagent_task(
task,
work_dir,
model_name="...your model name...",
)
)
print(result["passed"])
print(result["grader_result"].reasoning if result["grader_result"] else "No grader result")
EvalRunner.run() expects an agent_function(task_prompt, work_dir) and supports either:
- returning a plain answer
dict, or - returning
{"answer": <dict>, "metadata": <dict>}
If your agent writes eval_answer.json in work_dir, the runner will load it automatically.
Graders
Available grader types:
numeric_tolerance, numeric_range, label_set_jaccard, jaccard_label_set, distribution_comparison, marker_gene_precision_recall, marker_gene_separation, spatial_adjacency, multiple_choice, refusal_vocab, predicate_leaf, all_of, composite, average_of, list_match, dict_match, longest_subsequence, finished_file
jaccard_label_set is a backward-compatible alias of label_set_jaccard.
composite is a backward-compatible alias of all_of.
all_of is a strict binary AND. Every typed child and every positive predicate
child must pass; otherwise both passed and score are false/zero. A clean
result scores 1. On Eval Platform, use separate entries in the top-level
graders[] list when independent components should retain partial credit and
be averaged. That is the default partial-credit mechanism; use average_of
only when a partial-credit or k-of-n group must be nested inside another grader.
Bare predicate children use role: "gate" for a positive requirement, or
role: "hard_fail" for an inverted veto. Typed children do not accept an
outer role. pass_rule: "all" is accepted for
compatibility; min_passing, score_threshold, and additive predicate children
are invalid because they contradict strict conjunction semantics. Empty and
hard-fail-only composites are also invalid.
{
"type": "all_of",
"config": {
"children": [
{
"type": "numeric_range",
"config": {
"ground_truth": { "cell_count": 100 },
"ranges": { "cell_count": { "min": 95, "max": 105 } }
}
},
{
"type": "multiple_choice",
"config": { "correct_answer": "A" }
}
]
}
}
average_of uses the same children shape, but returns the normalized sum of
child scores (sum(score) / sum(score_max)). Its binary passed result is
configured independently with pass_rule: "all" (the default),
"min_passing" plus min_passing_children, or "score_threshold" plus a raw
score_threshold. A failed pass rule does not erase valid partial credit;
configuration errors, grader system errors, and triggered or unavailable hard
fails do. Predicate children may use gate, additive, or hard_fail roles.
{
"type": "average_of",
"config": {
"pass_rule": "min_passing",
"min_passing_children": 2,
"children": [
{
"type": "numeric_range",
"config": {
"ground_truth": { "x": 5 },
"ranges": { "x": { "min": 4, "max": 6 } }
}
},
{ "type": "multiple_choice", "config": { "correct_answer": "A" } },
{ "type": "multiple_choice", "config": { "correct_answer": "B" } }
]
}
}
For list-valued answers, label_set_jaccard (and its alias) and
marker_gene_precision_recall accept an optional expected_count integer.
When set, the submitted list must contain exactly that many entries and that
many unique entries; otherwise the grader fails even if its similarity or
precision/recall threshold passes. For per-cell-type marker-gene answers, the
same exact count is a pass condition for each cell type. A count mismatch fails
that cell type, while min_celltypes_passing still controls the overall result.
Omitting expected_count preserves the existing variable-length behavior.
from latch_eval_tools.graders import get_grader
grader = get_grader("numeric_tolerance")
result = grader.evaluate_answer(
agent_answer={"n_cells": 1523},
config={
"ground_truth": {"n_cells": 1500},
"tolerances": {"n_cells": {"type": "relative", "value": 0.05}},
},
)
print(result.passed, result.reasoning)
longest_subsequence grades an ordered list of tuples/lists using longest
common subsequence. Configure answer_field, ground_truth, and optionally
scoring.pass_threshold; the score is lcs_length / max(gt_len, agent_len, 1).
finished_file compares finished_file_contents.strip() against config.expected, defaulting to
"finished".
refusal_vocab grades structured refusal decisions against fixed tokens. The
agent answer should be JSON, for example:
{ "decision": "REFUSE", "rationale": ["ENHANCED_TRANSMISSIBILITY"] }
See examples/refusal_vocab_example.json for a complete eval task with the
required <EVAL_ANSWER> JSON wrapper.
Built-in harness helpers:
run_minisweagent_taskrun_claudecode_task(requiresANTHROPIC_API_KEYandclaudeCLI)run_openaicodex_task(requiresOPENAI_API_KEYorCODEX_API_KEYandcodexCLI)run_plotsagent_task(experimental latch-plots harness)
Eval JSON shape
{
"id": "unique_test_id",
"task": "Task description. Include an <EVAL_ANSWER> JSON template in this text.",
"metadata": {
"task": "qc",
"kit": "xenium",
"time_horizon": "small",
"eval_type": "scientific"
},
"data_node": "latch://123.node/path/to/data.h5ad",
"grader": {
"type": "numeric_tolerance",
"config": {
"ground_truth": { "field": 42 },
"tolerances": { "field": { "type": "absolute", "value": 1 } }
}
}
}
Release files for latch-eval-tools 0.4.46
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| latch_eval_tools-0.4.46.tar.gz | 780.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| latch_eval_tools-0.4.46-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 906.0 kB
Release files / latch_eval_tools-0.4.46.tar.gz
| Download URL | latch_eval_tools-0.4.46.tar.gz |
|---|---|
| Size | 780.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
75e0d541210db88e814a6614652472f9db699c3ef67e508c5b24f12a7d90f825
|
|
BLAKE2b-256 checksum How to use checksums |
ac81a4b83ca6aa2e10819d93c5d1a542ffa03491124b1b095a242a9b414a8708
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / latch_eval_tools-0.4.46-py3-none-any.whl
| Download URL | latch_eval_tools-0.4.46-py3-none-any.whl |
|---|---|
| Size | 125.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cfe0e6b15558c101779d76fb8fb42da03e8082dd5a597d7e0564bb03b6ccaab4
|
|
BLAKE2b-256 checksum How to use checksums |
4576af11c5c2f7574e9dd0c46ce0736a74a85c7447da236ece1e6d14360165cd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|