laya-evals
Judge cheap. Audit confidence.
Calibration-first evaluation for laya: score eval sets with a local System 1 decision model, verify whether its confidence is reliable, and protect CI from regressions.
Quick start · What you get · Measured evidence · Demo
Early-development community tool. This is not the
laya-evalsCLI distributed with the upstreamlayapackage. It is an independent, Apache-2.0 project built to make automated evaluation more honest.
The problem
An answer confidence is only safe to automate on if it has been calibrated against gold labels. That meaning can change with the question shape: in our measured runs, a 3-option XNLI question reached a 90% target at a 0.8074 gate, while a 20-option MASSIVE question needed 0.9944. One global min_confidence is not a policy.
laya-evals couples a cheap, batched rubric judge with the evidence needed to decide when to trust it.
What you get
| Capability | Why it matters |
|---|---|
| Batched rubric judge | Ask choice, ordered score, and boolean noul questions in one laya pass. |
| Calibration audit | Compute ECE, Brier score, reliability bins, and coverage/accuracy curves from (confidence, outcome) pairs. |
| Threshold advisor | Select the lowest gate that meets a target accuracy, separately for each question shape. |
| Judge comparison | Compare against gold or another judge with agreement, Cohen’s kappa, weighted kappa, accuracy, and cost-per-1k fields. |
| CI regression gate | Fail builds when accuracy drops or calibration worsens; deliberately exclude unstable wall-clock measurements. |
Scope and current status
- Implemented: calibration core, threshold advice, batched rubric judging, judge-comparison metrics, reproduction pack, and a reusable regression-gate Action.
- Measured, not guessed: the SST-2 laya run and the four reproduction claims below. A public reference-LLM comparison is intentionally still pending; its agreement and cost fields are
nulluntil a real run is recorded. - Use the right confidence: thresholds are shaped by option count. The audit is specifically designed to prevent a value that worked for one question type from silently governing another.
- Known upstream caveat: the laya checkpoint reports invalid temperature buckets for
choice:11+; the SST-2 measurement uses two options and is outside that bucket.
Quick start
Requires Python 3.10+ and uv.
git clone https://github.com/Gjusev/laya-evals.git
cd laya-evals
uv venv
uv pip install -e ".[dev]"
pytest
Start by turning known outcomes into calibration metrics:
from laya_evals import brier_score, coverage_accuracy_curve, ece, reliability_bins
confidences = [0.90, 0.90, 0.90, 0.20]
outcomes = [True, True, True, False] # prediction == gold
print(ece(confidences, outcomes))
print(brier_score(confidences, outcomes))
print(reliability_bins(confidences, outcomes))
print(coverage_accuracy_curve(confidences, outcomes))
For laya decisions, use answer_confidence—the model’s declared probability that the answer is correct. Judge accepts a laya Agent (or compatible object), runs a rubric over a batch, and returns judgments with that confidence attached.
from laya_evals import Judge, confidence_outcome_pairs, ece
judge = Judge(laya_agent, min_confidence=0.8)
questions = {
"relevance": {
"type": "choice",
"instructions": "Is the answer relevant to the question?",
"criteria": {
"relevant": "It addresses the question.",
"irrelevant": "It does not.",
},
},
}
judged = judge.judge_batch(model_outputs, questions)
confidences, outcomes = confidence_outcome_pairs(
[item["relevance"] for item in judged], gold_relevance
)
print(ece(confidences, outcomes))
Evidence
Independent reproduction pack
We re-ran four public laya benchmark claims using the upstream notebook’s protocol: seed 13, 20-option MASSIVE draws, the first 300 test rows per language, pinned checkpoints, and CPU inference on one machine. All four results were within 0.0003 of the published values.
| Claim | Published | Measured | Result |
|---|---|---|---|
| MASSIVE intent, English | 0.7830 | 0.7833 | Reproduces |
| MASSIVE intent, 13 other languages | 0.4510 | 0.4510 | Reproduces |
| XNLI, English | 0.8600 | 0.8600 | Reproduces |
| XNLI, 14 other languages | 0.7310 | 0.7307 | Reproduces |
The recorded reproduction artifact includes per-language detail, ECE, top-1 Brier score, protocol, and the ±0.05 verdict rule. Re-run it with:
uv run --extra compare python scripts/reproduction_pack.py --per-lang 300
SST-2 judge measurement
One measured CPU run on the balanced 872-item SST-2 development split, using laya 0.3.21 and convaiinnovations/laya with a two-option sentiment rubric:
| Metric | Measured |
|---|---|
| Accuracy against gold | 0.8968 |
| Cohen’s kappa | 0.7938 |
| ECE, 15 bins | 0.0253 |
| Brier score | 0.0779 |
| Throughput | 6.4 decisions/s |
The reference LLM judge has not been run, so its agreement and cost fields remain null rather than estimated. Full inputs and outputs are in results/sst2-comparison.json.
uv run --extra compare python scripts/sst2_judge_comparison.py
uv run --extra compare --extra dev pytest -m slow # real-checkpoint smoke test
Choose thresholds per shape
advise_thresholds finds the lowest confidence gate whose retained examples meet your target accuracy. If no gate can meet it, it returns the best available point and marks the advice as unachievable—never an invented threshold.
| Measured question shape | Target accuracy | Suggested gate | Coverage |
|---|---|---|---|
| SST-2 · 2 options | 0.95 | 0.8651 | 0.834 |
| XNLI English · 3 options | 0.90 | 0.8074 | 0.893 |
| MASSIVE intent English · 20 options | 0.90 | 0.9944 | 0.773 |
from laya_evals import advise_thresholds
advice = advise_thresholds(
{
"options=2": (confidences_2, outcomes_2),
"options=20": (confidences_20, outcomes_20),
},
target_accuracy=0.90,
)
CI regression gate
Compare a fresh measurement with a committed baseline. Accuracy-like metrics fail on drops; calibration metrics fail on rises. The command exits 0 for pass, 1 for regression, and 2 for invalid usage.
uv run python scripts/check_regression.py \
--current results/sst2-comparison.json \
--baseline results/baselines/sst2-comparison.json \
--metric accuracy_vs_gold:max \
--metric ece_15_bins:min
The reusable GitHub Action lives at .github/actions/regression-gate. It is designed for a manual benchmark run because checkpoints are large and CPU evaluation takes minutes.
Demo
Your browser does not support embedded video. Watch the 21-second demo.21 seconds: from a real benchmark reproduction to the threshold-calibration finding. If your README renderer does not support video, use the direct MP4 link or open the one-page visual overview.
Further reading
License
Apache-2.0. See LICENSE.
Metadata
Release files for laya-evals 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| laya_evals-0.1.0.tar.gz | 234.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| laya_evals-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 255.6 kB
Release files / laya_evals-0.1.0.tar.gz
| Download URL | laya_evals-0.1.0.tar.gz |
|---|---|
| Size | 234.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f32bacf6f643d0499e560bceb9c0f007fdc0af3c740c3aaae9e6b9c5f6e44cd2
|
|
BLAKE2b-256 checksum How to use checksums |
2c4ec3f69ab2fb2c7fe17ab07b4239871c6ebb611808c3f7dbd96364e0fd4876
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / laya_evals-0.1.0-py3-none-any.whl
| Download URL | laya_evals-0.1.0-py3-none-any.whl |
|---|---|
| Size | 21.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2eb15b66a56d2f5afbb8ac2527d4d29f5a86ab81d0d2d08307232df2d41e08db
|
|
BLAKE2b-256 checksum How to use checksums |
770fcab5449b5a7e2e41db07117d752254b6e4ebc2a91ea71601dc386a0f2622
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log