Skip to main content

laya-evals logo

laya-evals

Judge cheap. Audit confidence.

Calibration-first evaluation for laya: score eval sets with a local System 1 decision model, verify whether its confidence is reliable, and protect CI from regressions.

CI status Python 3.10 or later Apache 2.0 license

Quick start · What you get · Measured evidence · Demo

laya-evals social preview — calibration-first LLM evaluation for CI

Early-development community tool. This is not the laya-evals CLI distributed with the upstream laya package. It is an independent, Apache-2.0 project built to make automated evaluation more honest.

The problem

An answer confidence is only safe to automate on if it has been calibrated against gold labels. That meaning can change with the question shape: in our measured runs, a 3-option XNLI question reached a 90% target at a 0.8074 gate, while a 20-option MASSIVE question needed 0.9944. One global min_confidence is not a policy.

laya-evals couples a cheap, batched rubric judge with the evidence needed to decide when to trust it.

Eval set flows through laya judge, calibration audit, threshold advisor, and CI gate

What you get

Capability Why it matters
Batched rubric judge Ask choice, ordered score, and boolean noul questions in one laya pass.
Calibration audit Compute ECE, Brier score, reliability bins, and coverage/accuracy curves from (confidence, outcome) pairs.
Threshold advisor Select the lowest gate that meets a target accuracy, separately for each question shape.
Judge comparison Compare against gold or another judge with agreement, Cohen’s kappa, weighted kappa, accuracy, and cost-per-1k fields.
CI regression gate Fail builds when accuracy drops or calibration worsens; deliberately exclude unstable wall-clock measurements.

Scope and current status

  • Implemented: calibration core, threshold advice, batched rubric judging, judge-comparison metrics, reproduction pack, and a reusable regression-gate Action.
  • Measured, not guessed: the SST-2 laya run and the four reproduction claims below. A public reference-LLM comparison is intentionally still pending; its agreement and cost fields are null until a real run is recorded.
  • Use the right confidence: thresholds are shaped by option count. The audit is specifically designed to prevent a value that worked for one question type from silently governing another.
  • Known upstream caveat: the laya checkpoint reports invalid temperature buckets for choice:11+; the SST-2 measurement uses two options and is outside that bucket.

Quick start

Requires Python 3.10+ and uv.

git clone https://github.com/Gjusev/laya-evals.git
cd laya-evals
uv venv
uv pip install -e ".[dev]"
pytest

Start by turning known outcomes into calibration metrics:

from laya_evals import brier_score, coverage_accuracy_curve, ece, reliability_bins

confidences = [0.90, 0.90, 0.90, 0.20]
outcomes = [True, True, True, False]  # prediction == gold

print(ece(confidences, outcomes))
print(brier_score(confidences, outcomes))
print(reliability_bins(confidences, outcomes))
print(coverage_accuracy_curve(confidences, outcomes))

For laya decisions, use answer_confidence—the model’s declared probability that the answer is correct. Judge accepts a laya Agent (or compatible object), runs a rubric over a batch, and returns judgments with that confidence attached.

from laya_evals import Judge, confidence_outcome_pairs, ece

judge = Judge(laya_agent, min_confidence=0.8)
questions = {
    "relevance": {
        "type": "choice",
        "instructions": "Is the answer relevant to the question?",
        "criteria": {
            "relevant": "It addresses the question.",
            "irrelevant": "It does not.",
        },
    },
}

judged = judge.judge_batch(model_outputs, questions)
confidences, outcomes = confidence_outcome_pairs(
    [item["relevance"] for item in judged], gold_relevance
)
print(ece(confidences, outcomes))

Evidence

Independent reproduction pack

We re-ran four public laya benchmark claims using the upstream notebook’s protocol: seed 13, 20-option MASSIVE draws, the first 300 test rows per language, pinned checkpoints, and CPU inference on one machine. All four results were within 0.0003 of the published values.

Claim Published Measured Result
MASSIVE intent, English 0.7830 0.7833 Reproduces
MASSIVE intent, 13 other languages 0.4510 0.4510 Reproduces
XNLI, English 0.8600 0.8600 Reproduces
XNLI, 14 other languages 0.7310 0.7307 Reproduces

The recorded reproduction artifact includes per-language detail, ECE, top-1 Brier score, protocol, and the ±0.05 verdict rule. Re-run it with:

uv run --extra compare python scripts/reproduction_pack.py --per-lang 300

SST-2 judge measurement

One measured CPU run on the balanced 872-item SST-2 development split, using laya 0.3.21 and convaiinnovations/laya with a two-option sentiment rubric:

Metric Measured
Accuracy against gold 0.8968
Cohen’s kappa 0.7938
ECE, 15 bins 0.0253
Brier score 0.0779
Throughput 6.4 decisions/s

The reference LLM judge has not been run, so its agreement and cost fields remain null rather than estimated. Full inputs and outputs are in results/sst2-comparison.json.

uv run --extra compare python scripts/sst2_judge_comparison.py
uv run --extra compare --extra dev pytest -m slow  # real-checkpoint smoke test

Choose thresholds per shape

advise_thresholds finds the lowest confidence gate whose retained examples meet your target accuracy. If no gate can meet it, it returns the best available point and marks the advice as unachievable—never an invented threshold.

Measured question shape Target accuracy Suggested gate Coverage
SST-2 · 2 options 0.95 0.8651 0.834
XNLI English · 3 options 0.90 0.8074 0.893
MASSIVE intent English · 20 options 0.90 0.9944 0.773
from laya_evals import advise_thresholds

advice = advise_thresholds(
    {
        "options=2": (confidences_2, outcomes_2),
        "options=20": (confidences_20, outcomes_20),
    },
    target_accuracy=0.90,
)

CI regression gate

Compare a fresh measurement with a committed baseline. Accuracy-like metrics fail on drops; calibration metrics fail on rises. The command exits 0 for pass, 1 for regression, and 2 for invalid usage.

uv run python scripts/check_regression.py \
  --current results/sst2-comparison.json \
  --baseline results/baselines/sst2-comparison.json \
  --metric accuracy_vs_gold:max \
  --metric ece_15_bins:min

The reusable GitHub Action lives at .github/actions/regression-gate. It is designed for a manual benchmark run because checkpoints are large and CPU evaluation takes minutes.

Demo

Your browser does not support embedded video. Watch the 21-second demo.

21 seconds: from a real benchmark reproduction to the threshold-calibration finding. If your README renderer does not support video, use the direct MP4 link or open the one-page visual overview.

Further reading

License

Apache-2.0. See LICENSE.

Metadata

Release files for laya-evals 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for laya-evals 0.1.0
File Size Uploaded
laya_evals-0.1.0.tar.gz 234.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for laya-evals 0.1.0
File Interpreter ABI Platform
laya_evals-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 255.6 kB

Release files / laya_evals-0.1.0.tar.gz

Download URL laya_evals-0.1.0.tar.gz
Size 234.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f32bacf6f643d0499e560bceb9c0f007fdc0af3c740c3aaae9e6b9c5f6e44cd2
BLAKE2b-256 checksum
How to use checksums
2c4ec3f69ab2fb2c7fe17ab07b4239871c6ebb611808c3f7dbd96364e0fd4876
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / laya_evals-0.1.0-py3-none-any.whl

Download URL laya_evals-0.1.0-py3-none-any.whl
Size 21.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2eb15b66a56d2f5afbb8ac2527d4d29f5a86ab81d0d2d08307232df2d41e08db
BLAKE2b-256 checksum
How to use checksums
770fcab5449b5a7e2e41db07117d752254b6e4ebc2a91ea71601dc386a0f2622
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page