Skip to main content

Production-grade LLM evaluation and benchmark harness

Project description

Evalkit-bench

Production-grade LLM evaluation and benchmark harness

Tests PyPI Python 3.11+ License: MIT

Point it at a YAML test suite → get a live terminal table, a self-contained HTML report, regression tracking, and side-by-side model comparisons.


What it does

$ evalkit run evals/factual.yaml --model gpt-4o-mini

  ✓  [1/5]  capital-france
  ✓  [2/5]  capital-japan
  ✓  [3/5]  cell-powerhouse
  ✗  [4/5]  speed-of-light
  ✓  [5/5]  pythagorean

╭──────────────────┬──────────────────────────┬──────────────┬──────────┬────────╮
│ ID               │ Prompt                   │ Output       │ Scorers  │ Result │
├──────────────────┼──────────────────────────┼──────────────┼──────────┼────────┤
│ capital-france   │ What is the capital of…  │ Paris        │ exact ✓  │  PASS  │
│ speed-of-light   │ What is the speed of…    │ 299,792 km/s │ exact ✗  │  FAIL  │
╰──────────────────┴──────────────────────────┴──────────────┴──────────┴────────╯

  ████████████████░░░░  80%  4/5 passed · avg 340ms · 1,200 tokens

  Diff vs previous: 1 regression
  HTML report → .evalkit/runs/20260505T120000_factual-v1_report.html

Install

pip install evalkit-bench

Add the embedding scorer (optional, ~1 GB download):

pip install "evalkit-bench[embed]"

Set your API keys — create a .env file in your project:

OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...

Quick start

Run a suite:

evalkit run evals/example_factual.yaml

Override model at runtime:

evalkit run evals/example_factual.yaml --model gpt-4o
evalkit run evals/example_factual.yaml --model claude-haiku-4-5-20251001

Compare multiple models side-by-side:

evalkit compare evals/example_factual.yaml \
  --models gpt-4o-mini,gpt-4o,claude-haiku-4-5-20251001
╭──────────────────┬─────────────┬─────────┬──────────────╮
│ Case             │ gpt-4o-mini │  gpt-4o │  haiku-4-5   │
├──────────────────┼─────────────┼─────────┼──────────────┤
│ capital-france   │    PASS     │  PASS   │     PASS     │
│ cell-powerhouse  │    PASS     │  PASS   │     FAIL     │
│ speed-of-light   │    PASS     │  PASS   │     PASS     │
├──────────────────┼─────────────┼─────────┼──────────────┤
│ Pass rate        │    100%     │  100%   │     80%      │
│ Avg latency      │    340ms    │  820ms  │    280ms     │
│ Total tokens     │   1,200     │  2,400  │     950      │
╰──────────────────┴─────────────┴─────────┴──────────────╯

Writing a benchmark suite

Suites are plain YAML files:

name: "factual-qa-v1"
description: "Tests basic factual recall"
model: "gpt-4o-mini"             # default model
judge_model: "claude-haiku-4-5-20251001"  # used by llm_judge scorer
temperature: 0.0
max_tokens: 512

cases:
  - id: capital-france
    prompt: "What is the capital of France? Reply with just the city name."
    expected: "Paris"
    scorers: [exact]
    tags: [geography, easy]

  - id: summarise-article
    prompt: "Summarise in 2 sentences: {input}"
    input: "The Eiffel Tower was built between 1887 and 1889..."
    expected: "The Eiffel Tower was built in the late 1880s."
    scorers: [llm_judge, embed]
    rubric: "Does the summary capture the key facts without hallucinating?"
    tags: [summarisation, medium]

All case fields

Field Required Description
id Unique identifier within the suite
prompt Prompt sent to the model. Supports {input} substitution
expected Reference answer (required by exact, contains, embed)
input Replaces {input} in the prompt
scorers List of scorers (default: [exact])
rubric Evaluation guidance for llm_judge
tags Arbitrary labels for filtering
pattern Regex pattern for exact scorer

Scorers

Scorer Passes when Best for
exact Output matches expected (case-insensitive). Exact → 1.0, substring → 0.5 Short deterministic answers
contains Expected string appears anywhere in output Keywords, function names
llm_judge Claude rates the output ≥ 3.5 / 5 Open-ended quality, summarisation
embed Cosine similarity of sentence embeddings ≥ 0.7 Semantic equivalence

Multiple scorers are AND-ed — a case passes only if every scorer passes.


Regression tracking

Every run is saved automatically to .evalkit/runs/. After each run, evalkit diffs against the previous run for the same suite:

Diff vs previous: 2 regressions, 1 improvement

Compare any two runs manually:

evalkit diff .evalkit/runs/20260505T090000_factual-v1.json \
             .evalkit/runs/20260505T100000_factual-v1.json

List all runs:

evalkit list-runs

Inspect a specific run:

evalkit show 20260505T120000

CLI reference

evalkit run     <suite.yaml>  [--model MODEL] [--output-dir DIR] [--no-report]
evalkit compare <suite.yaml>  --models MODEL1,MODEL2,...
evalkit diff    <run_a.json>  <run_b.json>
evalkit list-runs             [--runs-dir DIR]
evalkit show    <run_id>      [--runs-dir DIR]

CI-friendly: evalkit run exits with 0 if all cases pass, 1 if any fail.


HTML report

Every run generates a self-contained HTML report — no CDN, works offline, sendable as an attachment.

  • Pass/fail per case with color coding
  • Click any row to expand full prompt, output, and judge reasoning
  • Regression and improvement badges when diffing against a previous run
  • Overall pass-rate bar and summary stats

Roadmap

  • Gemini and Ollama (local) providers
  • Parallel case execution
  • Web UI dashboard for run history
  • Custom scorer plugins via entry points
  • evalkit init — interactive suite generator

Contributing

git clone https://github.com/Arman176001/evalkit
cd evalkit
pip install -e ".[dev]"
python -m pytest tests/ -v

All 44 tests run in under 3 seconds with no API calls required.


Made by Arman · MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

evalkit_bench-0.1.1.tar.gz (23.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

evalkit_bench-0.1.1-py3-none-any.whl (24.9 kB view details)

Uploaded Python 3

File details

Details for the file evalkit_bench-0.1.1.tar.gz.

File metadata

  • Download URL: evalkit_bench-0.1.1.tar.gz
  • Upload date:
  • Size: 23.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for evalkit_bench-0.1.1.tar.gz
Algorithm Hash digest
SHA256 64a72cb13f4583aab1467ce3f221a37ed53e586e19dc60cb519c2a8ab7bc46d9
MD5 57888093cfd82e95837d4a3e8aef69b4
BLAKE2b-256 528c121bff26a234cc22101366b394f016b0104c3070de62611936fe813b6a92

See more details on using hashes here.

Provenance

The following attestation bundles were made for evalkit_bench-0.1.1.tar.gz:

Publisher: publish.yml on Arman176001/evalkit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file evalkit_bench-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: evalkit_bench-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 24.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for evalkit_bench-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 17f3008dc2a95224b82ae8898bf25aa2cddf9b05c79a19922508b110a8816524
MD5 4aba65194dbc98842218bc9c1697d7b6
BLAKE2b-256 f6020cbee183463345305fd59e5bdc1a0645eb9535142a217ce9981fbcc425fb

See more details on using hashes here.

Provenance

The following attestation bundles were made for evalkit_bench-0.1.1-py3-none-any.whl:

Publisher: publish.yml on Arman176001/evalkit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page