Skip to main content

Production-grade LLM evaluation and benchmark harness

Project description

Evalkit-bench

Production-grade LLM evaluation and benchmark harness

Tests PyPI Python 3.11+ License: MIT

Point it at a YAML test suite → get a live terminal table, a self-contained HTML report, regression tracking, and side-by-side model comparisons.


What it does

$ evalkit run evals/factual.yaml --model gpt-4o-mini

  ✓  [1/5]  capital-france
  ✓  [2/5]  capital-japan
  ✓  [3/5]  cell-powerhouse
  ✗  [4/5]  speed-of-light
  ✓  [5/5]  pythagorean

╭──────────────────┬──────────────────────────┬──────────────┬──────────┬────────╮
│ ID               │ Prompt                   │ Output       │ Scorers  │ Result │
├──────────────────┼──────────────────────────┼──────────────┼──────────┼────────┤
│ capital-france   │ What is the capital of…  │ Paris        │ exact ✓  │  PASS  │
│ speed-of-light   │ What is the speed of…    │ 299,792 km/s │ exact ✗  │  FAIL  │
╰──────────────────┴──────────────────────────┴──────────────┴──────────┴────────╯

  ████████████████░░░░  80%  4/5 passed · avg 340ms · 1,200 tokens

  Diff vs previous: 1 regression
  HTML report → .evalkit/runs/20260505T120000_factual-v1_report.html

Install

pip install evalkit-bench

Add the embedding scorer (optional, ~1 GB download):

pip install "evalkit-bench[embed]"

Set your API keys — create a .env file in your project:

OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...

Quick start

Run a suite:

evalkit run evals/example_factual.yaml

Override model at runtime:

evalkit run evals/example_factual.yaml --model gpt-4o
evalkit run evals/example_factual.yaml --model claude-haiku-4-5-20251001

Compare multiple models side-by-side:

evalkit compare evals/example_factual.yaml \
  --models gpt-4o-mini,gpt-4o,claude-haiku-4-5-20251001
╭──────────────────┬─────────────┬─────────┬──────────────╮
│ Case             │ gpt-4o-mini │  gpt-4o │  haiku-4-5   │
├──────────────────┼─────────────┼─────────┼──────────────┤
│ capital-france   │    PASS     │  PASS   │     PASS     │
│ cell-powerhouse  │    PASS     │  PASS   │     FAIL     │
│ speed-of-light   │    PASS     │  PASS   │     PASS     │
├──────────────────┼─────────────┼─────────┼──────────────┤
│ Pass rate        │    100%     │  100%   │     80%      │
│ Avg latency      │    340ms    │  820ms  │    280ms     │
│ Total tokens     │   1,200     │  2,400  │     950      │
╰──────────────────┴─────────────┴─────────┴──────────────╯

Writing a benchmark suite

Suites are plain YAML files:

name: "factual-qa-v1"
description: "Tests basic factual recall"
model: "gpt-4o-mini"             # default model
judge_model: "claude-haiku-4-5-20251001"  # used by llm_judge scorer
temperature: 0.0
max_tokens: 512

cases:
  - id: capital-france
    prompt: "What is the capital of France? Reply with just the city name."
    expected: "Paris"
    scorers: [exact]
    tags: [geography, easy]

  - id: summarise-article
    prompt: "Summarise in 2 sentences: {input}"
    input: "The Eiffel Tower was built between 1887 and 1889..."
    expected: "The Eiffel Tower was built in the late 1880s."
    scorers: [llm_judge, embed]
    rubric: "Does the summary capture the key facts without hallucinating?"
    tags: [summarisation, medium]

All case fields

Field Required Description
id Unique identifier within the suite
prompt Prompt sent to the model. Supports {input} substitution
expected Reference answer (required by exact, contains, embed)
input Replaces {input} in the prompt
scorers List of scorers (default: [exact])
rubric Evaluation guidance for llm_judge
tags Arbitrary labels for filtering
pattern Regex pattern for exact scorer

Scorers

Scorer Passes when Best for
exact Output matches expected (case-insensitive). Exact → 1.0, substring → 0.5 Short deterministic answers
contains Expected string appears anywhere in output Keywords, function names
llm_judge Claude rates the output ≥ 3.5 / 5 Open-ended quality, summarisation
embed Cosine similarity of sentence embeddings ≥ 0.7 Semantic equivalence

Multiple scorers are AND-ed — a case passes only if every scorer passes.


Regression tracking

Every run is saved automatically to .evalkit/runs/. After each run, evalkit diffs against the previous run for the same suite:

Diff vs previous: 2 regressions, 1 improvement

Compare any two runs manually:

evalkit diff .evalkit/runs/20260505T090000_factual-v1.json \
             .evalkit/runs/20260505T100000_factual-v1.json

List all runs:

evalkit list-runs

Inspect a specific run:

evalkit show 20260505T120000

CLI reference

evalkit run     <suite.yaml>  [--model MODEL] [--output-dir DIR] [--no-report]
evalkit compare <suite.yaml>  --models MODEL1,MODEL2,...
evalkit diff    <run_a.json>  <run_b.json>
evalkit list-runs             [--runs-dir DIR]
evalkit show    <run_id>      [--runs-dir DIR]

CI-friendly: evalkit run exits with 0 if all cases pass, 1 if any fail.


HTML report

Every run generates a self-contained HTML report — no CDN, works offline, sendable as an attachment.

  • Pass/fail per case with color coding
  • Click any row to expand full prompt, output, and judge reasoning
  • Regression and improvement badges when diffing against a previous run
  • Overall pass-rate bar and summary stats

Roadmap

  • Gemini and Ollama (local) providers
  • Parallel case execution
  • Web UI dashboard for run history
  • Custom scorer plugins via entry points
  • evalkit init — interactive suite generator

Contributing

git clone https://github.com/Arman176001/evalkit
cd evalkit
pip install -e ".[dev]"
python -m pytest tests/ -v

All 44 tests run in under 3 seconds with no API calls required.


Made by Arman · MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

evalkit_bench-0.1.2.tar.gz (23.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

evalkit_bench-0.1.2-py3-none-any.whl (24.9 kB view details)

Uploaded Python 3

File details

Details for the file evalkit_bench-0.1.2.tar.gz.

File metadata

  • Download URL: evalkit_bench-0.1.2.tar.gz
  • Upload date:
  • Size: 23.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for evalkit_bench-0.1.2.tar.gz
Algorithm Hash digest
SHA256 b7c3348163a8e532be67527b76bf85169df5ff8ec420b951c018e807cb91dff9
MD5 bd4f66c13f730b86232e0e3cda302edb
BLAKE2b-256 5bdbc7bad851c5adcbd110b3c41323fd2ebb09b35fa72ddbac40c63e748f9053

See more details on using hashes here.

Provenance

The following attestation bundles were made for evalkit_bench-0.1.2.tar.gz:

Publisher: publish.yml on Arman176001/evalkit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file evalkit_bench-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: evalkit_bench-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 24.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for evalkit_bench-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 9143f0d555e3900032a94ba829f4b7845cefb0259cd33b6874dfa5a516d2ea57
MD5 566c9d6cb64ce22ebc26f9c247418bca
BLAKE2b-256 cef5f9a24b9753085918bd13209b4d50121f284843fd6a17706037d9fd1fe9a6

See more details on using hashes here.

Provenance

The following attestation bundles were made for evalkit_bench-0.1.2-py3-none-any.whl:

Publisher: publish.yml on Arman176001/evalkit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page