kitty-evals
Evaluator toolkit for LLM outputs: an LLM-as-a-judge with pluggable providers, structured verdicts, and concurrent batches.
Quickstart
uvx kitty_evals init
Opens a blue multi-select of the available evaluators, asks where to put them
(default evaluation/), and copies the packages in:
evaluation/
└── llm_judge/
├── __init__.py
├── base.py
├── config.py
├── ...
Non-interactive / scripted:
uvx kitty_evals init --list # show the catalog
uvx kitty_evals init llm_judge # copy one evaluator
uvx kitty_evals init --yes # every ready evaluator
uvx kitty_evals init -d path/to/dir # explicit destination
Use a judge
from kitty_evals import Judge
judge = Judge("gpt-4o", temperature=0.2, score_range=(1, 5))
verdict = judge.evaluate("The model said ...", context="task prompt")
print(verdict.score, verdict.rationale)
# Batch: a JSON list (or any iterable of dicts) — results in input order
verdicts = judge.evaluate_many("cases.json", max_concurrency=20)
Each item needs a prompt; optional context, reference, plus any extra keys,
which are passed to the judge as a JSON variable block.
Install for development
uv sync
uv run pytest
Metadata
Release files for kitty-evals 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| kitty_evals-0.1.0.tar.gz | 10.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| kitty_evals-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.1 kB
Release files / kitty_evals-0.1.0.tar.gz
| Download URL | kitty_evals-0.1.0.tar.gz |
|---|---|
| Size | 10.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4fcb4baaf9544e1313435f3dd0f7cd9f3cf230298c9612593fca1d36d7f955b9
|
|
BLAKE2b-256 checksum How to use checksums |
171be57dfd56d93737921260d69df02bff0f6ca148216564ed2ab702bd3fabf7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / kitty_evals-0.1.0-py3-none-any.whl
| Download URL | kitty_evals-0.1.0-py3-none-any.whl |
|---|---|
| Size | 14.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7ad675c8e84c1a061a1c729919d34fec79cc1c45deb4d040b430287b94b0be42
|
|
BLAKE2b-256 checksum How to use checksums |
e334a2c5a822e1c6ae2c799b15fc2b9c3a919aee978e23eb53425880a753b0bc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log