Evaluate any model on any dataset and any task.
One library for benchmark, judge, code, red-team, security, and performance
evaluation — zero required deps.
Documentation (internal) · Quickstart · Examples · Contributing
TL;DR
| What | A unified Python library that evaluates any AI model on any dataset across any technique. |
| Architecture | 5-stage spine: Dataset → Adapter → Model → Metrics → RunResult. ak.evaluate() orchestrates it all. |
| Techniques | Benchmark, LLM-as-judge (GEval), code checks, RAG, hallucination, embedding similarity, toxicity/bias, pairwise/preference, red-teaming, security, performance |
| Models | 9 backends: OpenAI, Anthropic, HuggingFace, Lexsi, vLLM, LiteLLM, API, Groq, OpenRouter — all resolved via AutoModel.resolve(). Any list[str] → list[str] callable also works. |
| Zero deps | Core runs on stdlib. Backends and heavy metrics are optional extras (pip install auditkit[openai]). |
| Fingerprints | Every run gets a stable sha256 — results are cacheable, comparable, and reproducible by construction. |
| Status | 750+ tests, v1.0.0, LSAL-1.2 license (source-available, noncommercial). 10 metric families, 5 CLI subcommands, YAML config, MKDocs site. |
Why AuditKIT
Evaluating AI models is fragmented. Academic benchmarks (MMLU, GSM8K) use one tool. LLM-as-judge evaluations use another. Red-teaming and performance profiling each have their own frameworks. There is no single library that does all of them with a consistent API, zero required dependencies, and a provenance-first data model.
AuditKIT is that library. It provides a unified evaluation spine that supports every technique, so you can compare results across benchmarks, judge evaluations, red-team probes, and performance profiles — all from a single ak.evaluate() call.
Key features
- Metric families. Benchmark (exact match, F1, BLEU, ROUGE, ChrF), LLM-as-judge (GEval, rubric items), code/deterministic (contains, regex, JSON validation), RAG (lexical groundedness, context overlap/coverage), embedding similarity, hallucination detection, toxicity/bias, pairwise/preference (win rate, Elo, preference accuracy), security (DEFCON grade), performance (latency, throughput).
- Zero required deps. Core runs on the Python standard library alone. Heavy backends (BERTScore, vLLM, LiteLLM, OpenAI, Anthropic, HuggingFace) are optional extras.
- Any model backend.
echofor testing,openai:,anthropic:,hf:,lexsi:,vllm:,litellm:(which also reaches Ollama, e.g.litellm:ollama/llama3.1),api:,groq:,openrouter:, or any callable. Auto-resolved viaAutoModel.resolve(). - Many datasets, one model.
evaluate_many()runs one model across several datasets in a single call, returning oneRunResultper dataset. - Red teaming. Built-in adversarial probes (prompt injection, jailbreak, encoding, over-refusal) and detectors (keyword, refusal, injection success, system prompt leak) via
RedTeamRunner. - Experiment tracking. Named experiments with
ExperimentDB, MLflow logging, cross-run comparison with bootstrap significance tests. - Model comparison.
compare_models()runs the same dataset against multiple models and produces side-by-side results with pairwise significance. - CLI with YAML config.
auditkit eval,init,list,redteam,comparesubcommands. Define evaluations in YAML files withprompts:, model config, tags, and split strategies. - Provenance by default. Every run writes a
RunResultwith a stable fingerprint (sha256 over model+seed+tasks+config). Results are cacheable and comparable by fingerprint.
How it works
flowchart LR
A["Dataset<br/>Samples · Scenario · CSV"] --> B["Adapter<br/>generation · chat · instruction"]
B --> C["Model<br/>echo · openai · hf · vllm"]
C --> D{"Metrics"}
D -->|per sample| E["Score + Prediction"]
E --> F["Aggregate<br/>Stats · Headline"]
F --> G["RunResult<br/>fingerprint · cache"]
Install
pip install auditkit # core (zero deps)
pip install "auditkit[openai]" # OpenAI backend
pip install "auditkit[all]" # all backends + heavy metrics
Requires Python 3.10+ on Linux, macOS, or Windows.
More install options
pip install "auditkit[anthropic]" # Anthropic backend
pip install "auditkit[litellm]" # LiteLLM (Ollama, etc.)
pip install "auditkit[vllm]" # vLLM backend
pip install "auditkit[transformers]" # HuggingFace + hallucination + embedding similarity
pip install "auditkit[bert-score]" # BERTScore
pip install "auditkit[mlflow]" # MLflow experiment tracking
# From source
pip install "auditkit @ git+https://github.com/Lexsi-Labs/AuditKIT.git"
Quickstart
import auditkit as ak
samples = [
ak.Sample(input="What is 2+2?", target="4"),
ak.Sample(input="What is 3+3?", target="6"),
]
result = ak.evaluate(samples, model=lambda prompts: prompts)
print(result.summary())
From the CLI:
auditkit eval --model hf:gpt2 <<< "What is 2+2?"
With a YAML config:
model: openai:gpt-4o-mini
temperature: 0.0
prompts:
- "What is the capital of France?"
output: results.json
auditkit eval --config auditkit.yaml
What it evaluates
| Task | Technique | Example metrics |
|---|---|---|
| Benchmark text | MCQ, generation | ExactMatch, Bleu, F1Score |
| LLM output quality | LLM-as-judge | GEval, RubricItem |
| RAG pipelines | RAG | LexicalGroundedness, ContextOverlap |
| Safety | Red-team | KeywordDetector, DefconGrade, probes |
| Performance | Latency/throughput | LatencyStats, Throughput |
Output
Every evaluate() returns a RunResult:
| Field | Contents |
|---|---|
headline |
Per-metric averages |
predictions |
Per-sample scores with input, output, expected |
stats |
Full statistics (mean, std, min, max, count) |
fingerprint |
Stable sha256 for caching and comparison |
config |
The RunConfig used (seed, temperature, etc.) |
errors |
Any errors encountered during the run |
Examples
Colab-ready notebooks with real models and real datasets — see
examples/README.md for the full, current list.
| Example | File |
|---|---|
| Full pipeline: adapter, 5 metrics, LLM judge, LLM annotator | 01_full_evaluation_pipeline.ipynb |
| Generation across two real HF model families, compared | 02_generation_across_hf_families.ipynb |
| Custom annotators (regex, LLM-backed, fully custom) | 03_custom_annotators.ipynb |
| Metrics deep dive: built-in, custom, LLM-as-judge, RAG | 04_metrics_deep_dive.ipynb |
| Every data type/task kind: generative, MCQ, RAG, precomputed, chat | 05_data_types.ipynb |
Model comparison deep dive: compare_models() + RunComparison |
06_model_comparison.ipynb |
| Annotators across 4 real model families, then compared together | 07_annotators_across_models.ipynb |
CompareResult deep dive: what compare_models()'s native per-model support still can't express (different scorers per model), full method surface |
08_compare_result_deep_dive.ipynb |
GuardJudge implementation check: all 5 profiles, real and offline |
09_guard_judge_implementation_check.ipynb |
GuardJudge with real, flagship guard models |
10_guard_judge_legit_models.ipynb |
EncoderJudge: the base mechanism plus its two prebuilt subclasses |
11_encoder_judge_prebuilts.ipynb |
| Performance metrics on a real, large-scale evaluation | 12_performance_metrics_demo.ipynb |
Applications
Real-world scenarios answered end to end, not feature tours — see
examples/applications/README.md. Only 01 uses the
lm-evaluation-harness integration (ak.run_lmeval(), needs
auditkit[lmeval]), as an authoritative cross-check alongside the native
evaluation; none of the others do.
| Application | File |
|---|---|
How much does pruning severity (20%/40%/60%) degrade a model? Real BoolQ, real annotator, ship/no-ship verdicts, cross-checked via ak.run_lmeval() |
01_application_pruned_llama_boolq.ipynb |
| Healthcare: clinical QA correctness vs. grounding (PubMedQA) | 02_application_healthcare_pubmedqa.ipynb |
| Finance: QA grounded in real SEC 10-K filings | 03_application_finance_10k_qa.ipynb |
| E-commerce: review-sentiment triage at scale | 04_application_ecommerce_review_triage.ipynb |
| Education: auto-graded tutoring, correctness vs. explanation | 05_application_education_arc_tutor.ipynb |
| Enterprise search: internal knowledge assistant (real retrieval + RAG grounding) | 06_application_enterprise_search_rag.ipynb |
| LLM-as-judge via a real BERT NLI classifier, not a generative model | 07_application_bert_nli_judge.ipynb |
Repository map
| Directory | Description | README |
|---|---|---|
src/auditkit/ |
Core library — spine, runner, metrics, models, CLI | README |
src/auditkit/metrics/ |
Metric families | README |
src/auditkit/model/ |
Model backends (echo, openai, hf, vllm, etc.) | README |
src/auditkit/redteam/ |
Red team probes and detectors | README |
src/auditkit/scenarios/ |
Built-in benchmark datasets | README |
tests/ |
Test suite (750+ tests) | README |
examples/ |
Colab-ready example notebooks | README |
examples/applications/ |
Real-world application notebooks | README |
docs/ |
MkDocs documentation site (internal, requires Lexsi SSO) | index |
Developed by Lexsi Labs
Created by the team at Lexsi Labs, AuditKIT provides a unified evaluation spine for AI models across every technique and modality.
License
This project is released under the Lexsi Labs Source Available License (LSAL) v1.2 — free for academic research and teaching on MIT-like terms; use by any organization requires written acknowledgement or permission (Section 1A); a separate commercial license is required to sell it or embed it in a paid product.
Join Community / Contribute
- Issues and discussions are welcomed on the GitHub issue tracker.
- See the Contributing section for contribution standards, code reviews, and documentation tips.
Metadata
Release files for auditkit 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| auditkit-1.0.0.tar.gz | 2.8 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| auditkit-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.1 MB
Release files / auditkit-1.0.0.tar.gz
| Download URL | auditkit-1.0.0.tar.gz |
|---|---|
| Size | 2.8 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
73e8f4d878672514ebfe1049508d3ee1f87904c876765b57215c46587af7fe51
|
|
BLAKE2b-256 checksum How to use checksums |
78d672bf20b2a661bb836d3c2fe0bad54d825ad1792acc16839ee8aa2fb00b1c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|
Release files / auditkit-1.0.0-py3-none-any.whl
| Download URL | auditkit-1.0.0-py3-none-any.whl |
|---|---|
| Size | 292.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
393ba508c5818a25c8c1449ce3ad1ae9ec753ac773989f48d806d20c2fcaa928
|
|
BLAKE2b-256 checksum How to use checksums |
dea4450cf5bd7e418cf142ed7f2b91d9d89f18b09a299e305afa707f29e87cf9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|