Skip to main content

📘 llm-lab

A lightweight, notebook-first framework for running reproducible LLM experiments, comparing multiple models, and evaluating custom metrics across JSONL/CSV datasets.

🚀 Features

Multi-model evaluation — compare OpenAI models in one experiment

CSV + JSONL support — ideal for business analysts & product teams

Custom metrics — create your own quality checks with a one-line decorator

Reproducible runs — every run stored in a SQLite database

Built-in comparison tools — leaderboard, per-run summaries, charts

Notebook-first design — built to be used directly in Jupyter

📦 Installation

For now (local development):

pip install -e .

PyPI publishing will come later.

📄 Dataset Formats

llm-lab supports:

CSV files

JSONL (one example per line)

JSON lists

Each example must contain:

expected_output — the reference answer (used by metrics)

Other fields can be anything you want to use inside your prompt.

Example .csv question,context,expected_output "What is 2+2?", "", "4" "Capital of France?", "", "Paris"

🧪 Quickstart from llm_lab import Experiment, run_experiment, compare_models

exp = Experiment( name="qa_eval", dataset_path="data/qa_sample.csv", prompt_template="Q: {question}\nA:", model_names=["gpt-4o-mini", "gpt-3.5-turbo"], metrics=["exact_match"], model_params={"temperature": 0.0}, )

results = run_experiment(exp, max_examples=20)

compare_models(results)

📊 Comparing Models

Leaderboard:

from llm_lab import show_leaderboard show_leaderboard(metric="exact_match")

Detailed run summary:

from llm_lab import summarize_run summarize_run(results[0].run_id)

Plot metric across models:

from llm_lab import plot_model_metric plot_model_metric(results, metric="exact_match")

🧩 Custom Metrics

You can add your own metrics easily.

from llm_lab import register_metric

@register_metric def contains_expected(example, output): expected = example["expected_output"].lower() return {"contains_expected": float(expected in output.lower())}

Then include it in your experiment:

metrics=["exact_match", "contains_expected"]

🧰 CSV Normalization Utility

Business analysts often have arbitrary column names (Ideal_Answer, target, etc). Use the helper to normalize your raw CSV into an llm-lab-compatible dataset.

from llm_lab import prepare_csv_for_llm_lab

prepare_csv_for_llm_lab( src_path="raw/support_eval_raw.csv", dst_path="data/support_eval.csv", expected_col="Ideal_Answer", )

Now support_eval.csv can be used directly.

🗂 Project Structure llm-lab/ │ ├── data/ │ └── qa_sample.csv ├── notebooks/ │ └── 01_quickstart.ipynb ├── src/ │ └── llm_lab/ │ ├── experiment.py │ ├── metrics.py │ ├── model_client.py │ ├── storage.py │ ├── utils.py │ ├── analysis.py │ └── init.py ├── pyproject.toml └── README.md

🧠 Philosophy

llm-lab is designed with three principles:

Minimalism — no boilerplate, no YAML configs, no heavy framework

Reproducibility — all experiments are logged in SQLite

Accessibility — analysts and PMs should be able to use it with zero ML background

It aims to be the pytest + scikit-learn of LLM evaluation— simple, composable, and transparent.

🤝 Contributing

Pull requests are welcome.

For major changes, please open an issue first to discuss what you'd like to change.

📄 License

MIT License.

Metadata

Release files for llm-lab 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-lab 0.1.0
File Size Uploaded
llm_lab-0.1.0.tar.gz 12.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-lab 0.1.0
File Interpreter ABI Platform
llm_lab-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 24.4 kB

Release files / llm_lab-0.1.0.tar.gz

Download URL llm_lab-0.1.0.tar.gz
Size 12.0 kB
Tags Source
SHA-256 checksum
How to use checksums
e8d12fad9c3d494ff785b1c24404b035eede2661ff4ba4425a36cb59e6b8e85d
BLAKE2b-256 checksum
How to use checksums
e8609a76aecb87d2401a7bfb6435497f4ca9d1d4ac7be8aa524b6fdd1d2376ca
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.5

Release files / llm_lab-0.1.0-py3-none-any.whl

Download URL llm_lab-0.1.0-py3-none-any.whl
Size 12.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
68d65b9768c43598a4c085a532171e1adbd6caf7d786d0d70c5552a892904ae5
BLAKE2b-256 checksum
How to use checksums
8a77fd274bffd3581f3f5c58e4b170b466449244d0c90575cd3498fa1aef04a0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.5

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page