Robust, resumable LLM dataset annotation
llm-annotator is a Python 3.12+ library for robust, resumable
LLM-driven dataset annotation and generation.
It supports multiple providers through pluggable clients:
- vLLM offline inference (in-process):
VLLMOfflineClient - vLLM online inference (server API):
VLLMOnlineClient - OpenAI API:
OpenAIClient - Anthropic API:
ClaudeClient
Key capabilities:
- No-code config runs: describe prompts, schemas, model, dataset and
multiple chained annotation steps in one JSON/YAML file and run it with
llm-annotate my-pipeline.yaml. - Staged pipeline:
prepare_data+run_annotationseparates expensive template application and sorting from model inference, enabling SLURM and cluster restart workflows. - Multi-server vLLM:
VLLMQueueAnnotatorruns one workload over a pool of vLLM servers (e.g. one per GPU of a multi-node allocation); seeexamples/vllm-server-pool/for a config-driven and a Python-API example. - Resumable processing with JSONL checkpoints.
- Annotation of existing datasets and generation from scratch.
- Structured outputs via JSON schema.
- Retry and validation hooks for robust pipelines.
- Optional Hugging Face Hub upload cadence for both prepared data and outputs.
- Context-manager cleanup of client resources.
It is not intended for parallel, multi-node, multi-instance generation.
If that is what you are after, maybe datatrove
is something for you.
Documentation
Read the full documentation at bramvanroy.github.io/llm-annotator.
Provider setup reference: docs/provider-info.md
Installation
Recommended:
uv add llm-annotator
or
pip install llm-annotator
Install provider extras as needed:
uv add "llm-annotator[vllm]"
uv add "llm-annotator[openai]"
uv add "llm-annotator[anthropic]"
See docs/provider-info.md for auth environment variables and provider-specific setup notes.
Prebuilt vLLM kernels
In the context of SLURM it may be advisable to have the vLLM kernels prebuilt so that time is not wasted for JIT-compilation, no storage contention in the case of multiprocessing, etc. Installing the kernels up front avoids both, which matters most when you serve models on a cluster.
These kernels cannot be shipped as an extra of llm-annotator: flashinfer-jit-cache
is not on PyPI at all (it is published per CUDA version on FlashInfer's own
index) and the flashinfer-cubin on PyPI trails the releases vLLM pins
against. An extra would therefore fail to resolve for anyone installing
llm-annotator from PyPI. Install them next to the vllm extra instead,
matching the flashinfer-python version vLLM pulled in and the CUDA version
your torch wheel was built against:
version=$(python -c "import importlib.metadata as m; print(m.version('flashinfer-python'))")
cuda=cu$(python -c "import torch; print(torch.version.cuda.replace('.', ''))")
uv pip install "flashinfer-cubin==$version" --index-url https://flashinfer.ai/whl/
uv pip install "flashinfer-jit-cache==$version" --index-url "https://flashinfer.ai/whl/$cuda/"
Use pip install instead of uv pip install if you installed with pip. The
CUDA version comes from torch.version.cuda.
Usage
One-step convenience
Annotate an existing dataset:
from llm_annotator import Annotator, VLLMOfflineClient
client = VLLMOfflineClient(
model="meta-llama/Llama-3.2-3B-Instruct",
max_model_len=4096,
)
with Annotator(client=client, verbose=True) as anno:
ds = anno.annotate_dataset(
output_dir="outputs/sentiment",
prompt_template="Classify the sentiment of this text: {text}",
dataset_name="stanfordnlp/imdb",
dataset_split="test",
max_num_samples=100,
)
Generate a dataset from scratch:
from llm_annotator import Annotator, OpenAIClient
client = OpenAIClient(model="gpt-4o-mini")
with Annotator(client=client) as anno:
ds = anno.generate_dataset(
output_dir="outputs/generated-qa",
prompts="Write a short geography quiz question with answer.",
max_num_samples=200,
)
Two-step staged workflow
For large datasets or cluster (SLURM) environments, split the pipeline
explicitly into a preparation step and a generation step. prepare_data
applies prompt templates, optional sorting, and saves the prepared
artifacts locally and to Hugging Face Hub. run_annotation then handles
only model inference. If generation fails, re-run it with the same
output_dir and hub_id: the prepared data is restored and the samples
already recorded in the progress files are skipped.
A single hub_id drives every Hub destination: the prepared data and the
JSONL progress backup live on temporary branches of that repo, the final
dataset is pushed to its main branch, and both temporary branches are
deleted once the run completes.
from llm_annotator import Annotator, VLLMOfflineClient
client = VLLMOfflineClient(
model="meta-llama/Llama-3.2-3B-Instruct",
max_model_len=4096,
)
HUB_ID = "my-org/imdb-sentiment" # backups *and* the final dataset
with Annotator(client=client, verbose=True) as anno:
# Step 1: prepare data (reuses local cache or Hub backup if available)
prepared_dataset, local_path, hub_id = anno.prepare_data(
output_dir="outputs/imdb-sentiment",
prompt_template="Classify the sentiment of this text: {text}",
dataset_name="stanfordnlp/imdb",
dataset_split="test",
max_num_samples=100,
sort_by_length=True,
hub_id=HUB_ID,
)
# Step 2: run generation against the prepared data
ds = anno.run_annotation(
output_dir="outputs/imdb-sentiment",
prompt_template="Classify the sentiment of this text: {text}",
prepared_dataset=prepared_dataset,
hub_id=HUB_ID,
upload_every_n_samples=500,
)
To force a fresh preparation (ignoring any cached or Hub-stored artifacts),
pass force_data_preparation=True to prepare_data or to annotate_dataset.
Run from a config file
The same work can be described in a single JSON or YAML file and run without writing any Python:
llm-annotate my-pipeline.yaml
# or, from a checkout: python scripts/annotate.py my-pipeline.yaml
A config lists one or more steps that run in order, each annotating the dataset the previous one produced. That is what makes generate-then-judge workflows possible: one model writes question-answer pairs, a second rates them.
output_dir: outputs/pipeline-qa
dataset:
name: stanfordnlp/imdb
split: test
max_num_samples: 20
client:
provider: vllm_offline
model: Qwen/Qwen3-8B
options:
max_completion_tokens: 512
steps:
- name: write-qa
prompt_file: prompts/write_qa.md
output_schema_file: schemas/qa.json # produces `question`, `answer`
filter_invalid: true
rename:
question: question_v1
- name: rate-qa
prompt: "Rate this question about the text.\n\n{text}\n\nQ: {question_v1}"
output_schema_file: schemas/rating.json
client:
provider: claude # a different judge
model: claude-haiku-4-5
Paths inside the config resolve relative to the config file, so a config directory is self-contained. Finished steps write a snapshot and are skipped on a re-run, so an interrupted pipeline resumes rather than starting over.
A complete, runnable example lives in examples/pipeline-qa/, and the full key reference is in docs/pipeline.md.
See the documentation for more examples, including:
- Structured output with JSON schemas
- Custom validation and post-processing
- Generating datasets from scratch
Or check out the examples/ directory for complete working examples.
Testing
Install development dependencies first:
uv sync --dev
Run the default checks:
make style
make quality
make test
make typecheck
Pytest marker targets:
# Fast tests (same as `make test`)
make test-fast
# Slow tests only
make test-slow
# Integration tests only
make test-integration
# Entire suite (fast + slow)
make test-all
You can also run markers directly with pytest:
uv run pytest -m "not slow"
uv run pytest -m "slow"
uv run pytest -m "integration"
Slow and integration tests may load local models, require more runtime, or depend on optional components.
Building documentation
Local versioned docs preview (uses mike on a temporary local branch):
make serve-docs
Override version metadata when needed:
make serve-docs DOCS_VERSION=0.4.0 DOCS_ALIAS=latest DOCS_SOURCE_REF=v0.4.0
Docs are published with mike on release tags through
.github/workflows/docs.yml.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llm_annotator-0.12.1.tar.gz.
File metadata
- Download URL: llm_annotator-0.12.1.tar.gz
- Upload date:
- Size: 512.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9b844c2e4b0cf6ae73ed00246f85ea6415b42065b612fed8f24de9a471fda0b0
|
|
| MD5 |
781c8cdc38f5d5cd0d594891c11abb7b
|
|
| BLAKE2b-256 |
74183fa29e237c5ac4d4c49e99217bfb240ceecab863651d6f580f1e2d5dd4e0
|
File details
Details for the file llm_annotator-0.12.1-py3-none-any.whl.
File metadata
- Download URL: llm_annotator-0.12.1-py3-none-any.whl
- Upload date:
- Size: 115.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e7c1e87374f8391796f22ae9faeb2b6be5dd8ef8e3b88d1c7ea9673b580defbb
|
|
| MD5 |
1118cc3b935fc78c5762c91e3c7171f4
|
|
| BLAKE2b-256 |
88b0c091d2b645eb3e6e81ab74905463937293c7fcbb0635d72721521a641c1c
|