An easy-to-extend LLM annotator for robust, resumable data annotation.

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

BramVanroy

These details have not been verified by PyPI

Project description

Robust, resumable LLM dataset annotation

PyPI version

llm-annotator is a Python 3.12+ library for robust, resumable LLM-driven dataset annotation and generation.

It supports multiple providers through pluggable clients:

vLLM offline inference: VLLMOfflineClient
vLLM server API: VLLMClient
OpenAI API: OpenAIClient
Anthropic API: ClaudeClient

Key capabilities:

Staged pipeline: prepare_data + run_annotation separates expensive template application and sorting from model inference, enabling SLURM and cluster restart workflows.
Resumable processing with JSONL checkpoints.
Annotation of existing datasets and generation from scratch.
Structured outputs via JSON schema.
Retry and validation hooks for robust pipelines.
Optional Hugging Face Hub upload cadence for both prepared data and outputs.
Context-manager cleanup of client resources.

It is not intended for parallel, multi-node, multi-instance generation. If that is what you are after, maybe datatrove is something for you.

Documentation

Read the full documentation at bramvanroy.github.io/llm-annotator.

Provider setup reference: docs/provider-info.md

Installation

Recommended:

uv add llm-annotator

pip install llm-annotator

Install provider extras as needed:

uv add "llm-annotator[vllm]"
uv add "llm-annotator[vllm-flashinfer]"  # Faster if your hardware supports it
uv add "llm-annotator[openai]"
uv add "llm-annotator[anthropic]"

See docs/provider-info.md for auth environment variables and provider-specific setup notes.

Usage

One-step convenience

Annotate an existing dataset:

from llm_annotator import Annotator, VLLMOfflineClient

client = VLLMOfflineClient(
    model="meta-llama/Llama-3.2-3B-Instruct",
    max_model_len=4096,
)

with Annotator(client=client, verbose=True) as anno:
    ds = anno.annotate_dataset(
        output_dir="outputs/sentiment",
        prompt_template="Classify the sentiment of this text: {text}",
        dataset_name="stanfordnlp/imdb",
        dataset_split="test",
        max_num_samples=100,
    )

Generate a dataset from scratch:

from llm_annotator import Annotator, OpenAIClient

client = OpenAIClient(model="gpt-4o-mini")

with Annotator(client=client) as anno:
    ds = anno.generate_dataset(
        output_dir="outputs/generated-qa",
        prompts="Write a short geography quiz question with answer.",
        max_num_samples=200,
    )

Two-step staged workflow

For large datasets or cluster (SLURM) environments, split the pipeline explicitly into a preparation step and a generation step. prepare_data applies prompt templates, optional sorting, and saves the prepared artifacts locally and to Hugging Face Hub. run_annotation then handles only model inference. If generation fails, re-run run_annotation with prepared_hub_id pointing to the Hub backup: preparation is skipped.

from llm_annotator import Annotator, VLLMOfflineClient

client = VLLMOfflineClient(
    model="meta-llama/Llama-3.2-3B-Instruct",
    max_model_len=4096,
)

HUB_ID = "my-org/imdb-prepared"  # Hub repo for prepared data backup

with Annotator(client=client, verbose=True) as anno:
    # Step 1: prepare data (reuses local cache or Hub backup if available)
    prepared_dataset, local_path, hub_id = anno.prepare_data(
        output_dir="outputs/imdb-sentiment",
        prompt_template="Classify the sentiment of this text: {text}",
        dataset_name="stanfordnlp/imdb",
        dataset_split="test",
        max_num_samples=100,
        sort_by_length=True,
        prepared_hub_id=HUB_ID,
    )

    # Step 2: run generation against the prepared data
    ds = anno.run_annotation(
        output_dir="outputs/imdb-sentiment",
        prompt_template="Classify the sentiment of this text: {text}",
        prepared_dataset=prepared_dataset,
        new_hub_id="my-org/imdb-annotated",
        upload_every_n_samples=500,
    )

To force a fresh preparation (ignoring any cached or Hub-stored artifacts), pass force_data_preparation=True to prepare_data or to annotate_dataset.

See the documentation for more examples, including:

Structured output with JSON schemas
Custom validation and post-processing
Generating datasets from scratch

Or check out the examples/ directory for complete working examples.

Testing

Install development dependencies first:

uv sync --dev

Run the default checks:

make style
make quality
make test
make typecheck

Pytest marker targets:

# Fast tests (same as `make test`)
make test-fast

# Slow tests only
make test-slow

# Integration tests only
make test-integration

# Entire suite (fast + slow)
make test-all

You can also run markers directly with pytest:

uv run pytest -m "not slow"
uv run pytest -m "slow"
uv run pytest -m "integration"

Slow and integration tests may load local models, require more runtime, or depend on optional components.

Building documentation

Local versioned docs preview (uses mike on a temporary local branch):

make serve-docs

Override version metadata when needed:

make serve-docs DOCS_VERSION=0.4.0 DOCS_ALIAS=latest DOCS_SOURCE_REF=v0.4.0

Docs are published with mike on release tags through .github/workflows/docs.yml.

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

BramVanroy

These details have not been verified by PyPI

Release history Release notifications | RSS feed

0.10.7

Jun 2, 2026

0.10.6

Jun 2, 2026

0.10.5

Jun 2, 2026

0.10.4

Jun 2, 2026

This version

0.10.3

Jun 2, 2026

0.10.2

Jun 2, 2026

0.10.1

Jun 2, 2026

0.10.0

May 29, 2026

0.9.2

May 26, 2026

0.9.1

May 26, 2026

0.9.0

May 25, 2026

0.8.1

May 25, 2026

0.8.0

May 23, 2026

0.7.2

May 22, 2026

0.7.0

May 22, 2026

0.6.0

Dec 2, 2025

0.4.0

Nov 5, 2025

0.3.4

Nov 5, 2025

0.3.3

Nov 2, 2025

0.3.2

Oct 30, 2025

0.3.1

Oct 22, 2025

0.3.0.post1

Oct 22, 2025

0.3.0

Oct 22, 2025

0.2.9

Oct 22, 2025

0.2.8

Oct 8, 2025

0.2.7

Oct 7, 2025

0.2.6

Oct 4, 2025

0.2.5

Oct 4, 2025

0.2.4

Oct 4, 2025

0.2.3

Oct 4, 2025

0.2.2

Oct 4, 2025

0.2.1

Oct 3, 2025

0.2.0

Oct 3, 2025

0.1.1

Oct 2, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm_annotator-0.10.3.tar.gz (331.3 kB view details)

Uploaded Jun 2, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

llm_annotator-0.10.3-py3-none-any.whl (81.2 kB view details)

Uploaded Jun 2, 2026 Python 3

File details

Details for the file llm_annotator-0.10.3.tar.gz.

File metadata

Download URL: llm_annotator-0.10.3.tar.gz
Upload date: Jun 2, 2026
Size: 331.3 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: uv/0.11.18 {"installer":{"name":"uv","version":"0.11.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llm_annotator-0.10.3.tar.gz
Algorithm	Hash digest
SHA256	`7477276f15bb4687a554a5ce7f6ceda17525bdb9a3b2210476c138150b7993f2`
MD5	`0ebf998e01b39cc656420fca28ef45b7`
BLAKE2b-256	`fdf4ce266c12e6694c86e0f0eb9106415dc27da068d460d847c2d3f0ad8c56d2`

See more details on using hashes here.

File details

Details for the file llm_annotator-0.10.3-py3-none-any.whl.

File metadata

Download URL: llm_annotator-0.10.3-py3-none-any.whl
Upload date: Jun 2, 2026
Size: 81.2 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: uv/0.11.18 {"installer":{"name":"uv","version":"0.11.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llm_annotator-0.10.3-py3-none-any.whl
Algorithm	Hash digest
SHA256	`3c16ce117e0f2848d08f95549fe3463975a8c5651b92d09e8100b046343f5ca3`
MD5	`026f5257edf0a01d68a62768d23493c2`
BLAKE2b-256	`8712bf411be591478cba5724efb32fe1c7fcc44e41547b189e2d2f8e09b7f72c`

See more details on using hashes here.

llm-annotator 0.10.3

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Project description

Robust, resumable LLM dataset annotation

Documentation

Installation

Usage

One-step convenience

Two-step staged workflow

Testing

Building documentation

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes