Skip to main content

LLM Extractinator logo

LLM Extractinator

Turn unstructured text into Pydantic-validated structured data using local LLMs via Ollama.

Python License Docs Tests PyPI Streamlit

⚠️ Prototype. This project is under active development — interfaces, task formats, and directory names may still change. LLMs can also hallucinate, so always inspect and validate the output before using it downstream.

LLM Extractinator reads unstructured text (reports, clinical notes, the text column of a CSV) and returns structured, schema-validated data. You describe the shape you want as a Pydantic model, and the tool prompts a local LLM (via Ollama) to fill it in and validates the result against your schema.

It ships in three flavours that all do the same thing:

Interface Command Best for
Studio — a Streamlit app launch-extractinator Designing schemas and tasks with no code, then running and inspecting them
CLI extractinate Repeatable, scriptable runs on a workstation or server
Python API from llm_extractinator import extractinate Calling extraction from your own scripts or notebooks

📚 Full documentation: https://diagnijmegen.github.io/llm_extractinator/


How it works

        ┌────────────┐     ┌───────────────┐     ┌──────────────┐     ┌─────────────┐
Text →  │  Dataset   │  +  │ Output schema │  →  │  Local LLM   │  →  │  Validated  │
        │ (CSV/JSON) │     │  (Pydantic)   │     │  via Ollama  │     │ JSON output │
        └────────────┘     └───────────────┘     └──────────────┘     └─────────────┘
                    a Task JSON ties these together

You need three things, all described by a small task file:

  1. a dataset — a CSV or JSON file with a text field to read from,
  2. an output schema — a Pydantic model (OutputParser) describing the fields to extract, and
  3. a task JSON — pointing at the dataset, the text field, and the schema.

You can create all three in the Studio, or write them by hand.


1. Installation

🐳 Recommended: Docker

The easiest and most reliable way to run LLM Extractinator is the GPU-ready Docker image. It bundles Python, Ollama, and the Studio into a single container, so the only thing you install is Docker itself — no environments, no separate Ollama setup.

Create the folders that will be mounted into the container, then start it:

mkdir -p data examples tasks output ollama_models

docker run --rm --gpus all \
  -p 127.0.0.1:8501:8501 \
  -p 11434:11434 \
  -v $(pwd)/data:/app/data \
  -v $(pwd)/examples:/app/examples \
  -v $(pwd)/tasks:/app/tasks \
  -v $(pwd)/output:/app/output \
  -v $(pwd)/ollama_models:/root/.ollama \
  lmmasters/llm_extractinator:latest

This launches the Studio at http://127.0.0.1:8501. Drop --gpus all to run CPU-only. Full details — including the Windows/PowerShell command and a shell mode — are in the Docker guide.

Alternative: local install

Prefer your own Python environment? You'll need Ollama installed and running separately (curl -fsSL https://ollama.com/install.sh | sh on Linux, or the installer from ollama.com/download). Then:

conda create -n llm_extractinator python=3.11
conda activate llm_extractinator
pip install llm_extractinator

(Or pip install -e . from a clone to hack on it.) See Installation for details. You don't need to pull a model yourself — the tool pulls the one you ask for on first use.


2. Quick start

If you started the Docker image above, the Studio is already running — open http://127.0.0.1:8501. From a local install, launch it with:

launch-extractinator

Either way, the Studio follows a simple Task → Run → Results flow: configure or build a task, run it while watching the logs, then explore the output record by record.

Once a task file exists, you can also run it from the terminal…

extractinate --task_id 1 --model_name "phi4"

…or from Python:

from llm_extractinator import extractinate

extractinate(task_id=1, model_name="phi4")

For a complete first run — dataset, schema, task, output — follow the Quickstart.


3. Task files

A task describes what to extract and from where. Task files live in tasks/ and must be named Task<NNN>...json, where <NNN> is a zero-padded three-digit ID:

tasks/
├── Task001.json               # ID 1 — the Studio saves this form
├── Task002_reports.json       # ID 2 — an optional _name suffix is allowed
└── parsers/
    └── report.py              # the output schema referenced below

A minimal task:

{
  "Description": "Extract product name and price from each row of text",
  "Data_Path": "products.csv",
  "Input_Field": "text",
  "Parser_Format": "product_schema.py"
}
Field Meaning
Description Plain-language instruction shown to the model
Data_Path Dataset filename, relative to the data directory (--data_dir, default data/)
Input_Field Column/key holding the text to read
Parser_Format Python file in tasks/parsers/ defining the OutputParser schema
Example_Path (optional) Few-shot examples file, relative to --example_dir

You reference a task on the CLI by its ID: Task002_reports.json → --task_id 2.


4. Output schemas

The output schema defines the fields to extract and their types. It's a plain Pydantic model whose top-level class must be named OutputParser:

from pydantic import BaseModel

class OutputParser(BaseModel):
    product_name: str
    price: float

Don't want to write Python? Build one visually — the Studio has an Output Schema Builder (also available standalone via build-parser) that generates the file for you and saves it into tasks/parsers/.

See the output-schema guide for nested models, optional fields, and Literal choices.


5. Where results go

Each run writes to:

output/<run_name>/<TaskName>-run<N>/nlp-predictions-dataset.json

Every record contains your original input columns, the extracted fields, and a status of "success" or "failure". The Studio's Results tab lets you filter and inspect these interactively. Details in Understanding output.


6. Documentation map

Page What's in it
Quickstart A complete first extraction, end to end
Installation Python + Ollama setup
Preparing data What your CSV/JSON should look like
Output schema Designing the Pydantic model
Studio The Streamlit app, tab by tab
CLI usage & Settings reference Every flag and task field
Few-shot prompting Guiding the model with examples
Understanding output Folder layout and record shape
Troubleshooting Common errors and fixes
Docker & CPU-only Containers and no-GPU setups

7. Contributing

Pull requests are welcome. If you change task naming, required fields, or CLI flags, please update the docs in the same PR so the two stay in sync.

pip install -e ".[test]"
pytest                              # offline suite — no GPU, no Ollama, a few seconds
python devtests/run.py --model phi4  # the same pipeline against a real model

The two suites answer different questions: pytest proves the plumbing with the model faked out, and devtests/ covers what a faked model structurally cannot — whether the grammar holds an enum, whether the output budget is big enough, how far the token estimate drifts on your language. See Development & testing.


8. Citation & attribution

If you use this tool in your research, please cite:

https://doi.org/10.1093/jamiaopen/ooaf109

Developed by the Oncology Research Group at the Diagnostic Image Analysis Group (DIAG), Radboud University Medical Center. 🔗 diagnijmegen.nl/research/oncology

Contact:

Name Email
Luc Builtjes luc.builtjes@radboudumc.nl
Alessa Hering alessa.hering@radboudumc.nl

Metadata

Release files for llm-extractinator 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-extractinator 0.7.0
File Size Uploaded
llm_extractinator-0.7.0.tar.gz 340.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-extractinator 0.7.0
File Interpreter ABI Platform
llm_extractinator-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 646.9 kB

Release files / llm_extractinator-0.7.0.tar.gz

Download URL llm_extractinator-0.7.0.tar.gz
Size 340.3 kB
Tags Source
SHA-256 checksum
How to use checksums
cea36ae75c4e51980029d991a89b481643a0b4cce2a0ae70c5c09faef8047102
BLAKE2b-256 checksum
How to use checksums
268925a6edc3b129b54350c13e8cb626cc50f1657c3254b15dfb0e108b142887
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / llm_extractinator-0.7.0-py3-none-any.whl

Download URL llm_extractinator-0.7.0-py3-none-any.whl
Size 306.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ca69a4395d9664066082cd5be10b06e63c78f4c38188c1629d0a8e32d9204bea
BLAKE2b-256 checksum
How to use checksums
a1c74b7b6fa12e9d24cd247cbc1fa8e655619f287c7c476c897e589c1098ca75
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.7.1

2 release files

This release

0.7.0 This release

2 release files

0.6.0

2 release files

0.5.14

2 release files

0.5.10

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page