Skip to main content

Agentic system for automated extraction-schema generation from natural language descriptions

Project description

schematize logo

schematize

Turn a plain research problem into a typed, data-tested extraction schema.

PyPI version Docs Python CI License: MIT

Documentation · Examples · Quickstart · Pipeline · API reference


Department of Artificial Intelligence logo   Wrocław University of Science and Technology logo


schematize is a Python library that turns a natural-language description of what you want to extract into a typed, validated extraction schema — field names, types, descriptions, and enums — ready to drive structured information extraction from a document collection.

Instead of hand-writing JSON schemas and discovering their gaps in production, you describe the problem. A multi-agent LangGraph pipeline asks clarifying questions, drafts a schema, critiques and refines it, tests it against real documents from your corpus, and opens a chat for final tweaks.

See it in action first: worked examples show three real, unedited pipeline runs — the clarifying-question chat and the resulting schema — each run through five different LLMs.

Why schematize?

  • Designing extraction schemas by hand is slow and brittle. You guess the fields, miss edge cases, and only find out when extraction quality is poor.
  • Raw "ask an LLM for a schema" gives you an untested first draft. No critique loop, no contact with your actual data, no typing guarantees.
  • schematize closes the loop: clarify → draft → criteria-based refinement → data-grounded refinement against retrieved documents → interactive chat. The output is a Pydantic model you can plug straight into an extraction pipeline.
  • Use any LLM. schematize accepts any LangChain chat model and talks to it through the OpenAI-compatible API, so the official OpenAI API, LiteLLM, vLLM, Ollama, and many other providers/servers work out of the box. We ran our experiments through a LiteLLM proxy. No provider lock-in.
  • Bring your own data. Implementing a retriever is one async method; HuggingFace and Weaviate adapters are built in. A schema-coverage evaluator is included too.
Hand-written schema Ask-an-LLM once schematize
Clarifies an ambiguous request
Iterative critique & refinement
Validated against real documents
Typed Pydantic output manual
Built-in coverage evaluation

Install

pip install schematize

Optional adapters and tooling:

pip install "schematize[huggingface]"   # FAISS retriever over HuggingFace datasets
pip install "schematize[weaviate]"      # hybrid-search retriever for a Weaviate instance
pip install "schematize[scripts]"       # Hydra-based CLI runners

Requires Python 3.12+. Full options in the installation guide.

Quickstart

from langchain_openai import ChatOpenAI
from schematize import SchemaGenerator, load_prompts


# 1. Any object with this async __call__ is a valid retriever (DocumentRetriever protocol).
#    Return documents relevant to `query`; each is shown to the data-assessment agent.
class MyRetriever:
    async def __call__(self, query: str, max_docs: int = 100) -> list:
        return [
            {"text": "The court awarded 15,000 PLN in damages for breach of personal rights..."},
            {"text": "Claim dismissed; the plaintiff failed to prove the violation..."},
        ]


# 2. Load bundled prompts for your language/domain (en|pl × law|tax, plus en/general).
prompts = load_prompts(language="en", system_type="law")

llm = ChatOpenAI(model="gpt-4o", temperature=0.2)

generator = SchemaGenerator(llm=llm, retriever=MyRetriever(), **prompts)

# 3. Run the pipeline (interactive: it asks clarifying questions in the terminal).
state = generator.stream_graph_updates(
    "Study personal-rights violations in civil cases and assess their severity."
)
print(state["current_schema"])

A generated schema looks like this — a typed spec you can act on immediately:

{
    "fields": [
        {"name": "violation_type", "type_": "enum", "enum_name": "ViolationType",
         "enum_values": ["privacy", "reputation", "image", "bodily_integrity"],
         "description": "Category of personal right that was violated."},
        {"name": "severity", "type_": "integer",
         "description": "Severity of the violation on a 0–5 scale."},
        {"name": "compensation_awarded", "type_": "boolean",
         "description": "Whether monetary compensation was granted."},
        {"name": "compensation_amount", "type_": "float",
         "description": "Awarded amount in PLN, if any."},
    ]
}

Turn it into a Pydantic model and use it for extraction:

from schematize import SchemaFields, DynamicModelFactory

model_cls = DynamicModelFactory()(SchemaFields(**state["current_schema"]))

Use any LLM (via LiteLLM)

SchemaGenerator takes any LangChain BaseChatModel. Because it also honours an OpenAI-compatible base_url, the simplest way to reach any provider is to put a LiteLLM proxy in front and point schematize at it — the setup we used for our experiments:

from langchain_openai import ChatOpenAI

# Point at a LiteLLM proxy; the model name routes to OpenAI, Anthropic, Gemini, local, etc.
llm = ChatOpenAI(model="claude-opus-4-8", base_url="http://localhost:4000", api_key="sk-litellm")

One interface, 100+ providers, no lock-in. See Configuration.

Retrieval is pluggable

The core library has no retrieval dependency. Implementing your own retriever is a single async method — the DocumentRetriever protocol shown in the quickstart — so you can wrap Elasticsearch, Postgres FTS, a REST API, or any vector store.

Two adapters ship in the box:

HuggingFace ([huggingface]) — FAISS index over any HuggingFace dataset, cached to disk:

from schematize.retrieval.huggingface import HuggingFaceRetriever

retriever = HuggingFaceRetriever(
    dataset_name="JuDDGES/pl-court-raw", text_column="text", index_path=".cache/court-index"
)

Weaviate ([weaviate]) — hybrid search against a Weaviate instance:

from schematize.retrieval.weaviate import WeaviateRetriever

retriever = WeaviateRetriever(collection_name="LegalDocuments")

Needs WV_URL, WV_PORT, WV_GRPC_PORT, WV_API_KEY. See the Weaviate guide.

MMLWRobertaV2Retriever (Polish-optimised, built on the HuggingFace base) is a worked example of a custom retriever — read it as a template for specialising retrieval to your own model or language. See the custom retriever guide.

Evaluate a schema

Score how well a schema can answer a set of expert questions:

import yaml
from langchain_openai import ChatOpenAI
from schematize import SchemaEvaluator
from schematize.settings import PROMPTS_PATH

with open(PROMPTS_PATH / "eval" / "schema_evaluator.yaml") as f:
    evaluation_prompt = yaml.safe_load(f)["schema_evaluator_prompt"]

evaluator = SchemaEvaluator(ChatOpenAI(model="gpt-4o"), evaluation_prompt)
result = evaluator.evaluate_schema(schema, questions=["How severe was the violation?", ...])
print(result.covered_questions, "/", result.total_questions)

More in the evaluation guide.

Command-line runners

With the [scripts] extra you get three console scripts:

schematize-run                              # interactive pipeline
schematize-run-mocked +case=en_age          # replay a stored case (no live prompts)
schematize-evaluate +case_name=age          # evaluate against expert questions

See the CLI guide for mocked-case files and options.

Reproducing our study

The experiments from our paper are driven by the mocked runner (schema generation) and the evaluator (schema scoring against expert questions). Cases live in data/cases/ and expert question sets in data/eval/ (pl_age, pl_personal_rights, pl_medical_errors).

The paper's experiments use the pl/law domain (Polish legal judgments); tax and general are additional prompt sets for use beyond the paper.

git clone https://github.com/pwr-ai/schematize && cd schematize
uv sync --extra scripts --extra huggingface

# Configure the LLM in a .env file (we used a LiteLLM proxy — see "Use any LLM" above)
printf 'API_KEY=...\nAPI_URL=...\n' > .env

# 1. Generate schemas for every case (multiple runs per case)
bash scripts/experiments/search_params.sh

# 2. Ablation over pipeline components (no problem-definition / no refinement / no data-grounding)
bash scripts/experiments/ablation.sh <model>

# 3. Evaluate generated schemas against expert questions
bash scripts/experiments/eval_multirun.sh <eval_model> <generation_model>

The shell scripts in scripts/experiments/ are thin Hydra wrappers; edit the MODEL/CASES variables at the top to change the grid. Results are written to the Hydra output directory.

Note: exact model names, seeds, and hyperparameters used in the paper are documented in the reproduction guide — fill in once the study is published.

Citation

If you use schematize in your research, please cite:

TBA

Development

uv sync --extra dev
make check    # ruff lint
make test     # pytest + coverage
make fix      # ruff --fix

License

Released under the MIT License.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

schematize-0.1.8.tar.gz (83.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

schematize-0.1.8-py3-none-any.whl (139.4 kB view details)

Uploaded Python 3

File details

Details for the file schematize-0.1.8.tar.gz.

File metadata

  • Download URL: schematize-0.1.8.tar.gz
  • Upload date:
  • Size: 83.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for schematize-0.1.8.tar.gz
Algorithm Hash digest
SHA256 7413be3cd82ae0dd26934b0774a8fc17e3799afc48336f15750af0274fc61310
MD5 09798b3a0bd280b1e5e310b3619e3476
BLAKE2b-256 eb2eee36dfc6a7f51d895d2d415dae35cd6700dc70c97999ae61fd20326270f9

See more details on using hashes here.

Provenance

The following attestation bundles were made for schematize-0.1.8.tar.gz:

Publisher: publish.yml on pwr-ai/schematize

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file schematize-0.1.8-py3-none-any.whl.

File metadata

  • Download URL: schematize-0.1.8-py3-none-any.whl
  • Upload date:
  • Size: 139.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for schematize-0.1.8-py3-none-any.whl
Algorithm Hash digest
SHA256 d6acc36fdb9dc4ca2aa29ed104a69124605926a37a497962ea87e1e5121772ba
MD5 b9521902cbf7a94028a08583ca2a6738
BLAKE2b-256 f3578eff0789d7963c884cc0484732d2d9ccdfb09eeb5881301577eebbed298f

See more details on using hashes here.

Provenance

The following attestation bundles were made for schematize-0.1.8-py3-none-any.whl:

Publisher: publish.yml on pwr-ai/schematize

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page