Skip to main content

Agentic system for automated extraction-schema generation from natural language descriptions

Project description

schematize logo

schematize

Turn a plain research problem into a typed, data-tested extraction schema.

PyPI version Docs Python CI License: MIT

Documentation · Examples · Quickstart · Pipeline · API reference


schematize is a Python library that turns a natural-language description of what you want to extract into a typed, validated extraction schema — field names, types, descriptions, and enums — ready to drive structured information extraction from a document collection.

Instead of hand-writing JSON schemas and discovering their gaps in production, you describe the problem. A multi-agent LangGraph pipeline asks clarifying questions, drafts a schema, critiques and refines it, tests it against real documents from your corpus, and opens a chat for final tweaks.

See it in action first: worked examples show three real, unedited pipeline runs — the clarifying-question chat and the resulting schema — each run through five different LLMs.

Why schematize?

  • Designing extraction schemas by hand is slow and brittle. You guess the fields, miss edge cases, and only find out when extraction quality is poor.
  • Raw "ask an LLM for a schema" gives you an untested first draft. No critique loop, no contact with your actual data, no typing guarantees.
  • schematize closes the loop: clarify → draft → criteria-based refinement → data-grounded refinement against retrieved documents → interactive chat. The output is a Pydantic model you can plug straight into an extraction pipeline.
  • Use any LLM. schematize accepts any LangChain chat model and talks to it through the OpenAI-compatible API, so the official OpenAI API, LiteLLM, vLLM, Ollama, and many other providers/servers work out of the box. We ran our experiments through a LiteLLM proxy. No provider lock-in.
  • Bring your own data. Implementing a retriever is one async method; HuggingFace and Weaviate adapters are built in. A schema-coverage evaluator is included too.
Hand-written schema Ask-an-LLM once schematize
Clarifies an ambiguous request
Iterative critique & refinement
Validated against real documents
Typed Pydantic output manual
Built-in coverage evaluation

Install

pip install schematize

Optional adapters and tooling:

pip install "schematize[huggingface]"   # FAISS retriever over HuggingFace datasets
pip install "schematize[weaviate]"      # hybrid-search retriever for a Weaviate instance
pip install "schematize[scripts]"       # Hydra-based CLI runners

Requires Python 3.12+. Full options in the installation guide.

Quickstart

from langchain_openai import ChatOpenAI
from schematize import SchemaGenerator, load_prompts


# 1. Any object with this async __call__ is a valid retriever (DocumentRetriever protocol).
#    Return documents relevant to `query`; each is shown to the data-assessment agent.
class MyRetriever:
    async def __call__(self, query: str, max_docs: int = 100) -> list:
        return [
            {"text": "The court awarded 15,000 PLN in damages for breach of personal rights..."},
            {"text": "Claim dismissed; the plaintiff failed to prove the violation..."},
        ]


# 2. Load bundled prompts for your language/domain (en|pl × law|tax, plus en/general).
prompts = load_prompts(language="en", system_type="law")

llm = ChatOpenAI(model="gpt-4o", temperature=0.2)

generator = SchemaGenerator(llm=llm, retriever=MyRetriever(), **prompts)

# 3. Run the pipeline (interactive: it asks clarifying questions in the terminal).
state = generator.stream_graph_updates(
    "Study personal-rights violations in civil cases and assess their severity."
)
print(state["current_schema"])

A generated schema looks like this — a typed spec you can act on immediately:

{
    "fields": [
        {"name": "violation_type", "type_": "enum", "enum_name": "ViolationType",
         "enum_values": ["privacy", "reputation", "image", "bodily_integrity"],
         "description": "Category of personal right that was violated."},
        {"name": "severity", "type_": "integer",
         "description": "Severity of the violation on a 0–5 scale."},
        {"name": "compensation_awarded", "type_": "boolean",
         "description": "Whether monetary compensation was granted."},
        {"name": "compensation_amount", "type_": "float",
         "description": "Awarded amount in PLN, if any."},
    ]
}

Turn it into a Pydantic model and use it for extraction:

from schematize import SchemaFields, DynamicModelFactory

model_cls = DynamicModelFactory()(SchemaFields(**state["current_schema"]))

Use any LLM (via LiteLLM)

SchemaGenerator takes any LangChain BaseChatModel. Because it also honours an OpenAI-compatible base_url, the simplest way to reach any provider is to put a LiteLLM proxy in front and point schematize at it — the setup we used for our experiments:

from langchain_openai import ChatOpenAI

# Point at a LiteLLM proxy; the model name routes to OpenAI, Anthropic, Gemini, local, etc.
llm = ChatOpenAI(model="claude-opus-4-8", base_url="http://localhost:4000", api_key="sk-litellm")

One interface, 100+ providers, no lock-in. See Configuration.

Retrieval is pluggable

The core library has no retrieval dependency. Implementing your own retriever is a single async method — the DocumentRetriever protocol shown in the quickstart — so you can wrap Elasticsearch, Postgres FTS, a REST API, or any vector store.

Two adapters ship in the box:

HuggingFace ([huggingface]) — FAISS index over any HuggingFace dataset, cached to disk:

from schematize.retrieval.huggingface import HuggingFaceRetriever

retriever = HuggingFaceRetriever(
    dataset_name="JuDDGES/pl-court-raw", text_column="text", index_path=".cache/court-index"
)

Weaviate ([weaviate]) — hybrid search against a Weaviate instance:

from schematize.retrieval.weaviate import WeaviateRetriever

retriever = WeaviateRetriever(collection_name="LegalDocuments")

Needs WV_URL, WV_PORT, WV_GRPC_PORT, WV_API_KEY. See the Weaviate guide.

MMLWRobertaV2Retriever (Polish-optimised, built on the HuggingFace base) is a worked example of a custom retriever — read it as a template for specialising retrieval to your own model or language. See the custom retriever guide.

Evaluate a schema

Score how well a schema can answer a set of expert questions:

import yaml
from langchain_openai import ChatOpenAI
from schematize import SchemaEvaluator
from schematize.settings import PROMPTS_PATH

with open(PROMPTS_PATH / "eval" / "schema_evaluator.yaml") as f:
    evaluation_prompt = yaml.safe_load(f)["schema_evaluator_prompt"]

evaluator = SchemaEvaluator(ChatOpenAI(model="gpt-4o"), evaluation_prompt)
result = evaluator.evaluate_schema(schema, questions=["How severe was the violation?", ...])
print(result.covered_questions, "/", result.total_questions)

More in the evaluation guide.

Command-line runners

With the [scripts] extra you get three console scripts:

schematize-run                              # interactive pipeline
schematize-run-mocked +case=en_age          # replay a stored case (no live prompts)
schematize-evaluate +case_name=age          # evaluate against expert questions

See the CLI guide for mocked-case files and options.

Reproducing our study

The experiments from our paper are driven by the mocked runner (schema generation) and the evaluator (schema scoring against expert questions). Cases live in data/cases/ and expert question sets in data/eval/ (pl_age, pl_personal_rights, pl_medical_errors).

The paper's experiments use the pl/law domain (Polish legal judgments); tax and general are additional prompt sets for use beyond the paper.

git clone https://github.com/pwr-ai/schematize && cd schematize
uv sync --extra scripts --extra huggingface

# Configure the LLM in a .env file (we used a LiteLLM proxy — see "Use any LLM" above)
printf 'API_KEY=...\nAPI_URL=...\n' > .env

# 1. Generate schemas for every case (multiple runs per case)
bash scripts/experiments/search_params.sh

# 2. Ablation over pipeline components (no problem-definition / no refinement / no data-grounding)
bash scripts/experiments/ablation.sh <model>

# 3. Evaluate generated schemas against expert questions
bash scripts/experiments/eval_multirun.sh <eval_model> <generation_model>

The shell scripts in scripts/experiments/ are thin Hydra wrappers; edit the MODEL/CASES variables at the top to change the grid. Results are written to the Hydra output directory.

Note: exact model names, seeds, and hyperparameters used in the paper are documented in the reproduction guide — fill in once the study is published.

Citation

If you use schematize in your research, please cite:

TBA

Development

uv sync --extra dev
make check    # ruff lint
make test     # pytest + coverage
make fix      # ruff --fix

License

Released under the MIT License.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

schematize-0.1.7.tar.gz (82.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

schematize-0.1.7-py3-none-any.whl (138.2 kB view details)

Uploaded Python 3

File details

Details for the file schematize-0.1.7.tar.gz.

File metadata

  • Download URL: schematize-0.1.7.tar.gz
  • Upload date:
  • Size: 82.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for schematize-0.1.7.tar.gz
Algorithm Hash digest
SHA256 3ee49f4f8acbdc4d27d890f8f509bcf0b723bd44e36b76e274c79e0273e9ccdd
MD5 31ccd1d3f0c6e75ae2898148a21bae7d
BLAKE2b-256 3d05bca58ac514edca33f182dcb714c1e288ca1cf73c5535491fa25dcf6f72f5

See more details on using hashes here.

Provenance

The following attestation bundles were made for schematize-0.1.7.tar.gz:

Publisher: publish.yml on pwr-ai/schematize

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file schematize-0.1.7-py3-none-any.whl.

File metadata

  • Download URL: schematize-0.1.7-py3-none-any.whl
  • Upload date:
  • Size: 138.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for schematize-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 fde87a7bfca725dd0f078432ff1e22f8518628f92f80c6d14aa10964a8e0d076
MD5 92bcbff26dc434aed78584b964ab78b3
BLAKE2b-256 3eb6b0eb626c3de8712d90dc4e070d543484b5350f9a6a3079a166ae99f22040

See more details on using hashes here.

Provenance

The following attestation bundles were made for schematize-0.1.7-py3-none-any.whl:

Publisher: publish.yml on pwr-ai/schematize

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page