Skip to main content

dtxt

Schema-centric bidirectional conversion between text and structured data.

dtxt is not related to dtx (an AI red-teaming tool).

Core features

  1. Schema inference (SchemaInferer): derive a schema from a collection of texts
  2. T2D (StructuredEntityExtractor): convert text into a schema-conformant object
  3. D2T (StructuredEntityRenderer): convert an object into text
  4. Round-trip verification (check_roundtrip): check that extract(render(obj)) ≈ obj for a given schema and backend

Install

pip install dtxt              # core only
pip install dtxt[anthropic]   # + Anthropic backend
pip install dtxt[openai]      # + OpenAI backend
pip install dtxt[llamacpp]    # + local GGUF models via llama.cpp
pip install dtxt[all]         # everything

The core package depends only on pydantic and jsonschema. Backends are optional extras, imported lazily.

Usage

import dtxt
from dtxt import Schema
from dtxt.backends import MockBackend

schema = Schema({
    "type": "object",
    "properties": {
        "name": {"type": "string", "x-dtxt-description": "the person's full name"},
        "age": {"type": "integer"},
    },
    "required": ["name", "age"],
})

# Each class takes its backend at construction -- there is no global config.
extractor = dtxt.StructuredEntityExtractor(MockBackend(), schema)
renderer = dtxt.StructuredEntityRenderer(MockBackend(), schema)

obj = extractor.extract("Alice is 30 years old.")
text = renderer.render({"name": "Alice", "age": 30})

result = dtxt.check_roundtrip(
    {"name": "Alice", "age": 30}, schema, renderer=renderer, extractor=extractor
)
result.ok  # True if extract(render(obj)) == obj on every schema field

Swap MockBackend for a real one:

extractor = dtxt.StructuredEntityExtractor(dtxt.backends.LlamaCpp("model.gguf", n_ctx=8192), schema)
renderer = dtxt.StructuredEntityRenderer(dtxt.backends.Anthropic("claude-sonnet-4-6"), schema)
inferer = dtxt.SchemaInferer(dtxt.backends.Anthropic("claude-sonnet-4-6"))

LlamaCpp locates a model either by local path (model_path) or by pulling from the Hugging Face Hub (repo_id + filename, forwarded to Llama.from_pretrained); n_gpu_layers and flash_attn, among other llama-cpp-python constructor options, are also exposed:

dtxt.backends.LlamaCpp(
    repo_id="TheBloke/some-model-GGUF",
    filename="some-model.Q4_K_M.gguf",
    n_ctx=8192,
    n_gpu_layers=32,
    flash_attn=True,
)

Anthropic uses forced tool use to get structured output; OpenAI uses response_format={"type": "json_schema", ...}; LlamaCpp constrains decoding at the grammar level via GBNF. None of them guarantee full schema conformance on their own:

  • Anthropic/OpenAI guarantee valid JSON syntax, not every schema keyword.
  • LlamaCpp strips constructs GBNF can't reliably express (format, pattern, deeply nested objects/arrays) from the grammar-facing schema; the original schema is still checked afterwards.

So all three go through dtxt's retry + validation loop the same way. StructuredEntityExtractor.extract_many runs concurrently via asyncio for Anthropic/OpenAI, bounded by max_concurrency (default 8) to avoid tripping rate limits; LlamaCpp processes it sequentially in-process so its prompt cache stays warm. A partial batch failure raises one ParseError naming how many texts failed and the first failing index, rather than aborting on the first error.

Style is controllable at both the schema and call level:

schema = Schema({
    "type": "object",
    "properties": {"name": {"type": "string"}},
    "required": ["name"],
    "x-dtxt-style": "formal, third person",  # schema-wide default
})
renderer = dtxt.StructuredEntityRenderer(backend, schema)
renderer.render(obj)                       # uses "formal, third person"
renderer.render(obj, style="casual, upbeat")  # overrides it for this call

Schema inference (SchemaInferer) works schema-free first: each text is reduced to a tree of (type, value) entities (optionally with children for repeating structured records, e.g. a receipt's line items) via dtxt.entities, type names are reconciled across the corpus, and a field is kept only if it meets a min_coverage threshold -- applied recursively, so nested object fields go through the same coverage test as top-level ones:

inferer = dtxt.SchemaInferer(backend, min_coverage=0.6)
schema = inferer.infer(texts)

Status

Released as dtxt 0.7.0 on PyPI; 0.8.0 (unreleased) reworks the public API to be class-based -- see CHANGELOG.md. Schema, StructuredEntityExtractor (T2D, with extract_many asyncio batching), StructuredEntityRenderer (D2T, with schema-level and per-call style control), SchemaInferer (schema-free extraction + recursive coverage-based merge, min_coverage), check_roundtrip, a mock backend for testing, and the Anthropic / OpenAI / llama.cpp backends are implemented. dtxt.entities (FlatEntityExtractor, NestedEntityExtractor, EntityTypeNormalizer, EntityRenderer) is SchemaInferer's internal implementation. See CHANGELOG.md for release notes and CLAUDE.md for what's next.

Development

uv sync --dev
uv run pytest
uv run ruff check . && uv run ruff format --check .
uv run mypy src/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dtxt-0.11.0.tar.gz (24.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dtxt-0.11.0-py3-none-any.whl (33.1 kB view details)

Uploaded Python 3

File details

Details for the file dtxt-0.11.0.tar.gz.

File metadata

  • Download URL: dtxt-0.11.0.tar.gz
  • Upload date:
  • Size: 24.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.20

File hashes

Hashes for dtxt-0.11.0.tar.gz
Algorithm Hash digest
SHA256 8f81af3f2e0f33f7d380af21b554670e000c678b7dc298594dd586b49ef054d8
MD5 3e93f3560c4b0ea142dfb85d08ccd3b5
BLAKE2b-256 a733c5f1592201ad1a067aef491450fd0c4e4786c8d6f30b99f5470fe140547e

See more details on using hashes here.

File details

Details for the file dtxt-0.11.0-py3-none-any.whl.

File metadata

  • Download URL: dtxt-0.11.0-py3-none-any.whl
  • Upload date:
  • Size: 33.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.20

File hashes

Hashes for dtxt-0.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3e8a8f6bac922bab14c0d2af33be27c24be3bd26964ed8d409e56875aa1c6fcb
MD5 788daf631195fc2cbdca4adca83738ef
BLAKE2b-256 e728ebe5e333543ea1247d2d899000d65867ef512495abeac1816e0e9f9bcbb6

See more details on using hashes here.

Release history Release notifications | RSS feed

0.21.0

2 files

0.20.0

2 files

0.19.1

2 files

0.19.0

2 files

0.18.0

2 files

0.17.0

2 files

0.16.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

This release

0.11.0 This release

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page