dtxt
Schema-centric bidirectional conversion between text and structured data.
dtxt is not related to dtx (an AI red-teaming tool).
Core features
- Schema inference (
SchemaInferer): derive a schema from a collection of texts - T2D (
StructuredEntityExtractor): convert text into a schema-conformant object - D2T (
StructuredEntityRenderer): convert an object into text - Round-trip verification (
check_roundtrip): check thatextract(render(obj)) ≈ objfor a given schema and backend
Install
pip install dtxt # core only
pip install dtxt[anthropic] # + Anthropic backend
pip install dtxt[openai] # + OpenAI backend
pip install dtxt[llamacpp] # + local GGUF models via llama.cpp
pip install dtxt[all] # everything
The core package depends only on pydantic and jsonschema. Backends are
optional extras, imported lazily.
Usage
import dtxt
from dtxt import Schema
from dtxt.backends import MockBackend
schema = Schema({
"type": "object",
"properties": {
"name": {"type": "string", "x-dtxt-description": "the person's full name"},
"age": {"type": "integer"},
},
"required": ["name", "age"],
})
# Each class takes its backend at construction -- there is no global config.
extractor = dtxt.StructuredEntityExtractor(MockBackend(), schema)
renderer = dtxt.StructuredEntityRenderer(MockBackend(), schema)
obj = extractor.extract("Alice is 30 years old.")
text = renderer.render({"name": "Alice", "age": 30})
result = dtxt.check_roundtrip(
{"name": "Alice", "age": 30}, schema, renderer=renderer, extractor=extractor
)
result.ok # True if extract(render(obj)) == obj on every schema field
Swap MockBackend for a real one:
extractor = dtxt.StructuredEntityExtractor(dtxt.backends.LlamaCpp("model.gguf", n_ctx=8192), schema)
renderer = dtxt.StructuredEntityRenderer(dtxt.backends.Anthropic("claude-sonnet-4-6"), schema)
inferer = dtxt.SchemaInferer(dtxt.backends.Anthropic("claude-sonnet-4-6"))
LlamaCpp locates a model either by local path (model_path) or by
pulling from the Hugging Face Hub (repo_id + filename, forwarded to
Llama.from_pretrained); n_gpu_layers and flash_attn, among other
llama-cpp-python constructor options, are also exposed:
dtxt.backends.LlamaCpp(
repo_id="TheBloke/some-model-GGUF",
filename="some-model.Q4_K_M.gguf",
n_ctx=8192,
n_gpu_layers=32,
flash_attn=True,
)
Anthropic uses forced tool use to get structured output; OpenAI uses
response_format={"type": "json_schema", ...}; LlamaCpp constrains
decoding at the grammar level via GBNF. None of them guarantee full schema
conformance on their own:
- Anthropic/OpenAI guarantee valid JSON syntax, not every schema keyword.
LlamaCppstrips constructs GBNF can't reliably express (format,pattern, deeply nested objects/arrays) from the grammar-facing schema; the original schema is still checked afterwards.
So all three go through dtxt's retry + validation loop the same way.
StructuredEntityExtractor.extract_many runs concurrently via asyncio for
Anthropic/OpenAI, bounded by max_concurrency (default 8) to avoid
tripping rate limits; LlamaCpp processes it sequentially in-process so
its prompt cache stays warm. A partial batch failure raises one
ParseError naming how many texts failed and the first failing index,
rather than aborting on the first error.
Style is controllable at both the schema and call level:
schema = Schema({
"type": "object",
"properties": {"name": {"type": "string"}},
"required": ["name"],
"x-dtxt-style": "formal, third person", # schema-wide default
})
renderer = dtxt.StructuredEntityRenderer(backend, schema)
renderer.render(obj) # uses "formal, third person"
renderer.render(obj, style="casual, upbeat") # overrides it for this call
Schema inference (SchemaInferer) works schema-free first: each text is
reduced to a tree of (type, value) entities (optionally with children
for repeating structured records, e.g. a receipt's line items) via
dtxt.entities, type names are reconciled across the corpus, and a field
is kept only if it meets a min_coverage threshold -- applied recursively,
so nested object fields go through the same coverage test as top-level
ones:
inferer = dtxt.SchemaInferer(backend, min_coverage=0.6)
schema = inferer.infer(texts)
Status
Released as dtxt 0.7.0 on PyPI;
0.8.0 (unreleased) reworks the public API to be class-based -- see
CHANGELOG.md. Schema, StructuredEntityExtractor (T2D, with
extract_many asyncio batching), StructuredEntityRenderer (D2T, with
schema-level and per-call style control), SchemaInferer (schema-free
extraction + recursive coverage-based merge, min_coverage),
check_roundtrip, a mock backend for testing, and the Anthropic / OpenAI /
llama.cpp backends are implemented. dtxt.entities
(FlatEntityExtractor, NestedEntityExtractor, EntityTypeNormalizer,
EntityRenderer) is SchemaInferer's internal implementation. See
CHANGELOG.md for release notes and CLAUDE.md for what's next.
Development
uv sync --dev
uv run pytest
uv run ruff check . && uv run ruff format --check .
uv run mypy src/
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dtxt-0.10.0.tar.gz.
File metadata
- Download URL: dtxt-0.10.0.tar.gz
- Upload date:
- Size: 24.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.10.20
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4b9dbebdd9e5b67616499b23f0eee62aa8bb6b85b5ca615f98254af2b6a15b86
|
|
| MD5 |
0a97db50fde4a95d7a9b3e5879f9110a
|
|
| BLAKE2b-256 |
6d2f42760684de1d0e0ef04520b307f827ca3de77523971205c6dcc184bfc965
|
File details
Details for the file dtxt-0.10.0-py3-none-any.whl.
File metadata
- Download URL: dtxt-0.10.0-py3-none-any.whl
- Upload date:
- Size: 32.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.10.20
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1eaeea29ff530c2b0ea078169bf2ffee2bb7004f8539209235f2f0b55b4a8286
|
|
| MD5 |
764ac964a296472a5b653d1bc6936997
|
|
| BLAKE2b-256 |
556bba91ae77dacec359f8948b80b6cdb01dbb3b539eab4f451623fdb7851d65
|