llmparse
Robustly extract, repair, and coerce structured data (primarily JSON) out of messy LLM text output.
Part of the ragkit suite. Install with
pip install ragkit-llmparse, thenimport llmparse.
LLMs love to wrap JSON in markdown fences, add a friendly sentence before and after it, sprinkle in trailing commas, use single quotes, emit Python literals (True/False/None), and forget to quote object keys. llmparse cleans all of that up and hands you a real Python object — optionally coerced to a schema you expect.
Pure standard library (json, re, ast). No dependencies. Python 3.8+.
Note: this is a best-effort, heuristic library. It is designed to recover data from almost-JSON. It is not a strict validator and it can be fooled by sufficiently pathological input.
Install
pip install ragkit-llmparse
Local development (from llmparse/):
pip install -e .
Quick Start
Parse a messy LLM reply — fences plus prose — straight into a dict:
import llmparse
reply = """
Sure! Here is the data you asked for:
```json
{
"name": "Ada Lovelace",
"born": 1815,
"fields": ["math", "computing"]
}
Hope this helps! """
data = llmparse.loads(reply) print(data["name"]) # Ada Lovelace print(data["fields"]) # ['math', 'computing']
## Handling each kind of mess
`loads` runs a pipeline: try raw `json.loads`, then extract the first balanced JSON region, then repair it, then fall back to `ast.literal_eval`.
### Trailing commas
```python
llmparse.loads('{"a": 1, "b": 2,}') # {'a': 1, 'b': 2}
llmparse.loads('[1, 2, 3,]') # [1, 2, 3]
Single quotes
llmparse.loads("{'name': 'Alice', 'age': 30}")
# {'name': 'Alice', 'age': 30}
Apostrophes inside double-quoted strings are left alone:
llmparse.loads('{"msg": "it\'s fine"}') # {'msg': "it's fine"}
Python literals
llmparse.loads('{"a": True, "b": False, "c": None}')
# {'a': True, 'b': False, 'c': None}
Unquoted keys
llmparse.loads('{name: "Bob", age: 25}')
# {'name': 'Bob', 'age': 25}
Python-dict-style output (ast fallback)
When the text is valid Python but not valid JSON (tuples, etc.), loads falls back to ast.literal_eval:
llmparse.loads("{'a': (1, 2), 'b': {'nested': True}}")
# {'a': (1, 2), 'b': {'nested': True}}
Extracting without parsing
extract_json returns the raw JSON substring(s). The brace scanner respects string literals and escapes, so a } inside a string will not cut the object short:
llmparse.extract_json('{"a": "text with } brace"}')
# '{"a": "text with } brace"}'
# All top-level objects/arrays:
llmparse.extract_json('First {"a": 1} then {"b": 2} and [3, 4].', first=False)
# ['{"a": 1}', '{"b": 2}', '[3, 4]']
repair_json gives you the fixed-up string if you want to inspect it:
llmparse.repair_json("{name: 'Al', active: True, tags: ['x', 'y',],}")
# '{"name": "Al", "active": true, "tags": ["x", "y"]}'
Repairing already-valid JSON returns an equivalent, still-parseable string.
Schema coercion
Describe the shape you expect and let llmparse cast values into it. A schema maps each field to either a bare type or a spec dict.
schema = {
"name": str,
"price": float,
"in_stock": bool,
"qty": int,
# spec dict form:
"discount": {"type": float, "required": False, "default": 0.0},
}
obj = {"name": "Widget", "price": "19.99", "in_stock": "true", "qty": "5"}
clean = llmparse.coerce(obj, schema)
# {'name': 'Widget', 'price': 19.99, 'in_stock': True, 'qty': 5, 'discount': 0.0}
Coercion rules (when coerce is on, which is the default per field):
"3"->int3"3.5"->float3.5"true","false",1,0->bool- numbers/bools ->
str
Spec dict options
| key | meaning | default |
|---|---|---|
type |
target type (int, float, str, bool, list, dict) |
— |
required |
whether the field must be present | True |
default |
value to fill if the field is missing | — |
coerce |
cast the value, or require an exact type match | True |
Aggregated errors
coerce collects every problem and raises a single SchemaError whose .errors list holds them all — so you see all missing/invalid fields at once:
schema = {"a": int, "b": str, "c": float}
try:
llmparse.coerce({"a": "oops"}, schema)
except llmparse.SchemaError as e:
for problem in e.errors:
print(problem)
# field 'a': cannot coerce 'oops' to int
# missing required field 'b'
# missing required field 'c'
Extra fields
Extra fields not mentioned in the schema are kept by default. Pass strict=True to drop them:
llmparse.coerce({"a": 1, "extra": 2}, {"a": int}) # {'a': 1, 'extra': 2}
llmparse.coerce({"a": 1, "extra": 2}, {"a": int}, strict=True) # {'a': 1}
One-shot: parse + coerce
parse is loads followed by coerce when a schema is supplied. Extra kwargs (repair, fallback_ast) pass through to loads.
reply = """Here you go:
```json
{name: 'Widget', 'price': '19.99', in_stock: True, qty: '5',}
"""
schema = {"name": str, "price": float, "in_stock": bool, "qty": int} llmparse.parse(reply, schema)
{'name': 'Widget', 'price': 19.99, 'in_stock': True, 'qty': 5}
## Multiple objects in one blob
`extract_all_json` parses every balanced JSON object/array it can find, skipping the ones that do not parse:
```python
text = 'First user {"id": 1} and second {"id": 2}. Also a list [10, 20].'
llmparse.extract_all_json(text)
# [{'id': 1}, {'id': 2}, [10, 20]]
API summary
loads(text, repair=True, fallback_ast=True)— main entrypoint; returns a Python object or raisesParseError.parse(text, schema=None, strict=False, **loads_kwargs)—loadsthen optionalcoerce.extract_json(text, first=True)— raw JSON substring(s).extract_all_json(text)— list of parsed objects.repair_json(s)— best-effort near-JSON -> JSON string.coerce(obj, schema, strict=False)— type coercion/validation.ParseError— has.snippet.SchemaError— has.errors(list).
License
MIT
Metadata
Release files for ragkit-llmparse 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ragkit_llmparse-0.1.1.tar.gz | 14.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ragkit_llmparse-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.7 kB
Release files / ragkit_llmparse-0.1.1.tar.gz
| Download URL | ragkit_llmparse-0.1.1.tar.gz |
|---|---|
| Size | 14.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3fbf9b02654f8fb374b9dd00919c8a723489af1852e4b5f31c9ec8de651b185f
|
|
BLAKE2b-256 checksum How to use checksums |
f1e47782bf474b53e9f638dee58d3c5719007300a20481f0de64b86240b1ae87
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.
Transparency logRelease files / ragkit_llmparse-0.1.1-py3-none-any.whl
| Download URL | ragkit_llmparse-0.1.1-py3-none-any.whl |
|---|---|
| Size | 10.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9cdd38d474cd7a581ba262761f3170a347a4ad7b6392c638995247a9803e28e3
|
|
BLAKE2b-256 checksum How to use checksums |
6fd0ed3d2719c0786ca6e36fd6ac1c4cd798c3db324aa0241e59557e3087f553
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.
Transparency log