LLM Output Parser
A robust utility for extracting and parsing structured data (JSON and XML) from unstructured text outputs generated by Large Language Models (LLMs).
Features
- Extracts JSON and XML from plain text, code blocks, and mixed content
- Handles various JSON formats: objects, arrays, and nested structures
- Converts XML to JSON-compatible dictionary format
- Advanced extraction strategies for multiple JSON/XML objects in text
- Provides robust error handling and recovery strategies
- Works with markdown code blocks (
json ...andxml ...) - Intelligently selects the most comprehensive structure when multiple are found
Installation
Install from PyPI:
pip install llm-output-parser
Or install from source:
git clone https://github.com/KameniAlexNea/llm-output-parser.git
cd llm-output-parser
pip install -e .
Usage
JSON Parsing
from llm_output_parser import parse_json
# Parse JSON from an LLM response
llm_response = """
Here's the data you requested:
{
"name": "John Doe",
"age": 30,
"skills": ["Python", "Machine Learning", "NLP"]
}
Let me know if you need anything else!
"""
data = parse_json(llm_response)
print(data["name"]) # John Doe
print(data["skills"]) # ['Python', 'Machine Learning', 'NLP']
XML Parsing
from llm_output_parser import parse_xml
# Parse XML from an LLM response and convert to JSON
llm_response = """
Here's the user data in XML format:
```xml
<user id="123">
<name>Jane Smith</name>
<email>jane@example.com</email>
<roles>
<role>admin</role>
<role>editor</role>
</roles>
</user>
Let me know if you need any other information. """
data = parse_xml(llm_response) print(data["@id"]) # 123 print(data["name"]) # Jane Smith print(data["roles"]["role"]) # ['admin', 'editor']
### Handling Complex Cases
The library can handle various complex scenarios:
#### JSON Within Text
```python
text = 'The user profile is: {"name": "John", "email": "john@example.com"}'
data = parse_json(text) # -> {"name": "John", "email": "john@example.com"}
XML Within Text
text = 'The configuration is: <config><server>localhost</server><port>8080</port></config>'
data = parse_xml(text) # -> {"server": "localhost", "port": "8080"}
Multiple JSON/XML Objects
When multiple valid objects are present, the parser returns the most comprehensive one:
# For JSON
text = '''
Small object: {"id": 123}
Larger object:
{
"user": {
"id": 123,
"name": "John",
"email": "john@example.com",
"preferences": {
"theme": "dark",
"notifications": true
}
}
}
'''
data = parse_json(text) # Returns the larger, more complex object
# For XML
text = '''
Simple: <item>value</item>
Complex:
<product category="electronics">
<name>Smartphone</name>
<price currency="USD">999.99</price>
<features>
<feature>5G</feature>
<feature>High-res camera</feature>
</features>
</product>
'''
data = parse_xml(text) # Returns the more complex XML converted to JSON
XML to JSON Conversion Details
When parsing XML, the library converts it to a JSON-compatible dictionary with the following conventions:
- XML attributes are prefixed with
@(e.g.,<item id="123">becomes{"@id": "123"}) - Text content of elements with attributes or children is stored under
#textkey - Simple elements with only text become key-value pairs
- Repeated elements are automatically converted to arrays
Example:
xml_str = '''
<library>
<book category="fiction">
<title>The Great Gatsby</title>
<author>F. Scott Fitzgerald</author>
</book>
<book category="non-fiction">
<title>Sapiens</title>
<author>Yuval Noah Harari</author>
</book>
</library>
'''
data = parse_xml(xml_str)
# Results in:
# {
# "book": [
# {
# "@category": "fiction",
# "title": "The Great Gatsby",
# "author": "F. Scott Fitzgerald"
# },
# {
# "@category": "non-fiction",
# "title": "Sapiens",
# "author": "Yuval Noah Harari"
# }
# ]
# }
Error Handling
If no valid structure can be found, a ValueError is raised:
try:
data = parse_json("No JSON here!")
except ValueError as e:
print(f"Error: {e}") # "Error: Failed to parse JSON from the input string."
try:
data = parse_xml("No XML here!")
except ValueError as e:
print(f"Error: {e}") # "Error: Failed to parse XML from the input string."
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
- Fork the repository
- Create your feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add some amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Metadata
Release files for llm-output-parser 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_output_parser-0.3.0.tar.gz | 14.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_output_parser-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 29.0 kB
Release files / llm_output_parser-0.3.0.tar.gz
| Download URL | llm_output_parser-0.3.0.tar.gz |
|---|---|
| Size | 14.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2bf5c20b0da6460d4c7860c8a48d47d527bafd07f9af4f4fb9b1ff47d3c70c1f
|
|
BLAKE2b-256 checksum How to use checksums |
9efd3517db603e1fc124ce4a59cab568a6433cfba3733cafd9070bdf3334b32e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 19, 2025.
Transparency logRelease files / llm_output_parser-0.3.0-py3-none-any.whl
| Download URL | llm_output_parser-0.3.0-py3-none-any.whl |
|---|---|
| Size | 14.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b9d7c9cd469205aa698efb60fffad591f943e806858f9f77c60139295ad96ff9
|
|
BLAKE2b-256 checksum How to use checksums |
c26c888e9db804503ef31567b7b20bbe19263ec85c67c941a7beda2ac2797ae4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 19, 2025.
Transparency log