Skip to main content

LLM Output Parser

PyPI version GitHub stars codecov Build Status

A robust utility for extracting and parsing structured data (JSON and XML) from unstructured text outputs generated by Large Language Models (LLMs).

Features

  • Extracts JSON and XML from plain text, code blocks, and mixed content
  • Handles various JSON formats: objects, arrays, and nested structures
  • Converts XML to JSON-compatible dictionary format
  • Advanced extraction strategies for multiple JSON/XML objects in text
  • Provides robust error handling and recovery strategies
  • Works with markdown code blocks (json ... and xml ... )
  • Intelligently selects the most comprehensive structure when multiple are found

Installation

Install from PyPI:

pip install llm-output-parser

Or install from source:

git clone https://github.com/KameniAlexNea/llm-output-parser.git
cd llm-output-parser
pip install -e .

Usage

JSON Parsing

from llm_output_parser import parse_json

# Parse JSON from an LLM response
llm_response = """
Here's the data you requested:


{
    "name": "John Doe",
    "age": 30,
    "skills": ["Python", "Machine Learning", "NLP"]
}


Let me know if you need anything else!
"""

data = parse_json(llm_response)
print(data["name"])  # John Doe
print(data["skills"])  # ['Python', 'Machine Learning', 'NLP']

XML Parsing

from llm_output_parser import parse_xml

# Parse XML from an LLM response and convert to JSON
llm_response = """
Here's the user data in XML format:

```xml
<user id="123">
    <name>Jane Smith</name>
    <email>jane@example.com</email>
    <roles>
        <role>admin</role>
        <role>editor</role>
    </roles>
</user>

Let me know if you need any other information. """

data = parse_xml(llm_response) print(data["@id"]) # 123 print(data["name"]) # Jane Smith print(data["roles"]["role"]) # ['admin', 'editor']


### Handling Complex Cases

The library can handle various complex scenarios:

#### JSON Within Text

```python
text = 'The user profile is: {"name": "John", "email": "john@example.com"}'
data = parse_json(text)  # -> {"name": "John", "email": "john@example.com"}

XML Within Text

text = 'The configuration is: <config><server>localhost</server><port>8080</port></config>'
data = parse_xml(text)  # -> {"server": "localhost", "port": "8080"}

Multiple JSON/XML Objects

When multiple valid objects are present, the parser returns the most comprehensive one:

# For JSON
text = '''
Small object: {"id": 123}

Larger object:
{
    "user": {
        "id": 123,
        "name": "John",
        "email": "john@example.com",
        "preferences": {
            "theme": "dark",
            "notifications": true
        }
    }
}
'''
data = parse_json(text)  # Returns the larger, more complex object

# For XML
text = '''
Simple: <item>value</item>

Complex:
<product category="electronics">
    <name>Smartphone</name>
    <price currency="USD">999.99</price>
    <features>
        <feature>5G</feature>
        <feature>High-res camera</feature>
    </features>
</product>
'''
data = parse_xml(text)  # Returns the more complex XML converted to JSON

XML to JSON Conversion Details

When parsing XML, the library converts it to a JSON-compatible dictionary with the following conventions:

  • XML attributes are prefixed with @ (e.g., <item id="123"> becomes {"@id": "123"})
  • Text content of elements with attributes or children is stored under #text key
  • Simple elements with only text become key-value pairs
  • Repeated elements are automatically converted to arrays

Example:

xml_str = '''
<library>
    <book category="fiction">
        <title>The Great Gatsby</title>
        <author>F. Scott Fitzgerald</author>
    </book>
    <book category="non-fiction">
        <title>Sapiens</title>
        <author>Yuval Noah Harari</author>
    </book>
</library>
'''
data = parse_xml(xml_str)
# Results in:
# {
#     "book": [
#         {
#             "@category": "fiction",
#             "title": "The Great Gatsby",
#             "author": "F. Scott Fitzgerald"
#         },
#         {
#             "@category": "non-fiction",
#             "title": "Sapiens",
#             "author": "Yuval Noah Harari"
#         }
#     ]
# }

Error Handling

If no valid structure can be found, a ValueError is raised:

try:
    data = parse_json("No JSON here!")
except ValueError as e:
    print(f"Error: {e}")  # "Error: Failed to parse JSON from the input string."

try:
    data = parse_xml("No XML here!")
except ValueError as e:
    print(f"Error: {e}")  # "Error: Failed to parse XML from the input string."

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add some amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Metadata

Release files for llm-output-parser 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-output-parser 0.3.0
File Size Uploaded
llm_output_parser-0.3.0.tar.gz 14.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-output-parser 0.3.0
File Interpreter ABI Platform
llm_output_parser-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 29.0 kB

Release files / llm_output_parser-0.3.0.tar.gz

Download URL llm_output_parser-0.3.0.tar.gz
Size 14.5 kB
Tags Source
SHA-256 checksum
How to use checksums
2bf5c20b0da6460d4c7860c8a48d47d527bafd07f9af4f4fb9b1ff47d3c70c1f
BLAKE2b-256 checksum
How to use checksums
9efd3517db603e1fc124ce4a59cab568a6433cfba3733cafd9070bdf3334b32e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 19, 2025.

Transparency log

Release files / llm_output_parser-0.3.0-py3-none-any.whl

Download URL llm_output_parser-0.3.0-py3-none-any.whl
Size 14.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b9d7c9cd469205aa698efb60fffad591f943e806858f9f77c60139295ad96ff9
BLAKE2b-256 checksum
How to use checksums
c26c888e9db804503ef31567b7b20bbe19263ec85c67c941a7beda2ac2797ae4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 19, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page