Skip to main content

Amara is a general-purpose web data processing library with IRI handling and MicroXML/XML processing

Features

  • IRI (Internationalized Resource Identifier) processing - Complete implementation for handling IRIs, including percent encoding/decoding, joining, splitting, and validation
  • MicroXML/XML parsing and processing - Simplified XML data model based on MicroXML, with support for full XML 1.0
  • HTML5 parsing - Parse HTML5 documents with modern html5lib-modern
  • XPath-like queries - MicroXPath support for querying XML documents
  • Command-line tool - microx for rapid XML/MicroXML processing and extraction

PyPI - Version PyPI - Python Version

Installation

Requires Python 3.12 or later.

pip install amara

Or with uv (recommended):

uv pip install amara

You can also install directly from the latest source version:

git clone https://github.com/OoriData/Amara.git
cd Amara
pip install -U .

Quick Start

IRI Processing

from amara.iri import I, iri

# Create and manipulate IRIs
url = I('http://example.org/path/to/resource')
print(url.scheme)  # 'http'
print(url.host)    # 'example.org'

# Join relative paths with base URLs
joined = iri.join('http://example.org/a/b', '../c')
print(joined)  # 'http://example.org/a/c'

# Percent encoding/decoding
encoded = iri.percent_encode('hello world!')
print(encoded)  # 'hello%20world%21'

XML Processing

from amara.uxml import parse

SAMPLE_XML = '''<monty>
  <python spam="eggs">What do you mean "bleh"</python>
  <python ministry="abuse">But I was looking for argument</python>
</monty>'''

# Parse XML
root = parse(SAMPLE_XML)
print(root.xml_name)  # "monty"

# Access children and attributes
for child in root.xml_children:
    if hasattr(child, 'xml_attributes'):
        print(f'Element: {child.xml_name}')
        print(f'Spam attr: {child.xml_attributes.get('spam')}')
        print(f'Text: {child.xml_value}')

# Iterate through all elements
for elem in root.xml_descendants():
    print(f'Found element: {elem.xml_name}')

"MicroXML?" What's that?

MicroXML is a W3C Community Project and spec. A lot of XML veterans, including Uche, Amara's founder, had become fed up with the levels of unnecessary complexity in the XML stack, including XML Namespaces, which charges a huge technical cost in order to solve an overstated problem. Amara implements the MicroXML data model, and allows you to parse into this from tradiional XML and the MicroXML serialization.

In reality, most of the XML-like data you’ll be dealing with is full XML 1.0, so Amara package provides capabilities to parse legacy XML and reduce it to MicroXML. In many cases the biggest implication of this is that namespace information is stripped. You can get very far by just ignoring this, and it opens up the much simpler processing encouraged by MicroXML.

HTML5 Processing

from amara.uxml import html5

HTML_DOC = '''<!DOCTYPE html>
<html>
  <head><title>Example</title></head>
  <body><p class="plain">Hello World</p></body>
</html>'''

doc = html5.parse(HTML_DOC)
print(doc.xml_name)  # "html"

XPath-like Queries (MicroXPath)

from amara.uxml import parse

SAMPLE_XML = '''<catalog>
  <book id="1">
    <title>Python Programming</title>
    <author>John Doe</author>
  </book>
  <book id="2">
    <title>Web Development</title>
    <author>Jane Smith</author>
  </book>
</catalog>'''

root = parse(SAMPLE_XML)

# Find all book titles
titles = root.xml_xpath('//book/title')
for title in titles:
    print(title.xml_value)

# Find book by ID
book = list(root.xml_xpath("//book[@id='2']"))
if book:
    # First child is whitespace. 2nd is the "title" element
    print(f'Found: {book[0].xml_children[1].xml_value}')

Command-Line Tool

The microx command provides powerful XML/MicroXML querying and processing:

# Extract elements by name
microx file.xml --match=item

# XPath-like expressions
microx file.xml --expr="//item[@id='2']"

# Extract text content from specific elements
microx file.xml --match=name --foreach="text()"

# Process multiple files
microx *.xml --match=title --foreach="text()"

# Pretty-print XML
microx file.xml --pretty

# Convert to MicroXML
microx file.xml --microxml

For more options, run:

microx --help

Requirements

Development

Amara is primarily developed by the crew at Oori Data. We offer LLMOps, data pipelines and software engineering services around AI/LLM applications.

History

Amara was originally an open source project I created, renaming and expanding on Anobind 2003, looking to simplify and rethink XML and related technology processing, with an eye to Python. It went through a few evolutions and progress had slowed down since the late 2010s.

Quote from the revival ticket:

The Amara saga continues! I don't exactly remember why I decided to dead end the Amara PyPI project when it hit 2.0, but I moved to a series of Amara 3 generation projects (amara3.iri, amara3.xml & amara3-names). Those were far more lone wolf efforts, but at Oori Data we're seeing a lot of need for the sorts of capability that's inchoate in Amara 3.

Metadata

Release files for Amara 4.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for Amara 4.1.0
File Size Uploaded
amara-4.1.0.tar.gz 96.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for Amara 4.1.0
File Interpreter ABI Platform
amara-4.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 184.8 kB

Release files / amara-4.1.0.tar.gz

Download URL amara-4.1.0.tar.gz
Size 96.6 kB
Tags Source
SHA-256 checksum
How to use checksums
75d176d09090f03de0b48c43808321ce44896f8fb267350fd62714dd2c661b78
BLAKE2b-256 checksum
How to use checksums
3a1e9325e619b3087d35765d24edee3bb57e9f42a2f82f18cda19e44338f836c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 1, 2026.

Transparency log

Release files / amara-4.1.0-py3-none-any.whl

Download URL amara-4.1.0-py3-none-any.whl
Size 88.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
10153f68428c581a8e8e394648655da116671ff1efb4ae377400e9cdab4ec0c2
BLAKE2b-256 checksum
How to use checksums
5cfb4a5e2b1017eb1b90442f980ab12abcbebb99814276b56db0ccaaa710c59f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

4.1.0 This release

2 release files

4.0.2

2 release files

4.0.1

2 release files

2.0.0

1 release file

1.2

4 release files

1.1.9

6 release files

1.1.6

1.1.5

1.0

1.0b3

1.0b2

1.0b1

0.9.4

0.9.3

0.9.2

0.9.1

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page