Skip to main content

crxml

Stream Crystal Reports XML at memory bandwidth.

crxml

High-performance Crystal Reports XML → Arrow/DataFrame engine for Python.

Parse, filter, rename, cast, and project Crystal Reports XML directly into
columnar data, with Rust execution, parallel parsing, bounded-memory
processing, and automatic query fusion.

Python License Tests PyPI Docs


Quick start

from crxml import CrystalXMLSource

source = CrystalXMLSource("report.xml", row_tag="Details")

# Row iteration: yields dicts lazily
for row in source:
    print(row["invoice"], row["amount"])

# DataFrame (auto-routes to parallel engine)
df = source.to_dataframe()
print(df.head())

That is it. df is a pandas DataFrame with zero-copy ArrowDtype strings, built in under a second for a 100 MB file.

With pipeline stages fused into the Rust parse loop:

from crxml.stages import RenameFields, DropFields

pipeline = source | RenameFields({"f1": "invoice"}) | DropFields(["temp_id"])
df = pipeline.to_dataframe()

Why crxml

Crystal Reports XML exports are deeply nested: <Group> wraps <GroupHeader> wraps <Section> wraps <Details> wraps <Field>/<Text>/<FormattedValue>/ <Value>/<TextValue>. Standard XML libraries (ElementTree, SAX, lxml) spend most of their CPU time descending into children you do not need.

crxml skips the nesting:

  • The stream engine walks the XML once with quick-xml and yields flat dicts.
  • The parallel engine memory-maps the file, splits it at row boundaries, and parses each chunk on its own thread into Arrow buffers directly (no dicts).
  • Pipeline stages that rename, cast, drop, or filter fields execute in the Rust parse loop, before any Python object is created.

Comparison: stream vs parallel

Task stream parallel
Row iteration Yields dicts lazily Arrow table first, then dicts (slower)
DataFrame output Collects dicts, converts Direct Arrow buffers, zero-copy
100 MB synthetic 2.3 s 0.21 s
533 MB real export 12.8 s 1.13 s
Peak RSS ~1.07 GB ~534 MB (file size, mmap)
Pipeline fusion No (dict path) Yes (Rust BuildPlan)

The stream engine materializes one dict per row: fully consuming a large file costs roughly 10x its size in memory (~1 GB RSS for a 100 MB file). Use it for incremental processing; use table sinks for collection.

For files larger than RAM, add memory="500MB" to any engine for bounded mode: peak RSS tracks the budget, not the file.

Full benchmark details


Install

pip install crxml

The columnar and parallel engines are included by default. For performance profiling counters: pip install -e . --config-settings=--features=profile.


Features

Category What crxml handles
Stream engine Row-by-row XML parsing, yields dict[str, str], GIL-released batching
Columnar engine Single-threaded Arrow table output, zero-copy string columns
Parallel engine Multi-threaded (rayon), file split at row boundaries, off-GIL parse
Bounded mode memory="500MB" splits into chunks; RSS independent of file size
Pipeline fusion RenameFields, DropFields, CastTypes, FilterRows compile into Rust BuildPlan
mmap Memory-maps input files (default, zero-copy)
prefault MADV_WILLNEED vs MADV_SEQUENTIAL for RSS/speed trade-off
Arrow sinks to_arrow(), to_pandas() (ArrowDtype), to_polars(), to_parquet()
Auto-dict encoding auto_dict=True encodes low-cardinality string columns
Field typing field_types={"amount": "float64"} coerces at parse time
Filter pushdown filter={"field": "Status", "op": "==", "value": "Active"} in Rust
Correctness All engines validated byte-identical against stream oracle (29 test cases + 465k-row real cross-check)

Engine guide

Engine When to use
stream Row-by-row iteration (for row in source)
columnar Single-threaded Arrow output
parallel Fastest DataFrame output (default for files > 8 MB)
bounded Files larger than RAM (memory="500MB" with any engine)

Pass engine= explicitly, or let auto select the best engine per call.


Framework support

Framework Integration
FastAPI / Starlette / Litestar Parse in route handler, return DataFrame or Arrow table directly
Django / Flask Call source.to_dataframe() in view; pass to template or response
Pandas / Polars source.to_dataframe() / source.to_polars() for zero-copy analysis
Airflow / Prefect Parse in task, write to parquet with source.to_parquet()
CLI / ETL scripts Use to_csv() sink or iterate rows for line-by-line processing

Limitations

  • UTF-8 input only. UTF-16 exports (which Crystal Reports can produce) fail validation; convert first.
  • No compressed input. .gz/.zst files must be decompressed before parsing.
  • Crystal Reports grammar, not general XML. The flat-row model fits CR exports; arbitrary XML documents are out of scope.
  • Linux-tuned performance. madvise hints and thread-count ratios were measured on Linux; other platforms work but are untested territory.
  • No async API. Row iteration is synchronous.

Documentation

Full docs at crxml.emiliano-go.com covering:

  • All CrystalXMLSource parameters
  • Pipeline stages and fusion rules
  • Sink reference
  • Batch iteration and parallel distribution
  • Performance with phase breakdowns
  • Architecture and correctness

License

MIT

Release files for crxml 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for crxml 1.1.0
File Interpreter ABI Platform
crxml-1.1.0-cp313-cp313-manylinux_2_34_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.34+ x86-64 Details

Release files / crxml-1.1.0-cp313-cp313-manylinux_2_34_x86_64.whl

Download URL crxml-1.1.0-cp313-cp313-manylinux_2_34_x86_64.whl
Size 742.5 kB
Tags CPython 3.13 Linux glibc 2.34+ x86-64
SHA-256 checksum
How to use checksums
884aa1fe0c96c82a5fd5b448544b2284d589b8a53614f45606df98985a56c9c9
BLAKE2b-256 checksum
How to use checksums
370f6a7db26db28ac57746089e1561ad55c8eb339479ad9e5b9d13b352a27000
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.

Transparency log

Release history Release notifications | RSS feed

2.1.0

1 release file

1.2.0

1 release file

This release

1.1.0 This release

1 release file

1.0.0

1 release file

0.3.1

1 release file

0.3.0

1 release file

0.2.0

1 release file

0.1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page