Stream Crystal Reports XML at memory bandwidth.
crxml
High-performance Crystal Reports XML → Arrow/DataFrame engine for Python.
Parse, filter, rename, cast, and project Crystal Reports XML directly into
columnar data, with Rust execution, parallel parsing, bounded-memory
processing, and automatic query fusion.
Quick start
from crxml import CrystalXMLSource
source = CrystalXMLSource("report.xml", row_tag="Details")
# Row iteration: yields dicts lazily
for row in source:
print(row["invoice"], row["amount"])
# DataFrame (auto-routes to parallel engine)
df = source.to_dataframe()
print(df.head())
That is it. df is a pandas DataFrame with zero-copy ArrowDtype strings,
built in under a second for a 100 MB file.
For maximum throughput, declare the schema upfront:
schema = ["Level","Section","Field22","Field23","Field38","Field39",
"Field61","Field73","FieldG","Text20"]
src = CrystalXMLSource("report.xml", row_tag="Details", schema=schema)
batches = src.iter_record_batches(memory="64MB", threads=16)
With pipeline stages fused into the Rust parse loop:
from crxml.stages import RenameFields, DropFields
pipeline = source | RenameFields({"f1": "invoice"}) | DropFields(["temp_id"])
df = pipeline.to_dataframe()
Performance tip: When your column set is known, pass
schema=[...]to skip column discovery and enable the fast path. On production data this is the single largest performance lever (crxml goes from 4.2 GB/s to 7.6 GB/s on a 533 MB report), and therow_satisfiedprojection skip reaches 11 GB/s on benchmarks.
Why crxml
This library was originally inspired by carlosplanchon/xmlstreamer.
Crystal Reports XML exports are deeply nested: <Group> wraps <GroupHeader>
wraps <Section> wraps <Details> wraps <Field>/<Text>/<FormattedValue>/
<Value>/<TextValue>. Standard XML libraries (ElementTree, SAX, lxml)
spend most of their CPU time descending into children you do not need.
crxml skips the nesting:
- The stream engine walks the XML once with a hand-rolled
memchrscanner (src/crxml_core/src/xml/scanner.rs,scan_one_rowscanner.rs:119viaRowSinksrc/crxml_core/src/lib.rs:603) and yields flat dicts: 508 MB/s 100 MB. - The parallel engine memory-maps the file, splits it at row boundaries (
splitter.rs:57find_split_points), and parses each chunk on its own thread into Arrow buffers directly (no dicts): up to 4.2 GB/s on high-cardinality production reports (533 MB realpar1284231) and 4.2 GB/s on uniform exports (1 GBpar1284158) via rypipe (rypipe-coreVec<ColumnBuilder>+field_indexengine.rs:16,row_dirtyengine.rs:26). - Pipeline stages that rename, cast, drop, or filter fields execute in the Rust parse loop, before any Python object is created.
Comparison: stream vs parallel vs parallel streaming (bounded, schema)
| Task | stream (single) | parallel (full RAM) | parallel streaming (bounded) |
|---|---|---|---|
| Row iteration | Yields dicts lazily | Arrow table first, then dicts (slower) | Yields RecordBatches incrementally, stable schema |
| DataFrame / Table output | Collects dicts, converts | Direct Arrow buffers, zero-copy | Same, incremental + bounded |
| 533 MB real export (Table) | 953 MB/s (single) / 723 MB/s (1 MB) | 4231 MB/s par128 (4.16 MB) |
3828 auto / 7630 explicit schema=[...] (2 MB) |
| 1 GB (Table) | 940 MB/s | 4158 MB/s par128 |
3782 auto / ~4900 explicit |
| Peak RssAnon (533 MB) | 24 MB (1 MB) | 137 MB | 88 MB (auto or explicit) |
| Pipeline fusion | No (dict path) | Yes (Rust BuildPlan) | Yes (same plan, streamed) |
ParquetWriter |
N/A | N/A | write_batch succeeds (batches share schema schema.rs:14) |
Auto discovery (16x2 MiB windows for >128 MB) adds ~15% (19 ms on 533 MB) so auto is -10% vs par128 (3828 vs 4231) but still bounded and incremental. Explicit schema=[...] (FrozenSchema::from_plan) avoids Discovery and is +80% vs par128 (7630 vs 4231). Fastest bounded mode needs explicit schema; auto is safe and bounded but slightly slower. Use iter_record_batches(memory="64MB", threads=16, schema=[...]) for the fast path.
Full benchmark details: like-for-like Table vs Vec, chunk-per-cell, fixed-chunk isolation, and schema cost.
Install
pip install crxml
The columnar and parallel engines are included by default. For performance
profiling counters: pip install -e . --config-settings=--features=profile.
Features
| Category | What crxml handles |
|---|---|
| Stream engine | Row-by-row XML parsing, yields dict[str, str], GIL-released batching |
| Columnar engine | Single-threaded Arrow table output, zero-copy string columns |
| Parallel engine | Multi-threaded (rayon), file split at row boundaries, off-GIL parse |
| Bounded mode | memory="500MB" splits into chunks; RSS independent of file size |
| Pipeline fusion | RenameFields, DropFields, CastTypes, FilterRows compile into Rust BuildPlan |
| mmap | Memory-maps input files (default, zero-copy) |
| prefault | MADV_WILLNEED vs MADV_SEQUENTIAL for RSS/speed trade-off |
| Arrow sinks | to_arrow(), to_pandas() (ArrowDtype), to_polars(), to_parquet() |
| Auto-dict encoding | auto_dict=True encodes low-cardinality string columns |
| Field typing | field_types={"amount": "float64"} coerces at parse time |
| Filter pushdown | filter={"field": "Status", "op": "==", "value": "Active"} in Rust |
| Correctness | All engines validated byte-identical against stream oracle (29 test cases + 465k-row real cross-check) |
Engine guide: parallel streaming (explicit schema) is opt-in
| Engine / API | When to use | Throughput 533 MB / 1 GB | RssAnon |
|---|---|---|---|
stream (for row in source) |
Row-by-row dict iteration | 723 MB/s 1 MB budget (24 MB anon) | 24 MB |
columnar (single) |
Single-threaded Arrow Table | 953 / 940 MB/s | 134 MB |
parallel (par128 full RAM, 4 MB) |
Fastest full-RAM Table | 4231 / 4158 MB/s | 137 MB |
iter_record_batches(..., threads=16, schema=[...]) (explicit schema) |
Fastest bounded, stable schema, yields RecordBatches |
7630 / — MB/s | 88 MB |
iter_record_batches(memory="64MB", threads=16) auto |
Bounded + incremental, stable schema | 3828 / 3782 MB/s (-14% vs par, +15% Discovery) | 88 MB |
bounded (memory="64MB" single) |
Single-thread bounded | 645 / 546 MB/s | 133 MB |
Pass engine= explicitly, or let auto select per call. auto stays "parallel if it fits" (blocked: auto discovery adds 15% and would make auto slower until cheaper). Streaming is opt-in via iter_record_batches(..., threads=16), keeping 4 MB for par (src/crxml/source.py:164), 2 MB via budget/(threads*2) for streaming. Provide schema= for the fast path.
# Recommended bounded paths
from crxml import CrystalXMLSource
import pyarrow as pa, pyarrow.parquet as pq
src = CrystalXMLSource("report.xml", row_tag="Details")
# explicit schema: fastest, no Discovery, writer succeeds
schema = ["Level","Section","Field22","Field23","Field38","Field39","Field61","Field73","FieldG","Text20"]
src = CrystalXMLSource("report.xml", row_tag="Details", schema=schema)
batches = src.iter_record_batches(memory="64MB", threads=16)
# auto: stable but pays 15% Discovery (16×2 MiB windows for >128 MB)
batches = src.iter_record_batches(memory="64MB", threads=16)
# ParquetWriter (now works; batches share schema)
it = src.iter_record_batches(memory="64MB", threads=16)
first = next(it)
w = pq.ParquetWriter("out.parquet", first.schema)
w.write_batch(first)
for b in it: w.write_batch(b)
w.close()
Framework support
| Framework | Integration |
|---|---|
| FastAPI / Starlette / Litestar | Parse in route handler, return DataFrame or Arrow table directly |
| Django / Flask | Call source.to_dataframe() in view; pass to template or response |
| Pandas / Polars | source.to_dataframe() / source.to_polars() for zero-copy analysis |
| Airflow / Prefect | Parse in task, write to parquet with source.to_parquet() |
| CLI / ETL scripts | Use to_csv() sink or iterate rows for line-by-line processing |
Limitations
- UTF-8 input only. UTF-16 exports (which Crystal Reports can produce) fail validation; convert first.
- No compressed input.
.gz/.zstfiles must be decompressed before parsing. - Crystal Reports grammar, not general XML. The flat-row model fits CR exports; arbitrary XML documents are out of scope.
- Linux-tuned performance. madvise hints and thread-count ratios were measured on Linux; other platforms work but are untested territory.
- No async API. Row iteration is synchronous.
Documentation
Full docs at crxml.emiliano-go.com covering:
- All
CrystalXMLSourceparameters - Pipeline stages and fusion rules
- Sink reference
- Batch iteration and parallel distribution
- Performance with phase breakdowns
- Architecture and correctness
License
MIT
Release files for crxml 0.3.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| crxml-0.3.1-cp313-cp313-manylinux_2_34_x86_64.whl | CPython 3.13 | CPython 3.13 | Linux glibc 2.34+ x86-64 | Details |
Release files / crxml-0.3.1-cp313-cp313-manylinux_2_34_x86_64.whl
| Download URL | crxml-0.3.1-cp313-cp313-manylinux_2_34_x86_64.whl |
|---|---|
| Size | 2.9 MB |
| Tags | CPython 3.13 Linux glibc 2.34+ x86-64 |
|
SHA-256 checksum How to use checksums |
0ee3375964c0c38daec34302746bdf94ba5c58ff7744e783db5bb4e50457c5c1
|
|
BLAKE2b-256 checksum How to use checksums |
df588d895cc22dcb39dd52cd7a9c7429d9730ae811348faea8aedb524999854b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency log