Skip to main content

StreamXL

Problem

openpyxl and similar pure-Python Excel readers load the whole workbook into memory before you can touch a single row — fine for small files, a real ceiling for ETL pipelines and data engineering workloads working against large .xlsx exports.

Solution

Stream large .xlsx files row-by-row in constant memory, powered by a Rust core with no unsafe code.

pip installs as streamxl, import streamxl. Read multi-sheet Excel workbooks without loading them fully into memory, extract formulas and comments, write new .xlsx files, and append to existing ones — all through a small, plain Python API backed by a Rust engine.

PyPI CI Python 3.8+ License: Apache 2.0

Use cases

  • ETL against large Excel exports that don't fit comfortably in memory with openpyxl — read() keeps memory flat regardless of file size.
  • Extracting formulas/comments for audit or migration tooling, not just cell values.
  • Appending to a growing log-style .xlsx file without rewriting the whole workbook or losing other sheets.
  • Not yet a good fit for: SQL-style querying across sheets, formula evaluation (only extraction/classification), or pandas/Parquet/Arrow export built in — see Honest feature list for the full "what's not here" list.

Install

pip install streamxl

A prebuilt wheel is currently published only for macOS (arm64); other platforms install from the source distribution, which requires a Rust toolchain (see rust-toolchain.toml) and maturin to build. Every PyPI release to date (1.2.0 through 5.2.0) has shipped exactly one platform wheel plus an sdist — no Linux or Windows wheels have been published yet.

Quick start

import streamxl

for row in streamxl.read("data.xlsx"):
    print(row)  # ['Name', 'Age', 'Score']

read() streams rows one at a time — memory use stays flat regardless of file size.

Real, working examples

Read as dictionaries, keyed by header row:

import streamxl

for row in streamxl.read("sales.xlsx", as_dict=True):
    print(row["Customer"], row["Amount"])

Read only specific columns:

for row in streamxl.read("sales.xlsx", as_dict=True, columns=["Customer", "Amount"]):
    ...

Read every sheet in a workbook:

sheet_names = streamxl.sheets("workbook.xlsx")
all_data = streamxl.read_all("workbook.xlsx")  # {sheet_name: [rows...]}

Write a new .xlsx file:

import datetime
import streamxl

streamxl.write("report.xlsx", [
    ["Name", "Joined", "Score"],
    ["Alice", datetime.date(2024, 1, 15), 95.5],
    ["Bob", datetime.date(2024, 3, 2), 88.0],
])

Stream-write multiple sheets without holding the whole file in memory:

with streamxl.writer("report.xlsx") as w:
    w.write_row(["Name", "Age"])
    w.write_row(["Alice", 30])
    w.add_sheet("Summary")
    w.write_row(["Total", 1])

Append rows to an existing file (other sheets are preserved):

streamxl.write("log.xlsx", [["Date", "Event"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "started"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "finished"]])

Extract formulas and comments:

rows = list(streamxl.read("model.xlsx", with_formulas=True))
# each cell is a dict: {"value": ..., "formula": ..., "formula_type": ...,
#                        "comment": ..., "comment_author": ...}

from streamxl import FormulaSerializer
export = FormulaSerializer.export_formulas(rows)
FormulaSerializer.export_to_json(rows, "formulas.json")
FormulaSerializer.export_to_csv(rows, "formulas.csv")  # sanitized against CSV/formula injection

Export to CSV safely — untrusted cell content is never written to CSV verbatim (see Security below):

import csv
import streamxl
from streamxl.security import sanitize_csv_cell

with open("output.csv", "w", newline="") as f:
    writer = csv.writer(f)
    for row in streamxl.read("large.xlsx"):
        writer.writerow([sanitize_csv_cell(cell) for cell in row])

Validate a file and recover from bad cells instead of crashing:

from streamxl import validate_excel_file

report = validate_excel_file("questionable.xlsx")
if report.has_fatal_errors():
    print(report.format_summary())

More runnable examples live in examples/.

Honest feature list

What's here and real, backed by the Rust core and covered by the test suite:

  • Streaming reads — read() / stream(): a real Rust __iter__/__next__ iterator over the sheet — you get rows one at a time, not a pre-built Python list. Correction (2026-09-22 benchmark): this is not O(1) memory as previously claimed here. XlsxStream::open() decompresses the whole sheet XML into memory before iteration starts, so peak RSS scales with sheet size (measured ~1.3MB of RSS per 1MB of .xlsx, real data — see benchmarks/results.md). It's still far below openpyxl's full-load mode and the API shape (a lazy iterator) is real, but memory is not flat regardless of file size — tracked as gap #10 in ROADMAP_HONEST.md. read_rows_all_at_once()/read_rows_with_metadata_all_at_once() remain available as an explicit escape hatch for callers that need random access or to iterate the result more than once.
  • Multi-sheet support — sheets(), read_all(), and writer().add_sheet().
  • Streaming writes — write(), writer(), append(), all producing real .xlsx files.
  • Formula extraction — read formula text and a best-effort formula-type classification (with_formulas=True), plus FormulaReferenceMapper for shifting/rewriting cell references and FormulaSerializer for exporting/importing formulas as JSON or CSV.
  • Comment extraction — cell comments and authors, via with_formulas=True.
  • Conditional formatting rules — conditional_formats() reads every <conditionalFormatting>/<cfRule> in a sheet (type, operator, formulas, priority, stopIfTrue) and resolves each rule's dxfId against xl/styles.xml's <dxfs> into concrete font color/bold/italic and fill colors. colorScale/dataBar/iconSet rules are captured (type, sqref, priority) but their inline color-stop/threshold definitions aren't modeled — those rule types don't use dxfId in the first place.
  • Type-aware cells — strings, numbers, booleans, dates, datetimes, and empty cells round-trip correctly.
  • Error recovery & validation — validate_excel_file() and ErrorRecoveryHandler classify and (optionally) recover from malformed cells instead of hard-failing on the whole file.
  • Security hardening — path validation, file-size limits, and ZIP-bomb defenses (entry-size, compression-ratio, and total-decompressed-size limits) enforced before/while a file is opened. CSV export is sanitized against formula-injection (see below).
  • REST API (optional) — streamxl.server.StreamXLServer / create_flask_app() wrap the real streaming engine behind HTTP endpoints (/sources, /sources/<id>/query, /sources/<id>/export, ...). Requires pip install "streamxl[server]".

What's not here, so you don't have to find out the hard way:

  • No SQL-style query language — execute_query() in the REST API streams rows from a named sheet, it does not parse arbitrary queries.
  • No pandas/Parquet/Arrow export built in. Convert read()'s output yourself, or open an issue if this matters to you.
  • No formula evaluation — formula text is extracted and classified, not recalculated.
  • The pystreamxl dashboard CLI command renders sample data, not live telemetry — every mode (bare, --static, --alerts, --recommendations, --export) shows the same explicit "SAMPLE DATA — not live" warning.

Security

  • Path & size validation — validate_read_path() / validate_write_path() reject non-.xlsx paths, path traversal, and oversized files before any parsing happens.
  • ZIP-bomb defenses — the Rust core enforces a per-entry size limit, a compression-ratio limit, and a total-decompressed-size limit while unpacking a workbook (see core/src/zip_reader.rs), tested against real crafted archives in core/tests/zip_bomb_defense.rs.
  • CSV/formula-injection protection — streamxl.security.sanitize_csv_cell() neutralizes any string cell that starts with =, +, -, @, TAB, or CR (the standard CSV-injection trigger set) by prefixing it with ', so a malicious workbook can't turn a CSV export into an executable formula when reopened in Excel/LibreOffice/Google Sheets. FormulaSerializer.export_to_csv() applies this automatically; apply it yourself when writing CSV from read() output (see the example above).

Limits, enforced by default (no configuration needed):

Limit Value
Max file size 512 MB
Max size per ZIP entry 512 MB
Max total decompressed size 1 GB
Max compression ratio 30:1

Handle malformed or malicious files by catching SecurityError:

from streamxl import SecurityError, read

try:
    for row in read("data.xlsx"):
        process(row)
except SecurityError as e:
    print(f"Security violation: {e}")

Found a security issue? See SECURITY.md.

Performance

Rows are parsed and yielded one at a time rather than being collected into a Python list up front, and it's consistently faster than openpyxl on both reads and writes. Memory use is lower than openpyxl's full-load mode but currently scales with sheet size rather than staying flat — see the correction under "Honest feature list" above and ROADMAP_HONEST.md gap #10. See benchmarks/ for the scripts used to compare against openpyxl, and examples/memory_benchmark.py to measure it yourself against your own files:

python examples/memory_benchmark.py your_file.xlsx

Actual numbers depend heavily on your file's structure (shared strings, formulas, formatting) — measure on your own workloads rather than trusting a generic table.

vs openpyxl, on real data

Methodology: 150,000 real, live NYC 311 Service Request rows pulled from NYC Open Data's Socrata API (data.cityofnewyork.us/resource/erm2-nwe9, current as of 2026-09-22 — not synthetic/fabricated rows), 14 columns, written to a real 2-sheet .xlsx workbook (75k rows/sheet, 19MB) via openpyxl. Both libraries iterated every row of every sheet; row counts and a positional checksum matched exactly across all three methods (correctness verified, not just speed). 3 runs each, median reported, single-process wall-clock via time.perf_counter(), peak RSS via resource.getrusage(...).ru_maxrss on macOS/arm64, Python 3.13.

Rows streamxl read() openpyxl read_only=True openpyxl full load
10,000 0.07s · 28MB peak RSS 0.61s · 31MB peak RSS —
30,000 0.19s · 52MB peak RSS 1.89s · 32MB peak RSS —
75,000 0.48s · 103MB peak RSS 4.72s · 36MB peak RSS —
150,000 (2 sheets) 0.96s · 192MB peak RSS 9.27s · 43MB peak RSS 13.5s · 1,040MB peak RSS

streamxl is ~9.7x faster than openpyxl(read_only=True) and ~14x faster than openpyxl() full-load at 150k rows — but at that size it uses ~4.5x more peak memory than openpyxl(read_only=True) (192MB vs 43MB), because read() isn't actually O(1) yet (see above). If your bottleneck is wall-clock time, streamxl wins clearly. If your bottleneck is memory on a very large file and you don't need every column loaded at once, openpyxl(read_only=True) currently uses less RAM. Reproduce with benchmarks/openpyxl_vs_streamxl.py against any real .xlsx file.

CLI

pystreamxl dashboard          # sample extraction dashboard (unlabeled placeholder, see note above)
pystreamxl dashboard --static # same sample data, clearly labeled "SAMPLE DATA — not live"
pystreamxl --version

Development

git clone https://github.com/Mullassery/PyStreamXL.git
cd PyStreamXL
pip install -e ".[dev]"       # builds the Rust extension via maturin and installs test deps
pytest tests/ -v
cargo test --release --all-features     # Rust unit + integration tests (both core and python crates)

On macOS you may need RUSTFLAGS="-C link-args=-undefined -C link-args=dynamic_lookup" before cargo build/cargo test for the PyO3 extension crate to link outside of maturin/pip install.

See CONTRIBUTING.md before opening a PR.

Docs

License

This project is licensed under the Apache License 2.0.


StreamXL | Constant-memory Excel streaming | Rust core, Python API

Release files for streamxl 5.3.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for streamxl 5.3.2
File Size Uploaded
streamxl-5.3.2.tar.gz 84.7 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for streamxl 5.3.2
File Interpreter ABI Platform
streamxl-5.3.2-cp311-cp311-macosx_11_0_arm64.whl CPython 3.11 CPython 3.11 macOS 11.0+ ARM64 Details
streamxl-5.3.2-cp39-cp39-macosx_11_0_arm64.whl CPython 3.9 CPython 3.9 macOS 11.0+ ARM64 Details

Total release size: 2.0 MB

Release files / streamxl-5.3.2.tar.gz

Download URL streamxl-5.3.2.tar.gz
Size 84.7 kB
Tags Source
SHA-256 checksum
How to use checksums
6d28935904ea54276ab32ec145d2a30b11c0a504f13e781974b0506e6d010968
BLAKE2b-256 checksum
How to use checksums
d6f17c7eb435911946d56bd3bfb8c77e8b31975ab5e56e38c343b85073b49904
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / streamxl-5.3.2-cp311-cp311-macosx_11_0_arm64.whl

Download URL streamxl-5.3.2-cp311-cp311-macosx_11_0_arm64.whl
Size 980.6 kB
Tags CPython 3.11 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
63e582f51f349ae1156854f84746b40b0debff53f72b1a947d1cb05853b12069
BLAKE2b-256 checksum
How to use checksums
7a11c9f51f9c7f6de21348cb592636b247e4b5a921e7c5385e56bc02361b1a58
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / streamxl-5.3.2-cp39-cp39-macosx_11_0_arm64.whl

Download URL streamxl-5.3.2-cp39-cp39-macosx_11_0_arm64.whl
Size 981.0 kB
Tags CPython 3.9 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
c60dda5a87993cc98bfa522d608e80b8f8afea6a9951d611d63a99c7ad4a457d
BLAKE2b-256 checksum
How to use checksums
a9e22a81228fbf30f6b3528308c44856a605d8a120e8c40613752c35d3aa73ef
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

5.3.3

2 release files

This release

5.3.2 This release

3 release files

5.3.1

2 release files

5.3.0

2 release files

5.2.0

2 release files

5.1.0

2 release files

5.0.0

2 release files

4.0.0

2 release files

1.2.2

1 release file

1.2.1

1 release file

1.2.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page