StreamXL
Problem
openpyxl and similar pure-Python Excel readers load the whole workbook
into memory before you can touch a single row — fine for small files, a
real ceiling for ETL pipelines and data engineering workloads working
against large .xlsx exports.
Solution
Stream large .xlsx files row-by-row in constant memory, powered by a Rust core with no unsafe code.
pip installs as streamxl, import streamxl. Read multi-sheet Excel workbooks without loading them fully into memory, extract formulas and comments, write new .xlsx files, and append to existing ones — all through a small, plain Python API backed by a Rust engine.
Use cases
- ETL against large Excel exports that don't fit comfortably in memory
with
openpyxl—read()keeps memory flat regardless of file size. - Extracting formulas/comments for audit or migration tooling, not just cell values.
- Appending to a growing log-style
.xlsxfile without rewriting the whole workbook or losing other sheets. - Not yet a good fit for: SQL-style querying across sheets, formula evaluation (only extraction/classification), or pandas/Parquet/Arrow export built in — see Honest feature list for the full "what's not here" list.
Install
pip install streamxl
A prebuilt wheel is currently published only for macOS (arm64); other platforms install from the source distribution, which requires a Rust toolchain (see rust-toolchain.toml) and maturin to build. Every PyPI release to date (1.2.0 through 5.2.0) has shipped exactly one platform wheel plus an sdist — no Linux or Windows wheels have been published yet.
Quick start
import streamxl
for row in streamxl.read("data.xlsx"):
print(row) # ['Name', 'Age', 'Score']
read() streams rows one at a time — memory use stays flat regardless of file size.
Real, working examples
Read as dictionaries, keyed by header row:
import streamxl
for row in streamxl.read("sales.xlsx", as_dict=True):
print(row["Customer"], row["Amount"])
Read only specific columns:
for row in streamxl.read("sales.xlsx", as_dict=True, columns=["Customer", "Amount"]):
...
Read every sheet in a workbook:
sheet_names = streamxl.sheets("workbook.xlsx")
all_data = streamxl.read_all("workbook.xlsx") # {sheet_name: [rows...]}
Write a new .xlsx file:
import datetime
import streamxl
streamxl.write("report.xlsx", [
["Name", "Joined", "Score"],
["Alice", datetime.date(2024, 1, 15), 95.5],
["Bob", datetime.date(2024, 3, 2), 88.0],
])
Stream-write multiple sheets without holding the whole file in memory:
with streamxl.writer("report.xlsx") as w:
w.write_row(["Name", "Age"])
w.write_row(["Alice", 30])
w.add_sheet("Summary")
w.write_row(["Total", 1])
Append rows to an existing file (other sheets are preserved):
streamxl.write("log.xlsx", [["Date", "Event"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "started"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "finished"]])
Extract formulas and comments:
rows = list(streamxl.read("model.xlsx", with_formulas=True))
# each cell is a dict: {"value": ..., "formula": ..., "formula_type": ...,
# "comment": ..., "comment_author": ...}
from streamxl import FormulaSerializer
export = FormulaSerializer.export_formulas(rows)
FormulaSerializer.export_to_json(rows, "formulas.json")
FormulaSerializer.export_to_csv(rows, "formulas.csv") # sanitized against CSV/formula injection
Export to CSV safely — untrusted cell content is never written to CSV verbatim (see Security below):
import csv
import streamxl
from streamxl.security import sanitize_csv_cell
with open("output.csv", "w", newline="") as f:
writer = csv.writer(f)
for row in streamxl.read("large.xlsx"):
writer.writerow([sanitize_csv_cell(cell) for cell in row])
Validate a file and recover from bad cells instead of crashing:
from streamxl import validate_excel_file
report = validate_excel_file("questionable.xlsx")
if report.has_fatal_errors():
print(report.format_summary())
More runnable examples live in examples/.
Honest feature list
What's here and real, backed by the Rust core and covered by the test suite:
- Streaming reads —
read()/stream(): a real Rust__iter__/__next__iterator over the sheet — you get rows one at a time, not a pre-built Python list. Correction (2026-09-22 benchmark): this is not O(1) memory as previously claimed here.XlsxStream::open()decompresses the whole sheet XML into memory before iteration starts, so peak RSS scales with sheet size (measured ~1.3MB of RSS per 1MB of.xlsx, real data — seebenchmarks/results.md). It's still far belowopenpyxl's full-load mode and the API shape (a lazy iterator) is real, but memory is not flat regardless of file size — tracked as gap #10 inROADMAP_HONEST.md.read_rows_all_at_once()/read_rows_with_metadata_all_at_once()remain available as an explicit escape hatch for callers that need random access or to iterate the result more than once. - Multi-sheet support —
sheets(),read_all(), andwriter().add_sheet(). - Streaming writes —
write(),writer(),append(), all producing real.xlsxfiles. - Formula extraction — read formula text and a best-effort formula-type classification (
with_formulas=True), plusFormulaReferenceMapperfor shifting/rewriting cell references andFormulaSerializerfor exporting/importing formulas as JSON or CSV. - Comment extraction — cell comments and authors, via
with_formulas=True. - Conditional formatting rules —
conditional_formats()reads every<conditionalFormatting>/<cfRule>in a sheet (type, operator, formulas, priority,stopIfTrue) and resolves each rule'sdxfIdagainstxl/styles.xml's<dxfs>into concrete font color/bold/italic and fill colors.colorScale/dataBar/iconSetrules are captured (type, sqref, priority) but their inline color-stop/threshold definitions aren't modeled — those rule types don't usedxfIdin the first place. - Type-aware cells — strings, numbers, booleans, dates, datetimes, and empty cells round-trip correctly.
- Error recovery & validation —
validate_excel_file()andErrorRecoveryHandlerclassify and (optionally) recover from malformed cells instead of hard-failing on the whole file. - Security hardening — path validation, file-size limits, and ZIP-bomb defenses (entry-size, compression-ratio, and total-decompressed-size limits) enforced before/while a file is opened. CSV export is sanitized against formula-injection (see below).
- REST API (optional) —
streamxl.server.StreamXLServer/create_flask_app()wrap the real streaming engine behind HTTP endpoints (/sources,/sources/<id>/query,/sources/<id>/export, ...). Requirespip install "streamxl[server]".
What's not here, so you don't have to find out the hard way:
- No SQL-style query language —
execute_query()in the REST API streams rows from a named sheet, it does not parse arbitrary queries. - No pandas/Parquet/Arrow export built in. Convert
read()'s output yourself, or open an issue if this matters to you. - No formula evaluation — formula text is extracted and classified, not recalculated.
- The
pystreamxl dashboardCLI command renders sample data, not live telemetry — every mode (bare,--static,--alerts,--recommendations,--export) shows the same explicit "SAMPLE DATA — not live" warning.
Security
- Path & size validation —
validate_read_path()/validate_write_path()reject non-.xlsxpaths, path traversal, and oversized files before any parsing happens. - ZIP-bomb defenses — the Rust core enforces a per-entry size limit, a compression-ratio limit, and a total-decompressed-size limit while unpacking a workbook (see
core/src/zip_reader.rs), tested against real crafted archives incore/tests/zip_bomb_defense.rs. - CSV/formula-injection protection —
streamxl.security.sanitize_csv_cell()neutralizes any string cell that starts with=,+,-,@, TAB, or CR (the standard CSV-injection trigger set) by prefixing it with', so a malicious workbook can't turn a CSV export into an executable formula when reopened in Excel/LibreOffice/Google Sheets.FormulaSerializer.export_to_csv()applies this automatically; apply it yourself when writing CSV fromread()output (see the example above).
Limits, enforced by default (no configuration needed):
| Limit | Value |
|---|---|
| Max file size | 512 MB |
| Max size per ZIP entry | 512 MB |
| Max total decompressed size | 1 GB |
| Max compression ratio | 30:1 |
Handle malformed or malicious files by catching SecurityError:
from streamxl import SecurityError, read
try:
for row in read("data.xlsx"):
process(row)
except SecurityError as e:
print(f"Security violation: {e}")
Found a security issue? See SECURITY.md.
Performance
Rows are parsed and yielded one at a time rather than being collected into a Python list up front, and it's consistently faster than openpyxl on both reads and writes. Memory use is lower than openpyxl's full-load mode but currently scales with sheet size rather than staying flat — see the correction under "Honest feature list" above and ROADMAP_HONEST.md gap #10. See benchmarks/ for the scripts used to compare against openpyxl, and examples/memory_benchmark.py to measure it yourself against your own files:
python examples/memory_benchmark.py your_file.xlsx
Actual numbers depend heavily on your file's structure (shared strings, formulas, formatting) — measure on your own workloads rather than trusting a generic table.
vs openpyxl, on real data
Methodology: 150,000 real, live NYC 311 Service Request rows pulled from
NYC Open Data's Socrata API (data.cityofnewyork.us/resource/erm2-nwe9,
current as of 2026-09-22 — not synthetic/fabricated rows), 14 columns,
written to a real 2-sheet .xlsx workbook (75k rows/sheet, 19MB) via
openpyxl. Both libraries iterated every row of every sheet; row counts
and a positional checksum matched exactly across all three methods
(correctness verified, not just speed). 3 runs each, median reported,
single-process wall-clock via time.perf_counter(), peak RSS via
resource.getrusage(...).ru_maxrss on macOS/arm64, Python 3.13.
| Rows | streamxl read() |
openpyxl read_only=True |
openpyxl full load |
|---|---|---|---|
| 10,000 | 0.07s · 28MB peak RSS | 0.61s · 31MB peak RSS | — |
| 30,000 | 0.19s · 52MB peak RSS | 1.89s · 32MB peak RSS | — |
| 75,000 | 0.48s · 103MB peak RSS | 4.72s · 36MB peak RSS | — |
| 150,000 (2 sheets) | 0.96s · 192MB peak RSS | 9.27s · 43MB peak RSS | 13.5s · 1,040MB peak RSS |
streamxl is ~9.7x faster than openpyxl(read_only=True) and ~14x
faster than openpyxl() full-load at 150k rows — but at that size it
uses ~4.5x more peak memory than openpyxl(read_only=True) (192MB
vs 43MB), because read() isn't actually O(1) yet (see above). If your
bottleneck is wall-clock time, streamxl wins clearly. If your bottleneck
is memory on a very large file and you don't need every column loaded at
once, openpyxl(read_only=True) currently uses less RAM. Reproduce with
benchmarks/openpyxl_vs_streamxl.py against any real .xlsx file.
CLI
pystreamxl dashboard # sample extraction dashboard (unlabeled placeholder, see note above)
pystreamxl dashboard --static # same sample data, clearly labeled "SAMPLE DATA — not live"
pystreamxl --version
Development
git clone https://github.com/Mullassery/PyStreamXL.git
cd PyStreamXL
pip install -e ".[dev]" # builds the Rust extension via maturin and installs test deps
pytest tests/ -v
cargo test --release --all-features # Rust unit + integration tests (both core and python crates)
On macOS you may need RUSTFLAGS="-C link-args=-undefined -C link-args=dynamic_lookup" before cargo build/cargo test for the PyO3 extension crate to link outside of maturin/pip install.
See CONTRIBUTING.md before opening a PR.
Docs
docs/architecture/README.md— how the Rust engine and Python API fit together, including known dead codedocs/xlsx_format.md— XLSX/ZIP/XML format notesROADMAP_HONEST.md— unvarnished list of what's missing, broken, or technical debtCHANGELOG.md— release historySECURITY.md— security model, limits, and what it does not protect against
License
This project is licensed under the Apache License 2.0.
StreamXL | Constant-memory Excel streaming | Rust core, Python API
Release files for streamxl 5.3.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| streamxl-5.3.2.tar.gz | 84.7 kB | Details |
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| streamxl-5.3.2-cp311-cp311-macosx_11_0_arm64.whl | CPython 3.11 | CPython 3.11 | macOS 11.0+ ARM64 | Details |
| streamxl-5.3.2-cp39-cp39-macosx_11_0_arm64.whl | CPython 3.9 | CPython 3.9 | macOS 11.0+ ARM64 | Details |
Total release size: 2.0 MB
Release files / streamxl-5.3.2.tar.gz
| Download URL | streamxl-5.3.2.tar.gz |
|---|---|
| Size | 84.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6d28935904ea54276ab32ec145d2a30b11c0a504f13e781974b0506e6d010968
|
|
BLAKE2b-256 checksum How to use checksums |
d6f17c7eb435911946d56bd3bfb8c77e8b31975ab5e56e38c343b85073b49904
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|
Release files / streamxl-5.3.2-cp311-cp311-macosx_11_0_arm64.whl
| Download URL | streamxl-5.3.2-cp311-cp311-macosx_11_0_arm64.whl |
|---|---|
| Size | 980.6 kB |
| Tags | CPython 3.11 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
63e582f51f349ae1156854f84746b40b0debff53f72b1a947d1cb05853b12069
|
|
BLAKE2b-256 checksum How to use checksums |
7a11c9f51f9c7f6de21348cb592636b247e4b5a921e7c5385e56bc02361b1a58
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|
Release files / streamxl-5.3.2-cp39-cp39-macosx_11_0_arm64.whl
| Download URL | streamxl-5.3.2-cp39-cp39-macosx_11_0_arm64.whl |
|---|---|
| Size | 981.0 kB |
| Tags | CPython 3.9 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
c60dda5a87993cc98bfa522d608e80b8f8afea6a9951d611d63a99c7ad4a457d
|
|
BLAKE2b-256 checksum How to use checksums |
a9e22a81228fbf30f6b3528308c44856a605d8a120e8c40613752c35d3aa73ef
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|