StreamXL
Stream large .xlsx files row-by-row in constant memory, powered by a Rust core with no unsafe code.
pip installs as streamxl, import streamxl. Read multi-sheet Excel workbooks without loading them fully into memory, extract formulas and comments, write new .xlsx files, and append to existing ones — all through a small, plain Python API backed by a Rust engine.
Install
pip install streamxl
A prebuilt wheel is currently published only for macOS (arm64); other platforms install from the source distribution, which requires a Rust toolchain (see rust-toolchain.toml) and maturin to build. Every PyPI release to date (1.2.0 through 5.2.0) has shipped exactly one platform wheel plus an sdist — no Linux or Windows wheels have been published yet.
Quick start
import streamxl
for row in streamxl.read("data.xlsx"):
print(row) # ['Name', 'Age', 'Score']
read() streams rows one at a time — memory use stays flat regardless of file size.
Real, working examples
Read as dictionaries, keyed by header row:
import streamxl
for row in streamxl.read("sales.xlsx", as_dict=True):
print(row["Customer"], row["Amount"])
Read only specific columns:
for row in streamxl.read("sales.xlsx", as_dict=True, columns=["Customer", "Amount"]):
...
Read every sheet in a workbook:
sheet_names = streamxl.sheets("workbook.xlsx")
all_data = streamxl.read_all("workbook.xlsx") # {sheet_name: [rows...]}
Write a new .xlsx file:
import datetime
import streamxl
streamxl.write("report.xlsx", [
["Name", "Joined", "Score"],
["Alice", datetime.date(2024, 1, 15), 95.5],
["Bob", datetime.date(2024, 3, 2), 88.0],
])
Stream-write multiple sheets without holding the whole file in memory:
with streamxl.writer("report.xlsx") as w:
w.write_row(["Name", "Age"])
w.write_row(["Alice", 30])
w.add_sheet("Summary")
w.write_row(["Total", 1])
Append rows to an existing file (other sheets are preserved):
streamxl.write("log.xlsx", [["Date", "Event"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "started"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "finished"]])
Extract formulas and comments:
rows = list(streamxl.read("model.xlsx", with_formulas=True))
# each cell is a dict: {"value": ..., "formula": ..., "formula_type": ...,
# "comment": ..., "comment_author": ...}
from streamxl import FormulaSerializer
export = FormulaSerializer.export_formulas(rows)
FormulaSerializer.export_to_json(rows, "formulas.json")
FormulaSerializer.export_to_csv(rows, "formulas.csv") # sanitized against CSV/formula injection
Export to CSV safely — untrusted cell content is never written to CSV verbatim (see Security below):
import csv
import streamxl
from streamxl.security import sanitize_csv_cell
with open("output.csv", "w", newline="") as f:
writer = csv.writer(f)
for row in streamxl.read("large.xlsx"):
writer.writerow([sanitize_csv_cell(cell) for cell in row])
Validate a file and recover from bad cells instead of crashing:
from streamxl import validate_excel_file
report = validate_excel_file("questionable.xlsx")
if report.has_fatal_errors():
print(report.format_summary())
More runnable examples live in examples/.
Honest feature list
What's here and real, backed by the Rust core and covered by the test suite:
- Streaming reads —
read()/stream(): real, pull-based streaming backed by a Rust__iter__/__next__iterator over the sheet, not a full-sheet materialization dressed up as a generator — O(1) memory per row, regardless of file size.read_rows_all_at_once()/read_rows_with_metadata_all_at_once()remain available as an explicit escape hatch for callers that need random access or to iterate the result more than once. - Multi-sheet support —
sheets(),read_all(), andwriter().add_sheet(). - Streaming writes —
write(),writer(),append(), all producing real.xlsxfiles. - Formula extraction — read formula text and a best-effort formula-type classification (
with_formulas=True), plusFormulaReferenceMapperfor shifting/rewriting cell references andFormulaSerializerfor exporting/importing formulas as JSON or CSV. - Comment extraction — cell comments and authors, via
with_formulas=True. - Conditional formatting rules —
conditional_formats()reads every<conditionalFormatting>/<cfRule>in a sheet (type, operator, formulas, priority,stopIfTrue) and resolves each rule'sdxfIdagainstxl/styles.xml's<dxfs>into concrete font color/bold/italic and fill colors.colorScale/dataBar/iconSetrules are captured (type, sqref, priority) but their inline color-stop/threshold definitions aren't modeled — those rule types don't usedxfIdin the first place. - Type-aware cells — strings, numbers, booleans, dates, datetimes, and empty cells round-trip correctly.
- Error recovery & validation —
validate_excel_file()andErrorRecoveryHandlerclassify and (optionally) recover from malformed cells instead of hard-failing on the whole file. - Security hardening — path validation, file-size limits, and ZIP-bomb defenses (entry-size, compression-ratio, and total-decompressed-size limits) enforced before/while a file is opened. CSV export is sanitized against formula-injection (see below).
- REST API (optional) —
streamxl.server.StreamXLServer/create_flask_app()wrap the real streaming engine behind HTTP endpoints (/sources,/sources/<id>/query,/sources/<id>/export, ...). Requirespip install "streamxl[server]".
What's not here, so you don't have to find out the hard way:
- No SQL-style query language —
execute_query()in the REST API streams rows from a named sheet, it does not parse arbitrary queries. - No pandas/Parquet/Arrow export built in. Convert
read()'s output yourself, or open an issue if this matters to you. - No formula evaluation — formula text is extracted and classified, not recalculated.
- The
pystreamxl dashboardCLI command currently renders sample data, not live telemetry. Onlypystreamxl dashboard --static(and--alerts/--recommendations/--export) labels this clearly, with an explicit "SAMPLE DATA — not live" warning; the barepystreamxl dashboard(default interactive mode) currently prints an unlabeled placeholder (Status: Active) with no such disclaimer.
Security
- Path & size validation —
validate_read_path()/validate_write_path()reject non-.xlsxpaths, path traversal, and oversized files before any parsing happens. - ZIP-bomb defenses — the Rust core enforces a per-entry size limit, a compression-ratio limit, and a total-decompressed-size limit while unpacking a workbook (see
core/src/zip_reader.rs), tested against real crafted archives incore/tests/zip_bomb_defense.rs. - CSV/formula-injection protection —
streamxl.security.sanitize_csv_cell()neutralizes any string cell that starts with=,+,-,@, TAB, or CR (the standard CSV-injection trigger set) by prefixing it with', so a malicious workbook can't turn a CSV export into an executable formula when reopened in Excel/LibreOffice/Google Sheets.FormulaSerializer.export_to_csv()applies this automatically; apply it yourself when writing CSV fromread()output (see the example above).
Limits, enforced by default (no configuration needed):
| Limit | Value |
|---|---|
| Max file size | 512 MB |
| Max size per ZIP entry | 512 MB |
| Max total decompressed size | 1 GB |
| Max compression ratio | 30:1 |
Handle malformed or malicious files by catching SecurityError:
from streamxl import SecurityError, read
try:
for row in read("data.xlsx"):
process(row)
except SecurityError as e:
print(f"Security violation: {e}")
Found a security issue? See SECURITY.md.
Performance
Streaming keeps memory flat regardless of file size, since rows are parsed and yielded one at a time instead of materializing the whole workbook. See benchmarks/ for the scripts used to compare against openpyxl, and examples/memory_benchmark.py to measure it yourself against your own files:
python examples/memory_benchmark.py your_file.xlsx
Actual numbers depend heavily on your file's structure (shared strings, formulas, formatting) — measure on your own workloads rather than trusting a generic table.
CLI
pystreamxl dashboard # sample extraction dashboard (unlabeled placeholder, see note above)
pystreamxl dashboard --static # same sample data, clearly labeled "SAMPLE DATA — not live"
pystreamxl --version
Development
git clone https://github.com/Mullassery/StreamXL.git
cd StreamXL
pip install -e ".[dev]" # builds the Rust extension via maturin and installs test deps
pytest tests/ -v
cargo test --all-features # Rust unit + integration tests
License
Proprietary License — free to use with explicit attribution. See LICENSE.
StreamXL | Constant-memory Excel streaming | Rust core, Python API
Release files for streamxl 5.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| streamxl-5.2.0.tar.gz | 74.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| streamxl-5.2.0-cp311-cp311-macosx_11_0_arm64.whl | CPython 3.11 | CPython 3.11 | macOS 11.0+ ARM64 | Details |
Total release size: 915.5 kB
Release files / streamxl-5.2.0.tar.gz
| Download URL | streamxl-5.2.0.tar.gz |
|---|---|
| Size | 74.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
becc1ef3db38ada3e69db3b6704dbcd3b9830b40149185864574515c2c144eeb
|
|
BLAKE2b-256 checksum How to use checksums |
a72d8fa9f4115c14c3701e30091a3d9e1043583eb2e59ff4fcc03d30fb22159e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|
Release files / streamxl-5.2.0-cp311-cp311-macosx_11_0_arm64.whl
| Download URL | streamxl-5.2.0-cp311-cp311-macosx_11_0_arm64.whl |
|---|---|
| Size | 841.2 kB |
| Tags | CPython 3.11 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
1bad0edbc3017037b99240fa02e19db6636551d402f2fcfb84f5feda25dda302
|
|
BLAKE2b-256 checksum How to use checksums |
be6035262f3e794cd1fce1a305d257726ef176ad9fd648e65b08a4876e1ade1e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|