Skip to main content

polars_readstat

Polars plugin for SAS (.sas7bdat), Stata (.dta), and SPSS (.sav/.zsav) files.

The Python package wraps the Rust core in polars_readstat_rs and exposes a Polars-first API. The project includes cross-library parity tests and roundtrip checks to reduce regressions.

Why use this?

  • In project benchmarks, polars_readstat is typically faster than pandas/pyreadstat (usually at least 4x faster), and can be much, much faster for loading subsets of rows or columns.
  • It returns a Polars LazyFrame, so many polars operations that reduce the number of columns or rows you need can be pushed into the reader — only the data you actually need is read from disk.
  • Lots of customization and options for reading metadata (quickly), handling formats and labels, special null encoding, etc. to flexibly handle many needs when converting stat software data into a polars dataframe
  • Fast and low memory streaming of files to parquet, if you just want to get the data out of the original format to something more efficient (see code below).

Install

uv add polars-readstat

or if you prefer pip:

pip install polars-readstat

Docs/Core API

View the docs at https://jrothbaum.github.io/polars_readstat/ for more information than below on the options you can pass to the scan and write functions.

1) Lazy scan

import polars as pl
from polars_readstat import scan_readstat

lf = scan_readstat("/path/file.sas7bdat")
#   do something
df = lf.collect()

df = (
    scan_readstat("/path/file.sas7bdat")
    .select(["SERIALNO", "AGEP"])   # column pushdown — only these columns are read
    .head(1_000)                    # row limit is pushed down too
    .filter(pl.col("AGEP") >= 18)   # filters applied to streamed batches to avoid loading full file into memory
    .collect()
)

2) Getting metadata

from polars_readstat import ScanReadstat

reader = ScanReadstat(path="/path/file.sav")
schema = reader.schema           # polars.Schema
metadata = reader.metadata       # dict with file info and per-column details
lf = reader.df                   # LazyFrame — same as calling scan_readstat(path)

metadata is a dict with a variables (SPSS/Stata) or columns (SAS) list. Each entry includes:

  • "name" — column name
  • "label" — variable label (description), if present
  • "value_labels" — dict mapping coded values to label strings, if present

Polars lazy evaluation

scan_readstat returns a LazyFrame, so Polars can push operations into the reader before any data is loaded:

Read only specific columns — column selection is pushed into the reader; unselected columns are never read from disk:

lf = scan_readstat("file.sav")
df = lf.select(["id", "age", "income"]).collect()

Read the first N rows — head() / limit() stops the reader after N rows, so you never load the full file:

df = scan_readstat("file.sas7bdat").head(1000).collect()

Filter rows — filters are applied in Polars after reading, but still benefit from column pushdown if combined with .select():

df = scan_readstat("file.dta").select(["id", "age"]).filter(pl.col("age") >= 18).collect()

The benchmark numbers below reflect these optimizations — the large "Subset: True" speedups come from column pushdown.

3) Write

Writing support is experimental and compatibility varies across tools. Stata roundtrip tests are included; SPSS roundtrip coverage is limited. Please report issues.

from polars_readstat import write_readstat, write_sas_csv_import

write_readstat(df, "/path/out.dta")
write_readstat(df, "/path/out.sav")
write_sas_csv_import(df, "/path/out/sas_bundle", dataset_name="my_data")

write_readstat supports Stata (dta) and SPSS (sav).
Use write_sas_csv_import for SAS-ingestible output (.csv + .sas import script). Binary .sas7bdat writing is not currently supported.

4) You just want to get the data into parquet (or any other polars-supported file type):

Fast, Low-RAM Conversion

scan_readstat("/path/file.sas7bdat").sink_parquet("/path/file.parquet")

To give context, I converted a very tall sas file (over 8 billion rows) that was 500GB on disk to parquet and never exceeded 1GB of RAM.

Fast, Order-Preserving, but High-RAM Conversion

However, that code doesn't guarantee the row order is preserved, if you need that, use add preserve_order=True (at the cost of greater RAM usage).

scan_readstat("/path/file.sas7bdat",preserve_order=True).sink_parquet("/path/file.parquet")

Slower, Order-Preserving, Low-RAM Conversion

If you don't care as much about speed, but want the row order preserved and low RAM, set threads=1 and it will stream in order (somewhat slower) with super low ram usage.

scan_readstat("/path/file.sas7bdat",threads=1).sink_parquet("/path/file.parquet")

Benchmark

Benchmarks compare four scenarios: 1) load the full file, 2) load a subset of columns (Subset:True), 3) filter to a subset of rows (Filter: True), 4) load a subset of columns and filter to a subset of rows (Subset:True, Filter: True).

Benchmark context:

  • Machine: AMD Ryzen 7 8845HS (16 cores), 14 GiB RAM, Linux Mint 22
  • Storage: external SSD
  • Last run: May 14, 2026 — polars-readstat v0.17.0 vs pandas and pyreadstat
  • Method: wall-clock timings via Python time.time()

Compared to Pandas and Pyreadstat (using read_file_multiprocessing for parallel processing in Pyreadstat)

SAS

all times in seconds (speedup relative to pandas in parenthesis below each)

Library Full File Subset: True Filter: True Subset: True, Filter: True
polars_readstat 0.55
(3.9×)
0.07
(28.4×)
1.46
(2.0×)
0.08
(39.4×)
pandas 2.16 1.99 2.93 3.15
pyreadstat 6.76
(0.3×)
1.64
(1.2×)
7.86
(0.4×)
2.18
(1.4×)

Stata

all times in seconds (speedup relative to pandas in parenthesis below each)

Library Full File Subset: True Filter: True Subset: True, Filter: True
polars_readstat 0.16
(7.3×)
0.10
(11.7×)
0.18
(7.3×)
0.09
(13.8×)
pandas 1.17 1.17 1.31 1.24
pyreadstat 5.48
(0.2×)
4.57
(0.3×)
5.67
(0.2×)
7.69
(0.2×)

SPSS

all times in seconds (speedup relative to pandas in parenthesis below each)

Library Full File Subset: True Filter: True Subset: True, Filter: True
polars_readstat 1.09
(62.5×)
0.15
(3.9×)
1.10
(62.4×)
0.15
(3.9×)
pandas 68.12 0.59 68.67 0.59
pyreadstat 3.06
(22.3×)
1.15
(0.5×)
7.09
(9.7×)
1.23
(0.5×)

zsav

all times in seconds (speedup relative to pandas in parenthesis below each)

Library Full File Subset: True Filter: True Subset: True, Filter: True
polars_readstat 3.97
(5.9×)
1.04
(2.1×)
4.77
(4.7×)
1.15
(2.0×)
pandas 23.47 2.20 22.40 2.29

Detailed benchmark notes and dataset descriptions are in BENCHMARKS.md.

Tests run

Test coverage includes:

  • Cross-library comparisons on the pyreadstat and pandas test data, checking results against polars-readstat==0.11.1, pyreadstat, and pandas.
  • Stata/SPSS read/write roundtrip tests.
  • Large-file read/write benchmark runs on real-world data (results below).

If you want to run the same checks locally, helper scripts and tests are in scripts/ and tests/.

Release files for polars-readstat 0.23.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for polars-readstat 0.23.0
File
polars_readstat-0.23.0-cp39-abi3-win_amd64.whl CPython 3.9 abi3 Windows x86-64 Details
polars_readstat-0.23.0-cp39-abi3-manylinux_2_28_x86_64.whl CPython 3.9 abi3 Linux glibc 2.28+ x86-64 Details
polars_readstat-0.23.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.9 abi3 Linux glibc 2.17+ x86-64 Details
polars_readstat-0.23.0-cp39-abi3-macosx_11_0_arm64.whl CPython 3.9 abi3 macOS 11.0+ ARM64 Details
polars_readstat-0.23.0-cp39-abi3-macosx_10_15_x86_64.whl CPython 3.9 abi3 macOS 10.15+ x86-64 Details

Total release size: 117.2 MB

Release files / polars_readstat-0.23.0-cp39-abi3-win_amd64.whl

Download URL polars_readstat-0.23.0-cp39-abi3-win_amd64.whl
Size 25.5 MB
Tags CPython 3.9 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
4554c30437c3bd379627fda00ffacfbab6667c662c880b3df7c4ab4f0a153eb9
BLAKE2b-256 checksum
How to use checksums
0b2c5dbdf4693bcaab5cb6b6e6bb91cba964c908c615a07d5da3deea39cb2153
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / polars_readstat-0.23.0-cp39-abi3-manylinux_2_28_x86_64.whl

Download URL polars_readstat-0.23.0-cp39-abi3-manylinux_2_28_x86_64.whl
Size 24.0 MB
Tags CPython 3.9 Linux glibc 2.28+ x86-64 abi3
SHA-256 checksum
How to use checksums
c70a09f1620c005ad5f335b0925264671e98f688e4c37e0845751b19910bf201
BLAKE2b-256 checksum
How to use checksums
6a88ac7f3f7a9044c6b1a3665c45d038f991eed6c23cf27eebbe9412c0e8c3af
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / polars_readstat-0.23.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL polars_readstat-0.23.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 23.9 MB
Tags CPython 3.9 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
abf396d5c3261a8af3729af7f53503b13126db2e15d6ff0f2ddb26e524200d33
BLAKE2b-256 checksum
How to use checksums
cf5ee7c1102c48499c3a9251ef359731bd1306743347b176ccb669e47e848f7b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / polars_readstat-0.23.0-cp39-abi3-macosx_11_0_arm64.whl

Download URL polars_readstat-0.23.0-cp39-abi3-macosx_11_0_arm64.whl
Size 20.8 MB
Tags CPython 3.9 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
f66fa3050ad7d31295d98496f442f305510c14a9db1ae08f600f11087e2340ba
BLAKE2b-256 checksum
How to use checksums
eaf4feee688f5615e9153bddd0fa8ac5239ca7b8dff2c9d5a5fa600092597aa3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / polars_readstat-0.23.0-cp39-abi3-macosx_10_15_x86_64.whl

Download URL polars_readstat-0.23.0-cp39-abi3-macosx_10_15_x86_64.whl
Size 23.1 MB
Tags CPython 3.9 abi3 macOS 10.15+ x86-64
SHA-256 checksum
How to use checksums
69db6e0733bda62c850ed5dc9e621d08ec5e11563072738644dede5fa1763b04
BLAKE2b-256 checksum
How to use checksums
62f56263681d4f251b81c7e5f8b9cbd9da3f89a2c4cf0eb0b4084cc6786fa4b7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.23.0 This release

5 release files

0.21.1

5 release files

0.21.0

5 release files

0.20.2

5 release files

0.20.1

5 release files

0.19.4

5 release files

0.19.3

5 release files

0.19.2

5 release files

0.19.1

5 release files

0.19.0

5 release files

0.18.1

5 release files

0.17.0

5 release files

0.16.1

5 release files

0.15.0

5 release files

0.14.8

5 release files

0.14.7

5 release files

0.14.6

5 release files

0.14.5

5 release files

0.13.0

5 release files

0.12.4

5 release files

0.12.3

5 release files

0.12.2

5 release files

0.12.1

5 release files

0.12.0

5 release files

0.11.1

5 release files

0.11.0

5 release files

0.9.0

4 release files

0.8.2

4 release files

0.8.1

4 release files

0.7.2

4 release files

0.7.1

4 release files

0.7.0

4 release files

0.6.1

4 release files

0.6.0

4 release files

0.5.1

4 release files

0.4.2

4 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page