Fast extraction of CMS Transparency-in-Coverage machine-readable files to parquet
Project description
cms-mrf-extractor
Extract CMS Transparency-in-Coverage machine-readable files (MRFs) to parquet.
Payer MRFs are single JSON objects that reach tens of gigabytes decompressed. This
package reads one in a single pass, splits its arrays into byte blobs at exact item
boundaries, and parses those straight into Arrow buffers — around 160–200 MB/s of
decompressed JSON on one machine, against 55 MB/s for the obvious
ijson + json.loads implementation.
Local paths and s3:// URIs work the same for both input and output.
Install
pip install cms-mrf-extractor
Pure Python, no compiler needed. Reads .json, .gz, and ordinary .zip.
The parallel paths use fork(); on Windows the extractor says so and runs
sequentially in one process.
Deflate64-compressed .zip files need one extra:
pip install 'cms-mrf-extractor[deflate64]'
This one does build a C extension — it has no wheels for Python 3.11+ or arm64, so it needs a compiler (MSVC on Windows, Xcode CLT on macOS). Skip it unless you hit the deflate64 error, which names the extra when it fires.
Use it from the shell
mrf-extract /data/payer # -> ./output-payer/{pr,nr,nrpr}/
mrf-extract /data/payer -o /data/out-payer
mrf-extract s3://bucket/payer/ -o s3://bucket/parquet/
mrf-extract /data/payer --dry-run --json # which sections each file holds
mrf-extract --help lists every flag. The ones that matter most:
| Flag | What it does |
|---|---|
--rows-per-file N |
rows per parquet file (default 5000) |
--id-type, --npi-type |
force a column type instead of detecting it |
--workers N |
processes when there are more files than cores |
--blob-workers N |
decode workers inside a single file |
--no-arrow-json |
decode with json.loads (reference path, ~3x slower) |
--json |
run summary as JSON on stdout; logs stay on stderr |
Use it from Python
from cms_mrf_extractor import run
summary = run('/data/payer', output='/data/out-payer', rows_per_file=50_000)
# {'payer_in-network': {'nr': (1_284_530, 257), 'pr': (2_411, 1)}}
configure() takes the same keywords and sets them without extracting, and
discover() reports what a directory holds:
from cms_mrf_extractor import configure, discover, run
configure('/data/payer', id_type='string', progress=False)
for path, base_name, kinds, pr_key in discover():
print(base_name, kinds, pr_key) # e.g. payer_01 ['nr'] provider_references
run()
One configured extractor per process — settings are module state that the worker
processes inherit through fork().
What it writes
Three section types, each into its own subdirectory of the output, as
{source_name}_{kind}_{NNN}.parquet with zstd compression:
| Directory | Section | Shape |
|---|---|---|
pr/ |
provider_references |
provider_group_id, provider_groups[] (npi list + tin) |
nr/ |
in_network referencing provider ids |
rates carry provider_references[] |
nrpr/ |
in_network with groups inline |
rates carry provider_groups[] |
Which one a file produces is discovered from the file, not from its name. The nested list/struct shape of the source is preserved; nothing is exploded or joined.
Some payers ship the provider definitions under provider_group_reference with
business_name beside tin rather than inside it. That layout is detected and
reshaped to the same schema. Items carrying only a location URL are fetched over
HTTP and inlined.
Column types
provider_group_id / provider_reference and npi are not the same JSON type in
every payer's files — some quote npi, some ship ids too large for int64. The
types are read out of the files during discovery and widened one way
(int64 → decimal128 → string for ids, int64 → string for npi), then the widest
answer across the directory is used, because one run writes one schema.
Detection reads forward from byte 0 under a bounded event budget. A payer that puts
its provider section after a multi-gigabyte in_network array hides it behind
more events than that budget allows; the run warns and names those files. Pass
--npi-type string if the warning applies to you.
Correctness ladder
Three tiers, each strictly more general than the last, per file:
- Arrow-native —
pyarrow.jsonparses blobs into Arrow buffers directly. - Per-blob Python —
json.JSONDecoder.raw_decode, one blob at a time. - ijson — structure-agnostic streaming, for layouts the byte splitter cannot prove boundaries for.
A failure at any tier discards that file's output and retries at the next one, so a
layout the fast path cannot handle costs speed rather than a run. raw_decode
consumes exactly one JSON value and reports where it ended, so a bad split surfaces
as a JSONDecodeError and a retry — never as corrupted parquet. Truncated files are
reported and listed in bad_files.log in the output directory.
Environment variables
Equivalent to the CLI flags, useful for containers:
| Variable | Effect |
|---|---|
MRF_SOURCE_DIR, MRF_OUTPUT_DIR |
source and output when none is passed |
MRF_ID_TYPE, MRF_NPI_TYPE |
force a column type |
MRF_TYPE_DETECT=0 |
skip detection, keep the int64 defaults |
MRF_ARROW_JSON=0 |
decode with json.loads |
MRF_STREAM_JSON=0 |
decode blob by blob rather than streaming a section |
The MRF_ARROW_JSON / MRF_STREAM_JSON switches exist so a suspected regression
can be bisected across the three tiers without editing code.
Why it is fast
PERF_NOTES.md in the source distribution carries the measurements. The short
version:
ijson.items()never stops early, so pulling an 80 MB array out of a 34 GB file read all 34 GB — 399 s instead of 2 s. Structure discovery still uses ijson, where incremental parsing is the right tool; extraction does not.- Items are located by
bytes.find()on the first item's first key (~3.2 GB/s), so a section becomes independent byte blobs without parsing anything. - Building Python dicts and walking them with
pa.Table.from_pylistran at 81 MB/s per core and scaled at 35% efficiency across 12 processes — a 10 MB blob becomes 50–100 MB of Python objects and the workers end up bound by the allocator.pyarrow.jsonon the same bytes creates no Python objects: 161–343 MB/s per core. - One pass covers every section of a file, in the order the sections appear, so a
file with
provider_referenceslast costs the same as one with it first.
License
MIT — see LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cms_mrf_extractor-0.1.1.tar.gz.
File metadata
- Download URL: cms_mrf_extractor-0.1.1.tar.gz
- Upload date:
- Size: 53.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c4dcea794e25dd9d23a3ef03b24dfd6a936f6d0275088081fca1fa65472a3fff
|
|
| MD5 |
427c16e000818de0e31106209b740dfd
|
|
| BLAKE2b-256 |
0b2a3bdf23671c64b4e857deea3c6fa6f0a811c28904ac0f350c7e3fd7cb6390
|
File details
Details for the file cms_mrf_extractor-0.1.1-py3-none-any.whl.
File metadata
- Download URL: cms_mrf_extractor-0.1.1-py3-none-any.whl
- Upload date:
- Size: 35.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
37704bd76dbf94896fabebc7d9454ed2d6ae5f6dd676bbdb70346fcc7365ba72
|
|
| MD5 |
07bdd7b84b86a996eafc487fb13608e8
|
|
| BLAKE2b-256 |
3e0efed9971375e98803a035b370c3d5b9dc28776eedf272e4407d3656f490eb
|