Skip to main content

Universal MRF CSV parser for hospital price transparency data

Project description

mrf-etl

Universal MRF CSV parser for hospital price transparency data.

Parses any hospital Machine-Readable File (MRF) CSV — regardless of format, column naming, layout, or number of payers — and loads it into MySQL, Postgres, or normalized CSV files with zero data loss.


Install

# Core (no dependencies)
pip install -e .

# With MySQL support
pip install -e ".[mysql]"

# With Postgres support
pip install -e ".[postgres]"

# Both
pip install -e ".[all]"

Quick Start

Parse a single file → CSV

mrf-etl parse --input hospital_standardcharges.csv --output csv --out-dir ./output

Parse a single file → MySQL

mrf-etl parse --input hospital_standardcharges.csv --output mysql \
  --db-host localhost --db-user root --db-pass secret --db-name mrf_db

Inspect a file (no loading)

mrf-etl inspect --input hospital_standardcharges.csv

Bulk: process many hospitals

# Create a sources file (one path or URL per line)
mrf-etl bulk --input sources.txt --output csv --out-dir ./output --workers 4

Python API

from mrf_etl import parse_file, parse_metadata, CSVLoader

# Stream rows from any MRF file
for row in parse_file("hospital_standardcharges.csv"):
    print(row.description, row.gross_charge)
    for code in row.billing_codes:
        print(f"  {code.code_type}: {code.code}")
    for rate in row.rates:
        print(f"  {rate.payer_name_raw}: ${rate.negotiated_dollar}")

# Load to CSV
meta   = parse_metadata("hospital_standardcharges.csv")
rows   = parse_file("hospital_standardcharges.csv")
loader = CSVLoader(output_dir="./output")
stats  = loader.load(rows, meta, source_file="hospital_standardcharges.csv")
loader.close()
print(stats)  # {'items': ..., 'codes': ..., 'rates': ..., 'raw': ...}

# Load to MySQL
from mrf_etl import MySQLLoader
loader = MySQLLoader(host="localhost", user="root", password="secret", database="mrf_db")
stats  = loader.load(rows, meta, source_file="hospital_standardcharges.csv")

# Bulk processing
from mrf_etl import run_bulk, read_sources
sources = read_sources("sources.txt")
stats   = run_bulk(sources=sources, loader=loader, workers=4)
print(stats.summary())

Output Schema (5 tables)

mrf_hospitals       — one row per hospital file
mrf_items           — one row per procedure per hospital
mrf_item_codes      — all billing codes per item (CPT, HCPCS, RC, NDC, CDM...)
mrf_rates           — all payer/plan rates per item
mrf_raw             — original CSV row preserved (audit trail)

Supported Formats

  • Horizontal layout (payers as columns — up to 500+ payer columns)
  • Vertical layout (payer_name / plan_name columns)
  • Mixed layout
  • Up to N billing codes per procedure (code|1 through code|N)
  • Code types: CPT, HCPCS, MS-DRG, RC, NDC, CDM, LOCAL
  • Gzipped CSV (.csv.gz)
  • Zipped CSV (.zip)
  • Local files and HTTP/HTTPS URLs

Bulk Sources File Format

# sources.txt — one path or URL per line
# Comments and blank lines are ignored

/path/to/hospital1_standardcharges.csv
/path/to/hospital2_standardcharges.csv
https://hospital.org/standardcharges.csv
https://hospital.org/standardcharges.csv.gz

Run Tests

python tests/test_phase1.py
python tests/test_phase2.py
python tests/test_phase3.py
python tests/test_phase5.py

CLI Reference

mrf-etl inspect --input <file>
mrf-etl parse   --input <file> --output [csv|mysql|postgres] [options]
mrf-etl bulk    --input <sources.txt> --output [csv|mysql|postgres] [options]

Options:
  --out-dir       Output directory for CSV (default: ./mrf_output)
  --db-host       Database host (default: localhost)
  --db-port       Database port (default: 3306 MySQL / 5432 Postgres)
  --db-user       Database user
  --db-pass       Database password
  --db-name       Database name (default: mrf_db)
  --workers       Parallel threads for bulk (default: 4)
  --chunk-size    Batch insert size (default: 500)
  --max-retries   Retry attempts per file (default: 3)
  --force         Re-load even if file already loaded
  --error-report  Write error details to file
  --failed-sources Write failed paths to file for re-run

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mrf_etl-0.1.1.tar.gz (48.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mrf_etl-0.1.1-py3-none-any.whl (46.8 kB view details)

Uploaded Python 3

File details

Details for the file mrf_etl-0.1.1.tar.gz.

File metadata

  • Download URL: mrf_etl-0.1.1.tar.gz
  • Upload date:
  • Size: 48.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.4

File hashes

Hashes for mrf_etl-0.1.1.tar.gz
Algorithm Hash digest
SHA256 b4ccdd0079c034048509334b4c32200df42548236f136fa08392cef42d740ae5
MD5 cc08b1921c3988b9a93cf80340b159b7
BLAKE2b-256 c48ab51dfeff359d3fec823d80655ca8b1fee429418a3a541b67265b24640daa

See more details on using hashes here.

File details

Details for the file mrf_etl-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: mrf_etl-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 46.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.4

File hashes

Hashes for mrf_etl-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 7de91b9c10181b94c384e89f71b1a7251561f302be97ac7f59ec219d30016000
MD5 b1f220bfee7ee5ba5dc274e4fddaeb38
BLAKE2b-256 8dddd36b77f743961f665b4814ed3f21a4751cbf87ef173ac2cfad53c7803c87

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page