Skip to main content

PySpark Datasource for ROOT

Apache Spark 4 Python Datasource for reading files in the ROOT data format used in High-Energy Physics (HEP).

Author and version
Luca.Canali@cern.ch · v0.1 (Sep 2025)

Highlights

  • ✅ Allows to read ROOT data using Apache Spark using a custom Spark 4 Python DataSource.
  • ✅ Works with local files, directories, and globs; optional XRootD (root://) support.
  • ✅ Implements partitioning and optional schema inference.
  • ✅ Powered by uproot, awkward, PyArrow and Spark's Python Datasource.

Related work & acknowledgments


Install

# From PyPI
pip install pyspark-root-datasource

# Or, local for development
pip install -e .

Quick start

from pyspark.sql import SparkSession
from pyspark_root_datasource import register

spark = (SparkSession.builder
         .appName("Read ROOT via PySpark + uproot")
         .getOrCreate())

# Register the datasource (short name = "root")
register(spark)

# Get the example ROOT file (2 GB)
# xrdcp root://eospublic.cern.ch//eos/opendata/cms/derived-data/AOD2NanoAODOutreachTool/Run2012BC_DoubleMuParked_Muons.root .
# if you don't have xrdcp installed, on Linux use wget or curl -O
wget https://sparkdltrigger.web.cern.ch/sparkdltrigger/Run2012BC_DoubleMuParked_Muons.root

# Best practice: provide a schema to prune branches early
schema = "nMuon int, Muon_pt array<float>, Muon_eta array<float>, Muon_phi array<float>, Muon_mass array<float>, Muon_charge array<int>"

df = (spark.read.format("root")
      .schema(schema)
      .option("path", "/data/Run2012BC_DoubleMuParked_Muons.root")
      .option("tree", "Events")
      .option("step_size", "1000000")
      .load())

df.show(5, truncate=False)
print("Count:", df.count())

# Use schema inference
df2 = (spark.read.format("root")
       .option("path", "/data/Run2012BC_DoubleMuParked_Muons.root")
       .option("tree", "Events")
       .option("sample_rows", "1000")   # default 1000
       .load())
df2.printSchema()

Examples and tests


Options

  • "path" (required) – file path, URL, comma-separated list, directory, or glob (e.g. "/data/*.root")
  • "tree" (default: "Events") – TTree name
  • "step_size" (default: "1000000") – entries per Spark partition (per file)
  • "num_partitions" (optional, per file) – overrides step_size
  • "entry_start", "entry_stop" (optional, per file) – index bounds
  • "columns" – comma-separated branch names (if not providing a Spark schema)
  • "list_to32" (default: "true") – Arrow list offset width
  • "extensionarray" (default: "false") – Arrow extension array support
  • "cast_unsigned" (default: "true") – cast uint* → signed (Spark lacks unsigned)
  • "recursive" (default: "false") – expand directories recursively
  • "ext" (default: "*.root") – filter pattern when path is a directory
  • "sample_rows" (default: "1000") – rows for schema inference
  • "arrow_max_chunksize" (default: "0") – if >0, limit rows per Arrow RecordBatch

Reading over XRootD (root://)

# fsspec plugins for xrootd
pip install fsspec fsspec-xrootd

# XRootD client libs + Python bindings
conda install -c conda-forge xrootd

Install the extras, then:

remote_file = "root://eospublic.cern.ch//eos/opendata/cms/derived-data/AOD2NanoAODOutreachTool/Run2012BC_DoubleMuParked_Muons.root"
df = (spark.read.format("root")
      .option("path", remote_file)
      .option("tree", "Events")
      .load())
df.show(3, truncate=False)

Reading folders, globs, recursion

# All .root files in a directory (non-recursive)
df = (spark.read.format("root")
      .option("path", "/data/myfolder")
      .load())

# Recursive directory expansion
df = (spark.read.format("root")
      .option("path", "/data/myfolder")
      .option("recursive", "true")
      .load())

# Custom extension used when 'path' is a directory
df = (spark.read.format("root")
      .option("path", "/data/myfolder")
      .option("ext", "*.parquet.root")
      .load())

# Glob
df = (spark.read.format("root")
      .option("path", "/data/*/atlas/*.root")
      .load())

Tips and troubleshooting

  • Prefer explicit schemas to prune early and minimize I/O.
  • Tune partitioning:
    • step_size = entries per Spark partition.
    • num_partitions (per file) overrides step_size.
  • Large jagged arrays benefit from reasonable step_size (e.g., 100k–1M).
  • If necessary, use arrow_max_chunksize to keep batch sizes moderate for downstream stages.
  • cast_unsigned=true normalizes uint* to signed widths (Spark-friendly).
  • Fixed-size lists are preserved as Arrow fixed_size_list (no silent downgrade).
  • XRootD errors: install both fsspec and fsspec-xrootd, and the XRootD client libs. Conda is often the smoothest:
    pip install fsspec fsspec-xrootd
    conda install -c conda-forge xrootd
    
  • Tree not found: double-check .option("tree", "..."); error messages list available keys.
  • Different schemas across files: ensure compatible branch types or read by subsets, then reconcile in Spark.
  • Driver vs executors env mismatch: set both spark.pyspark.python and spark.pyspark.driver.python to your Python.

Metadata

Release files for pyspark-root-datasource 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyspark-root-datasource 0.1.0
File Size Uploaded
pyspark_root_datasource-0.1.0.tar.gz 20.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyspark-root-datasource 0.1.0
File Interpreter ABI Platform
pyspark_root_datasource-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 35.6 kB

Release files / pyspark_root_datasource-0.1.0.tar.gz

Download URL pyspark_root_datasource-0.1.0.tar.gz
Size 20.4 kB
Tags Source
SHA-256 checksum
How to use checksums
39ef97649b4979caf91b335239330aa649aa553171a1f194a63c98b17748b786
BLAKE2b-256 checksum
How to use checksums
87c2d4ef2666c08007ef6998ed1f05f3d3e77bdce24c6bf66e52a1758a64e74d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release files / pyspark_root_datasource-0.1.0-py3-none-any.whl

Download URL pyspark_root_datasource-0.1.0-py3-none-any.whl
Size 15.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
27b8c6a827566a6e4d45bd220de5df1ac99d017dcf7c379eb3a61b23ff02af9f
BLAKE2b-256 checksum
How to use checksums
38af652ebd0799fe9228a7a80f806c6da4d9356df33d7a5339e3a31480a3a370
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page