Skip to main content

sparkpl

A lightweight, pandas-free Python package for seamless conversion between PySpark and Polars DataFrames.

Installation

pip install sparkpl

Features

  • 🚀 Direct Arrow conversion - Uses native Arrow for maximum performance (Spark 4.0+)
  • ⚡ Zero pandas dependency - Pure Polars ↔ Spark conversion
  • 🔄 Bidirectional conversion - Seamless data exchange between frameworks
  • 🛡️ Type preservation - Maintains data types during conversion
  • 📊 Batch processing - Handles large datasets efficiently
  • 🔍 Smart logging - Structured logging with loguru
  • 🎯 Simple API - Both functional and class-based interfaces
  • 💾 Minimal footprint - Lightweight with essential dependencies only

Quick Start

import polars as pl
from pyspark.sql import SparkSession
from sparkpl.converter import spark_to_polars, polars_to_spark

# Initialize Spark
spark = SparkSession.builder.appName("example").getOrCreate()

# Create sample data
spark_df = spark.createDataFrame([(1, "Alice"), (2, "Bob")], ["id", "name"])

# Convert Spark → Polars
polars_df = spark_to_polars(spark_df)
print(polars_df)

# Convert Polars → Spark
spark_df_back = polars_to_spark(polars_df)
spark_df_back.show()

Advanced Usage

Class-based API

from sparkpl.converter import DataFrameConverter

converter = DataFrameConverter(spark)

# With Arrow optimization (default)
polars_df = converter.spark_to_polars(spark_df, use_arrow=True)

# Native fallback for compatibility
polars_df = converter.spark_to_polars(spark_df, use_arrow=False)

# Batch processing for large datasets
polars_df = converter.spark_to_polars(large_spark_df, batch_size=100000)

# Create temporary view
spark_df = converter.polars_to_spark(polars_df, table_name="my_table")

Error Handling

from sparkpl.converter import DataFrameConverterError

try:
    polars_df = spark_to_polars(spark_df)
except DataFrameConverterError as e:
    print(f"Conversion failed: {e}")

Logging Configuration

from loguru import logger

# Configure structured logging
logger.add("sparkpl.log", rotation="10 MB", level="INFO")

# Conversions automatically log progress
polars_df = spark_to_polars(spark_df)  # Logs conversion metrics

Performance

SparkPL automatically selects the optimal conversion method:

  • Spark 4.0+: Direct Arrow conversion (toArrow() → createDataFrame(arrow_table))
  • Older versions: Native collection methods with fallback
  • Large datasets: Automatic batching to manage memory

Type Support

Polars Type Spark Type Notes
pl.Utf8 StringType
pl.Int32 IntegerType
pl.Int64 LongType
pl.Float32 FloatType
pl.Float64 DoubleType
pl.Boolean BooleanType
pl.Date DateType
pl.Datetime TimestampType
pl.Binary BinaryType
pl.Time StringType Converted to string
pl.Duration LongType Microseconds

Requirements

  • Python >=3.8
  • polars >=0.18.0
  • pyspark >=3.0.0
  • pyarrow >=5.0.0
  • loguru >=0.6.0

API Reference

Functions

  • spark_to_polars(spark_df, **kwargs) - Convert Spark DataFrame to Polars
  • polars_to_spark(polars_df, **kwargs) - Convert Polars DataFrame to Spark

DataFrameConverter Class

  • spark_to_polars(spark_df, use_arrow=True, batch_size=None)
  • polars_to_spark(polars_df, use_arrow=True, table_name=None)
  • validate_conversion(original_df, converted_df, check_data=False)

Why No Pandas?

SparkPL eliminates pandas dependency for:

  • Reduced footprint - Fewer dependencies to manage
  • Better performance - Direct conversion without intermediate steps
  • Simplified deployment - No pandas version conflicts
  • Pure workflow - Stay within Polars/Spark ecosystem

Examples

Basic Conversion

# Sample data
data = [("Alice", 25), ("Bob", 30), ("Charlie", 35)]
spark_df = spark.createDataFrame(data, ["name", "age"])

# Convert and process
polars_df = spark_to_polars(spark_df)
filtered = polars_df.filter(pl.col("age") > 28)
result_spark = polars_to_spark(filtered)

Working with Large Data

# Process large dataset in chunks
converter = DataFrameConverter(spark)
large_polars = converter.spark_to_polars(
    huge_spark_df, 
    batch_size=50000  # Process 50k rows at a time
)

Contributing

  1. Fork the repository
  2. Create feature branch: git checkout -b feature/my-feature
  3. Make changes with tests
  4. Commit: git commit -am 'Add feature'
  5. Push: git push origin feature/my-feature
  6. Create pull request

Development Setup

git clone https://github.com/yourusername/sparkpl.git
cd sparkpl
pip install -e ".[dev]"
pytest tests/

License

MIT License - see LICENSE file.

Support

  • Issues: GitHub Issues
  • Documentation: Coming soon
  • Community: Discussions welcome

Built with ❤️ for the Python data community.

Metadata

Release files for sparkpl 2.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sparkpl 2.0.1
File Size Uploaded
sparkpl-2.0.1.tar.gz 7.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sparkpl 2.0.1
File Interpreter ABI Platform
sparkpl-2.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 15.4 kB

Release files / sparkpl-2.0.1.tar.gz

Download URL sparkpl-2.0.1.tar.gz
Size 7.7 kB
Tags Source
SHA-256 checksum
How to use checksums
d6f9fcdad0da8ef490576ad76fb1cd558a07d0bb6f5009218d08df5586d5434a
BLAKE2b-256 checksum
How to use checksums
b8aa9bd97f55cf2c71dc1c151ad1f4c8ab8d99c48a18713e5b510c9bd80370f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.11.12

Release files / sparkpl-2.0.1-py3-none-any.whl

Download URL sparkpl-2.0.1-py3-none-any.whl
Size 7.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f44d1946cd9c3ddb9ecc99b6579af5c3f6b8ec0d67890392e2753f58a68a163e
BLAKE2b-256 checksum
How to use checksums
ad0a4b58241a4c681f7b87613875ce1034251442ce899f9650229d4b34654404
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.11.12

Release history Release notifications | RSS feed

This release

2.0.1 This release

2 release files

2.0.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page