Skip to main content

Sparkleframe

Project description

SparkleFrame

SparkleFrame implements the PySpark DataFrame API in order to enable running transformation pipelines directly on Polars Dataframe - no Spark clusters or dependencies required.

Apache Spark is designed for distributed, large-scale data processing, but it is not optimized for low-latency use cases. There are scenarios, however, where you need to quickly re-compute certain data—for example, regenerating features for a machine learning model in real time or near-real time.

SparkleFrame is great for:

  • Users who want to run PySpark code quickly locally without the overhead of starting a Spark session
  • Users who want to run PySpark DataFrame code without the complexity of using Spark for processing
  • Useful for unit testing, feature prototyping, or serving small pipelines in microservices.

You can learn more about the design motivation behind Sparkleframe in this discussion thread.

Documentation

Full documentation is available at https://flypipe.github.io/sparkleframe/.

Installation

pip install sparkleframe

Usage

SparkleFrame can be used in two ways:

  • Directly importing the sparkleframe.polarsdf package
  • Using the activate function to allow for continuing to use pyspark.sql but have it use SparkleFrame behind the scenes.

Directly importing

If converting a PySpark pipeline, all pyspark.sql should be replaced with sparkleframe.polarsdf.

# PySpark import
# from pyspark.sql import SparkSession
# from pyspark.sql import functions as F
# from pyspark.sql.dataframe import DataFrame
# SparkleFrame import
from sparkleframe.polarsdf.session import SparkSession
from sparkleframe.polarsdf import functions as F
from sparkleframe.polarsdf.dataframe import DataFrame

Activating SparkleFrame

SparkleFrame can either replace pyspark imports or be used alongside them. To replace pyspark imports, use the activate function to set the engine to use.

from sparkleframe.activate import activate

# Activate SparkleFrame
activate()

from pyspark.sql import SparkSession
session = SparkSession.builder.getOrCreate()

SparkSession will now be a SparkleFrame Session object and everything will be run on Polars Dataframe directly.

SparkleFrame can also be directly imported which both maintains pyspark imports:

from sparkleframe.polarsdf.session import SparkSession
session = SparkSession.builder.getOrCreate()

Example Usage

from sparkleframe.activate import activate

# Activate SparkleFrame
activate()

from pyspark.sql import SparkSession
from pyspark.sql import functions as F

session = SparkSession.builder.getOrCreate()
df = session.createDataFrame(data=[{"col1": 1, "col2": 2}])
df = df.withColumn("col3", F.col("col2") + F.col("col1"))
>>> print(type(df))
<class 'sparkleframe.polarsdf.dataframe.DataFrame'>
>>> df.show()
shape: (1, 3)
┌──────┬──────┬──────┐
 col1  col2  col3 
 ---   ---   ---  
 i64   i64   i64  
╞══════╪══════╪══════╡
 1     2     3    
└──────┴──────┴──────┘

!!! note

If you encounter any transformation that is not implemented, please open an [issue on GitHub](https://github.com/flypipe/sparkleframe/issues/new) so it can be prioritized.

Source Code

API code is available at https://github.com/flypipe/sparkleframe.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sparkleframe-0.3.0.tar.gz (694.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sparkleframe-0.3.0-py3-none-any.whl (43.6 kB view details)

Uploaded Python 3

File details

Details for the file sparkleframe-0.3.0.tar.gz.

File metadata

  • Download URL: sparkleframe-0.3.0.tar.gz
  • Upload date:
  • Size: 694.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: python-requests/2.32.3

File hashes

Hashes for sparkleframe-0.3.0.tar.gz
Algorithm Hash digest
SHA256 d56946729fe4443e847832d55cfcf5c4792e85ef1b1618fb2a2b596466283f38
MD5 fde8a7bd6d3d0c301e629bfde3eeb38f
BLAKE2b-256 83996647e89939a0cd013d26a6a9fa06d2bda2052fab5f6039b348cace5d07fa

See more details on using hashes here.

File details

Details for the file sparkleframe-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: sparkleframe-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 43.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: python-requests/2.32.3

File hashes

Hashes for sparkleframe-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 9410d4239bfa4880dd2567790ae3b0c0620d6d8ccb27f5e89028c1f949a636db
MD5 ba94600e5dd87b55d3c17d43fcd1391a
BLAKE2b-256 5ac9bda06ea212f45045f9c58ad70444e3d5e984b580e8696669b99e87191dbf

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page