Skip to main content

scikit-play

Rethinking machine learning pipelines a bit.

What does scikit-play do?

I was wondering if there might be an easier way to construct scikit-learn pipelines. Don't get me wrong, scikit-learn is amazing when you want elaborate pipelines (exhibit A, exhibit B) but maybe there is also a place for something more lightweight and playful. This library is all about exploring that.

Imagine that you are dealing with the titanic dataset.

import pandas as pd

df = pd.read_csv("https://calmcode.io/static/data/titanic.csv")
df.head()

Here's what the dataset looks like.

survived pclass name sex age fare sibsp parch
0 3 Braund, Mr. Owen Harris male 22 7.25 1 0
1 1 Cumings, Mrs. John Bradley (Florence Briggs Thayer) female 38 71.2833 1 0
1 3 Heikkinen, Miss. Laina female 26 7.925 0 0
1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35 53.1 1 0
0 3 Allen, Mr. William Henry male 35 8.05 0 0

The goal of this dataset is to predict who survived, so survived is the target column for a classification task. But in order to make the right predictions you would need to encode the features in the right way. So to do that, you might construct a preprocessing pipeline like this:

from sklearn.pipeline import make_union, make_pipeline
from sklearn.preprocessing import OneHotEncoder
from skrub import SelectCols

pipe = make_union(
    SelectCols(["age", "fare", "sibsp", "parch"]),
    make_pipeline(
        SelectCols(["sex", "pclass"]),
        OneHotEncoder()
    )
)

This pipeline takes the age, fare, sibsp and parch features as-is. These features are already numeric so these do not need to be changed. But the sex and pclass features are candidates to one-hot encode first. These are categorical features, so it helps to encode them as such.

The pipeline works, and it's fine, but you could wonder if this is easy. After all, you do need to know scikit-learn fairly well in order to build a pipeline this way and you may also need to appreciate Python. There's some nesting happening in here as well, so for a novice or somebody who just immediately wants to make a quick model ... there's some stuff that gets in the way. All of this is fine when you consider that scikit-learn needs to allow for elaborate pipelines ... but if you just want something dead simple ... then you may appreciate another syntax instead.

Enter skplay.

Skplay offers an API that allows you to declare the aforementioned pipeline by doing this instead:

from skplay import feats, onehot

formula = feats("age", "fare", "sibsp", "parch") + onehot("sex", "pclass")

This formula object is just an object that can accumulate components.

# This object is a scikit-learn pipeline but with operator support!
formula

skplay

It's pretty much the same pipeline as before, but it's a lot easier to go ahead and declare. You're mostly dealing with column names and how to encode them, instead of thinking about how scikit-learn constructs a pipeline.

This is what scikit-play is all about, but this is just the start of what it can do. If that sounds interest you can read more on the documentation page.

Alternative you may also explore this tool by installing it via:

uv pip install scikit-play

Metadata

Release files for scikit-play 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scikit-play 0.1.2
File Size Uploaded
scikit_play-0.1.2.tar.gz 8.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scikit-play 0.1.2
File Interpreter ABI Platform
scikit_play-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 18.1 kB

Release files / scikit_play-0.1.2.tar.gz

Download URL scikit_play-0.1.2.tar.gz
Size 8.9 kB
Tags Source
SHA-256 checksum
How to use checksums
adc117806fa3fba487cfd1eb0d49740a0c1f961c895eae1c3400f6078a1ab3e4
BLAKE2b-256 checksum
How to use checksums
0956c63e6c79d050aa661e5407c4fc46cec42cb7c3e3618853776db847681f9d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.13

Release files / scikit_play-0.1.2-py3-none-any.whl

Download URL scikit_play-0.1.2-py3-none-any.whl
Size 9.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
57fd11fc142ed36eca9e369589f7aef871669cf5c02cd9b388bfdb2da99f546d
BLAKE2b-256 checksum
How to use checksums
48732f1713bde1c4cf79e58f88dab0d4039379e84a5522e5c5bc08d03bf75cd2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.13

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page