Feature engineering on polars and pandas dataframes for machine learning!
tubular implements pre-processing steps for tabular data commonly used in machine learning pipelines.
The transformers are compatible with scikit-learn Pipelines. Each has a transform method to apply the pre-processing step to data and a fit method to learn the relevant information from the data, if applicable.
The transformers in tubular are written in narwhals narwhals, so are agnostic between pandas and polars dataframes, and will utilise the chosen (pandas/polars) API under the hood.
There are a variety of transformers to assist with;
- capping
- dates
- imputation
- mapping
- categorical encoding
- numeric operations
Here is a simple example of applying capping to two columns;
import polars as pl
transformer = CappingTransformer(
capping_values={"a": [10, 20], "b": [1, 3]},
)
test_df = pl.DataFrame({"a": [1, 15, 18, 25], "b": [6, 2, 7, 1], "c": [1, 2, 3, 4]})
transformer.transform(test_df)
# ->
# shape: (4, 3)
# ┌─────┬─────┬─────┐
# │ a ┆ b ┆ c │
# │ --- ┆ --- ┆ --- │
# │ i64 ┆ i64 ┆ i64 │
# ╞═════╪═════╪═════╡
# │ 10 ┆ 3 ┆ 1 │
# │ 15 ┆ 2 ┆ 2 │
# │ 18 ┆ 3 ┆ 3 │
# │ 20 ┆ 1 ┆ 4 │
# └─────┴─────┴─────┘
Tubular also supports saving/reading transformers and pipelines to/from json format (goodbye .pkls!), which we demo below:
import polars as pl
from tubular.imputers import MeanImputer, MedianImputer
from sklearn.pipeline import Pipeline
from tubular.pipeline import dump_pipeline_to_json, load_pipeline_from_json
# Create a simple dataframe
df = pl.DataFrame({"a": [1, 5], "b": [10, None]})
# Add imputers
median_imputer = MedianImputer(columns=["b"])
mean_imputer = MeanImputer(columns=["b"])
# Create and fit the pipeline
original_pipeline = Pipeline(
[("MedianImputer", median_imputer), ("MeanImputer", mean_imputer)]
)
original_pipeline = original_pipeline.fit(df)
# Dumping the pipeline to JSON
pipeline_json = dump_pipeline_to_json(original_pipeline)
pipeline_json
# Printed value:
# ->
# {
# 'MedianImputer': {
# 'tubular_version': '2.6.1',
# 'classname': 'MedianImputer',
# 'init': {
# 'columns': ['b'],
# 'copy': False,
# 'verbose': False,
# 'return_native': True,
# 'weights_column': None
# },
# 'fit': {
# 'impute_values_': {'b': 10.0}
# }
# },
# 'MeanImputer': {
# 'tubular_version': '2.6.1',
# 'classname': 'MeanImputer',
# 'init': {
# 'columns': ['b'],
# 'copy': False,
# 'verbose': False,
# 'return_native': True,
# 'weights_column': None
# },
# 'fit': {
# 'impute_values_': {
# 'b': 10.0
# }
# }
# }
# Load the pipeline from JSON
pipeline = load_pipeline_from_json(pipeline_json)
# Verify the reconstructed pipeline
print(pipeline)
# Printed value:
# Pipeline(steps=[('MedianImputer', MedianImputer(columns=['b'])),
# ('MeanImputer', MeanImputer(columns=['b']))])
We are currently in the process of rolling out support for polars lazyframes!
track our progress below:
| polars_compatible | pandas_compatible | jsonable | lazyframe_compatible | |
|---|---|---|---|---|
| AggregateColumnsOverRowTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| AggregateRowsOverColumnTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| ArbitraryImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| BetweenDatesTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| BooleanImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| CappingTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| CategoricalImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| ColumnDtypeSetter | ✔️ | ✔️ | ✔️ | ✔️ |
| CompareTwoColumnsTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| DateDifferenceTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| DatetimeComponentExtractor | ✔️ | ✔️ | ✔️ | ✔️ |
| DatetimeInfoExtractor | ✔️ | ✔️ | ✔️ | ✔️ |
| DatetimeSinusoidCalculator | ✔️ | ✔️ | ✔️ | ✔️ |
| DifferenceTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| ExtractStringComponentsTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| GroupRareLevelsTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| LowerCaseTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| MappingTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| MeanImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| MeanResponseTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| MedianImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| ModeImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| NullIndicator | ✔️ | ✔️ | ✔️ | ✔️ |
| NumberImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| OneDKmeansTransformer | ✔️ | ✔️ | ✔️ | ❌ |
| OneHotEncodingTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| OutOfRangeNullTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| RatioTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| RemoveCharactersTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| RenameColumnsTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| SetValueTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| StringContainsTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| StringImputer | ✔️ | ✔️ | ✔️ | ✔️ |
| ToDatetimeTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
| WhenThenOtherwiseTransformer | ✔️ | ✔️ | ✔️ | ✔️ |
Installation
The easiest way to get tubular is directly from pypi with;
pip install tubular
Documentation
The documentation for tubular can be found on readthedocs.
Instructions for building the docs locally can be found in docs/README.
Examples
We utilise doctest to keep valid usage examples in the docstrings of transformers in the package, so please see these for getting started!
Issues
For bugs and feature requests please open an issue.
Build and test
The test framework we are using for this project is pytest. To build the package locally and run the tests follow the steps below.
First clone the repo and move to the root directory;
git clone https://github.com/azukds/tubular.git
cd tubular
Next install tubular and development dependencies;
pip install . -r requirements-dev.txt
Finally run the test suite with pytest;
pytest
Contribute
tubular is under active development, we're super excited if you're interested in contributing!
See the CONTRIBUTING file for the full details of our working practices.
Metadata
Release files for tubular 4.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tubular-4.1.0.tar.gz | 254.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tubular-4.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 344.2 kB
Release files / tubular-4.1.0.tar.gz
| Download URL | tubular-4.1.0.tar.gz |
|---|---|
| Size | 254.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
167f3ac38836e6ba71d5caa1a79945ced7f2c63b2f6d560aebb0776dff3a17a2
|
|
BLAKE2b-256 checksum How to use checksums |
18f396105b3b2bb9e4a082f5b1cbdd576d98e3978769dd99bdafb442fe8d4dad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / tubular-4.1.0-py3-none-any.whl
| Download URL | tubular-4.1.0-py3-none-any.whl |
|---|---|
| Size | 89.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
483b19f00bf8ee76420f9eb0e4d2100eff9f888a0447af75d20cd7d3d62db81f
|
|
BLAKE2b-256 checksum How to use checksums |
ba2cb7cd02302d8f1b71ab47ff89db70409e1b3cd2aa3ac636d04b3b6e302dbe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log