viper
Simple, expressive pipeline syntax to transform and manipulate data with ease
Overview
viper is a Python package that provides a simple, expressive way to work with data. It allows you to easily manipulate and transform data using a pipeline syntax similar to that of dplyr.
Pipelining your DataFrame manipulation operations offers several benefits:
- improved code readability (no need to 'comment the what')
- no need to save intermediate dataframes
- ability to chain a long sequence of operations in a single command
- thinking of coding as a series of transformations between the input and the desired output can improve the design and make it less coupled
Docs
Complete documentation and reference are available on the package's site.
Quick Start
Installation:
pip install viper-df
Here is an example of how to use viper to analyze the famed mtcars dataset.
We want to find:
- the average consumption, expressed in Miles/(US) gallon
- the average power
Furthermore:
- only consider those cars that weigh more than 2000lbs
- group the results by the number of cylinders and number of gears
- arrange in descending orders by the grouping variables
import viper as v
from viper.data import mtcars
v.pipeline(
mtcars,
v.rename(
"hp = power",
"mpg = consumption",
),
v.mutate(
consumption=lambda r: 1 / r["consumption"]
),
v.filter(
lambda r: r["wt"] > 2
),
v.group_by("cyl", "gear"),
v.summarize(
"power = mean()",
"consumption = mean()"
),
v.arrange(
"cyl desc",
"gear desc"
),
)
# power consumption
# cyl gear
# 8 5 299.500000 0.064979
# 3 194.166667 0.068824
# 6 5 175.000000 0.050761
# 4 116.500000 0.050875
# 3 107.500000 0.050989
# 4 5 91.000000 0.038462
# 4 85.000000 0.041259
# 3 97.000000 0.046512
Here you can find more examples, particularly on joins.
Roadmap
The future development of the package will probably focus on:
- adding
pivot_longerandpivot_widerfunctions - adding more
join_*functions
Contributions
You are welcome to contribute to the project or open issues if you have any ideas.
Release files for viper-df 0.0.7
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| viper_df-0.0.7.tar.gz | 6.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| viper_df-0.0.7-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 14.9 kB
Release files / viper_df-0.0.7.tar.gz
| Download URL | viper_df-0.0.7.tar.gz |
|---|---|
| Size | 6.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2006a101abfbb49abbff371f8f7ffdfa4737a9908cd454b30e26c09892447741
|
|
BLAKE2b-256 checksum How to use checksums |
d14d32961c648b1d9c3951c401bd621018d074c3952ca4508ae320d46ad896b0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.8.10
|
Release files / viper_df-0.0.7-py3-none-any.whl
| Download URL | viper_df-0.0.7-py3-none-any.whl |
|---|---|
| Size | 8.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
623b116e706100be38057b23d69a980ca45dcf541c3b77fa2f6a9794dd770b26
|
|
BLAKE2b-256 checksum How to use checksums |
3cbcca27080425e31df008b8fed2cbcb44ac4572589c773b278ded2b2ec1e587
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.8.10
|