A Python implementation of R's tidyplots for creating publication-ready plots
Project description
Tidyplots for Python
A Python library for creating publication-ready plots with a fluent, chainable interface, inspired by tidyplots R version.
Built on top of plotnine, it provides a pandas DataFrame extension method for easy and intuitive plot creation.
todo list
- complete the whole set of plotting that maps to R version
- more themes, such as ggprism
- more palettes, including palettes from famous journals, for instance, ggsci
- more plotting variations from ggpubr
Features
- Fluent, chainable interface using pandas DataFrame extension
- Publication-ready plots with Nature Publishing Group (NPG) color palette by default
- Comprehensive set of plot types and statistical visualizations
- Flexible faceting with
split_byparameter for both single and multi-variable splits - Statistical annotations (p-values, correlations, etc.)
- Multiple scientific journal color palettes (NPG, AAAS, NEJM, etc.)
- Easy customization of colors, labels, and themes
- Prism-style publication themes
Installation
pip install tidyplots-python
for development version:
pip install git+https://github.com/JNU-Tangyin/tidyplots-python.git
Quick Start
import pandas as pd
import numpy as np
import seaborn as sns
from tidyplots import TidyPlot
# Create sample data
iris = sns.load_dataset("iris")
# Create a scatter plot with groups
(iris.tidyplot(x='sepal_length', y='sepal_width', fill='species')
.add_scatter(size=5, alpha=0.5)
.add_density_2d(alpha=0.1)
.adjust_labels(title='Customization Test',
x='Sepal Length', y='Sepal Width')
.adjust_colors(['#1f77b4', '#ff7f0e', '#2ca02c'])
.adjust_legend_position('right')
.show())
API Reference
Core Plot Creation
-
df.tidyplot(x, y=None, color=None): Initialize a plot with aestheticsdf.tidyplot(x='column1', y='column2', color='group')
Plot Types
.add_scatter(alpha=0.6, size=3): Add scatter points.add_line(alpha=0.8): Add line plot.add_boxplot(alpha=0.3): Add box plot.add_violin(draw_quantiles=[0.25, 0.5, 0.75]): Add violin plot.add_density(alpha=0.5): Add density plot.add_density_2d(): Add 2D density contour plot.add_bar(): Add bar plot.add_errorbar(ymin, ymax): Add error bars.add_hex(bins=20): Add hexbin plot.add_data_points(alpha=0.3): Add jittered data points
Statistical Features
.add_smooth(method='lm'): Add smoothing line.add_correlation_text(): Add correlation coefficient.add_pvalue(p_value, x1, x2, y): Add p-value annotation
Customization
.adjust_labels(title=None, x=None, y=None): Set plot labels.adjust_colors(palette): Change color palette- Available palettes: 'npg' (default), 'aaas', 'nejm', 'lancet', 'jama', 'd3', 'material', 'igv'
.scale_color_gradient(low, high): Set color gradient.adjust_axis_text_angle(angle): Rotate axis text
Themes
-
Default theme: Prism-style with NPG colors
-
Customizable base theme:
.adjust_theme(base_size=11, base_family="Arial")
Examples
Basic Plots
import pandas as pd
import seaborn as sns
from tidyplots import TidyPlot
# Load sample data
iris = sns.load_dataset("iris")
# Simple scatter plot with groups
(iris.tidyplot(x='sepal_length', y='sepal_width', fill='species')
.add_scatter()
.show())
# Faceted plot using single variable
(iris.tidyplot(x='sepal_length', y='sepal_width', fill='species', split_by='species')
.add_scatter()
.show())
# Grid faceted plot using two variables
(iris.tidyplot(x='sepal_length', y='sepal_width', fill='species',
split_by=['species', 'petal_size_category']) # petal_size_category is an example
.add_scatter()
.show())
Scatter Plot
import seaborn as sns
from tidyplots import TidyPlot
# Load iris dataset
iris = sns.load_dataset('iris')
# Create a scatter plot
(iris.tidyplot(x='sepal_length', y='sepal_width', fill='species')
.add_scatter(size=5, alpha=0.7)
.adjust_labels(title='Iris Sepal Dimensions',
x='Sepal Length', y='Sepal Width'))
Box Plot
# Create a box plot
(iris.tidyplot(x='species', y='sepal_length', fill='species')
.add_boxplot(alpha=0.5)
.add_jitter(width=0.2, size=3, alpha=0.7)
.adjust_labels(title='Sepal Length by Species',
x='Species', y='Sepal Length'))
Violin Plot
# Create a violin plot
(iris.tidyplot(x='species', y='petal_length', fill='species')
.add_violin(draw_quantiles=[0.25, 0.5, 0.75], alpha=0.6)
.adjust_labels(title='Petal Length Distribution',
x='Species', y='Petal Length'))
Density Plot
# Create a density plot
(iris.tidyplot(x='sepal_width', fill='species')
.add_density(alpha=0.3)
.adjust_labels(title='Sepal Width Density',
x='Sepal Width', y='Density'))
Line Plot
# Load flights dataset
flights = sns.load_dataset('flights')
# Create a line plot
(flights.tidyplot(x='year', y='passengers', fill='month')
.add_line(size=2)
.adjust_labels(title='Passenger Trends',
x='Year', y='Number of Passengers'))
Bar Plot
# Create a bar plot
tips = sns.load_dataset('tips')
# Average tip by day
(tips.tidyplot(x='day', y='tip', fill='sex')
.add_bar(stat='mean')
.adjust_labels(title='Average Tip by Day',
x='Day', y='Average Tip'))
Dot Plot
# Create a dot plot
(iris.tidyplot(x='species', y='sepal_length', fill='species')
.add_dotplot()
.adjust_labels(title='Sepal Length Dot Plot',
x='Species', y='Sepal Length'))
Mean Bar Plot
# Mean bar plot with error bars
(tips.tidyplot(x='day', y='total_bill', fill='sex')
.add_mean_bar(alpha=0.5)
.add_sem_errorbar()
.adjust_labels(title='Mean Total Bill by Day',
x='Day', y='Total Bill'))
Curve Fitting
# Scatter plot with curve fitting
(tips.tidyplot(x='total_bill', y='tip', fill='sex')
.add_scatter()
.add_smooth(method='lm')
.adjust_labels(title='Tip vs Total Bill',
x='Total Bill', y='Tip'))
Ribbon Plot (SEM)
# Ribbon plot showing standard error
(tips.groupby('day')['total_bill']
.apply(lambda x: pd.DataFrame({
'mean': x.mean(),
'sem': x.sem()
}))
.reset_index()
.tidyplot(x='day', y='mean', fill='sex')
.add_line(color='cyan')
.add_ribbon(ymin='mean - sem', ymax='mean + sem', alpha=0.3)
.adjust_labels(title='Total Bill with SEM',
x='Day', y='Total Bill'))
Hexbin Plot
# Hexbin plot for density visualization
(tips.tidyplot(x='total_bill', y='tip', fill='sex')
.add_hex(bins=20)
.adjust_labels(title='Total Bill vs Tip Density',
x='Total Bill', y='Tip'))
Count Plot
# Count plot
titanic = sns.load_dataset('titanic')
(titanic.tidyplot(x='class', fill='survived')
.add_bar(stat='count')
.adjust_labels(title='Passenger Count by Class and Survival',
x='Class', y='Count'))
Pie Charts
TidyPlots supports creating customizable pie charts using the add_pie() method. You can customize colors, explode slices, add shadows, and format percentages.
Basic Pie Chart
# Create a simple pie chart
titanic = sns.load_dataset("titanic")
survival_counts = titanic['survived'].value_counts().reset_index()
survival_counts.columns = ['Status', 'Count']
survival_counts['Status'] = survival_counts['Status'].map({0: 'Did Not Survive', 1: 'Survived'})
(survival_counts.tidyplot(x='Status', y='Count')
.add_pie(colors=['#ff9999', '#66b3ff'])
.show())
Exploded Pie Chart with Shadow
# Create pie chart with exploded slices and shadow effect
class_counts = titanic['class'].value_counts().reset_index()
class_counts.columns = ['Class', 'Count']
(class_counts.tidyplot(x='Class', y='Count')
.add_pie(colors=['#ff9999', '#66b3ff', '#99ff99'],
explode=(0.1, 0, 0), # Explode first slice
shadow=True) # Add shadow effect
.show())
Custom Percentage Formatting
# Show both percentages and counts
tips = sns.load_dataset("tips")
day_counts = tips['day'].value_counts().reset_index()
day_counts.columns = ['Day', 'Count']
(day_counts.tidyplot(x='Day', y='Count')
.add_pie(colors=['#ff9999', '#66b3ff', '#99ff99', '#ffcc99'],
autopct=lambda pct: f'{pct:.1f}%\n({int(pct/100.*sum(day_counts["Count"])):.0f})',
startangle=45)
.show())
Advanced Label Positioning
# Adjust label positions and add white edges
diamonds = sns.load_dataset("diamonds")
cut_counts = diamonds['cut'].value_counts().reset_index()
cut_counts.columns = ['Cut', 'Count']
(cut_counts.tidyplot(x='Cut', y='Count')
.add_pie(colors=['#ff9999', '#66b3ff', '#99ff99', '#ffcc99', '#ff99cc'],
pctdistance=0.85, # Move percentage labels outward
labeldistance=1.1, # Move labels outward
wedgeprops={'linewidth': 3, 'edgecolor': 'white'}) # Add white edges
.show())
The add_pie() method supports various customization options:
colors: List of colors for pie slicesexplode: Tuple of floats to offset each sliceautopct: String or function for percentage display (default: '%1.1f%%')pctdistance: Float for percentage label distance from centerlabeldistance: Float for label distance from centershadow: Bool to add shadow effectstartangle: Int for starting angle in degrees (default: 90)wedgeprops: Dict of properties for pie wedges
Note: Pie charts do not support faceting with split_by. For faceted visualizations, consider using bar plots or other plot types.
Faceting
The split_by parameter allows you to create faceted plots in two ways:
- Single variable faceting using
facet_wrap - Two-variable faceting using
facet_grid
Single Variable Faceting (facet_wrap)
# Create faceted scatter plot by species
(iris.tidyplot(x='sepal_length', y='sepal_width', split_by='species', fill='species')
.add_scatter(alpha=0.6)
.adjust_labels(title='Iris Measurements by Species',
x='Sepal Length', y='Sepal Width')
.show())
# Create faceted violin plot by island
penguins = sns.load_dataset("penguins")
(penguins.tidyplot(x='species', y='body_mass_g', split_by='island', fill='species')
.add_violin(alpha=0.7)
.adjust_labels(title='Penguin Body Mass by Island',
x='Species', y='Body Mass (g)')
.show())
Two-Variable Faceting (facet_grid)
# Create faceted scatter plot by day and time
tips = sns.load_dataset("tips")
(tips.tidyplot(x='total_bill', y='tip', split_by=['day', 'time'], fill='smoker')
.add_scatter(alpha=0.6)
.adjust_labels(title='Tips by Day and Time',
x='Total Bill', y='Tip')
.show())
# Create faceted boxplot by color and clarity
diamonds = sns.load_dataset("diamonds")
diamonds_subset = diamonds.sample(n=1000, random_state=42)
(diamonds_subset.tidyplot(x='cut', y='price', split_by=['color', 'clarity'], fill='color')
.add_boxplot(alpha=0.7)
.adjust_labels(title='Diamond Prices by Cut, Color, and Clarity',
x='Cut', y='Price')
.show())
Bar Plot with Faceting
# Create faceted bar plot showing survival counts
titanic = sns.load_dataset("titanic")
survival_data = titanic.groupby(['class', 'sex', 'survived']).size().reset_index(name='count')
(survival_data.tidyplot(x='class', y='count', fill='survived', split_by='sex')
.add_bar(position='dodge', alpha=0.7)
.adjust_labels(title='Titanic Survival by Class and Sex',
x='Class', y='Count')
.show())
Color Palettes
Default color palettes from scientific journals:
# Change color palette to Nature Publishing Group (default)
df.tidyplot(...).adjust_colors('npg') #
Available palettes (thanks to ggsci):
- 'npg': Nature Publishing Group colors (default)
- 'aaas': Science/AAAS colors
- 'nejm': New England Journal of Medicine colors
- 'lancet': The Lancet colors
- 'jama': Journal of American Medical Association colors
- 'd3': D3.js colors
- 'material': Material Design colors
- 'igv': Integrative Genomics Viewer colors
and it supports the all color schemes in matplotlib by name:
cmaps = [
('Perceptually Uniform Sequential', ['viridis', 'plasma', 'inferno', 'magma']),
('Sequential', [ 'Greys', 'Purples', 'Blues', 'Greens', 'Oranges', 'Reds', 'YlOrBr', 'YlOrRd', 'OrRd', 'PuRd', 'RdPu', 'BuPu', 'GnBu', 'PuBu', 'YlGnBu', 'PuBuGn', 'BuGn', 'YlGn']),
('Sequential (2)', ['binary', 'gist_yarg', 'gist_gray', 'gray', 'bone', 'pink', 'spring', 'summer', 'autumn', 'winter', 'cool', 'Wistia','hot', 'afmhot', 'gist_heat', 'copper']),
('Diverging', ['PiYG', 'PRGn', 'BrBG', 'PuOr', 'RdGy', 'RdBu', 'RdYlBu', 'RdYlGn', 'Spectral', 'coolwarm', 'bwr', 'seismic']), ('Qualitative', [ 'Pastel1', 'Pastel2', 'Paired', 'Accent', 'Dark2', 'Set1', 'Set2', 'Set3', 'tab10', 'tab20', 'tab20b', 'tab20c']),
('Miscellaneous', [ 'flag', 'prism', 'ocean', 'gist_earth', 'terrain', 'gist_stern', 'gnuplot', 'gnuplot2', 'CMRmap', 'cubehelix', 'brg', 'hsv', 'gist_rainbow', 'rainbow', 'jet', 'nipy_spectral', 'gist_ncar'])
]
# Change color palette using Matplotlib colormaps
df.tidyplot(...).adjust_colors('Set1')
more to come ...
Dependencies
- pandas >= 1.0.0
- numpy >= 1.18.0
- plotnine >= 0.8.0
Notices
This project is a Python transplant from tidyplots R version, and it is supported by Windsurf.
Althought I believe the philosophy we present in this work is convenient, pythonic, and easy to use, it is however rudimentary starting point, therefore it may contain bugs and missing features. Please forgive me if you find any.
And most importantly, contributions are more than welcome! Please feel free to submit a Pull Request.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Missing functions
Please note that not all the functions are available because of:
- plotnine being a younger library with fewer contributors
- Some R-specific features being harder to implement in Python
- Different underlying graphics engines (R's grid vs Python's matplotlib)
The list is provided as below:
- Removed Functions :
geom_hex()- Hexagonal binning not available in plotninegeom_text_repel()- Text repulsion for avoiding overlaps not availablecoord_polar()- Polar coordinate system not fully supportedstat_density_2d_filled()- Filled 2D density plots not availablegeom_raster()- High-performance raster rendering not availablegeom_sf()- Simple features (geographic data) not supportedgeom_spoke()- Spoke plots not availablestat_ellipse()- Statistical ellipses not supportedgeom_label()- Labels with backgrounds not availablegeom_curve()- Curved line segments not available
- Modified Functions :
add_data_labels_repel()- Modified to use regulargeom_text()instead ofgeom_text_repel()add_pie_chart()- Implemented using bar plots and coordinate transformations since native pie charts aren't supportedadd_donut_chart()- Similar to pie charts, implemented through workaroundsadd_density_2d_filled()- Had to use regular density_2d with modified aesthetics
- Limited Functionality :
facet_grid()- More limited options compared to ggplot2scale_*_gradient2()- Diverging color scales have limited optionstheme()- Some theme elements and customizations not availablecoord_fixed()- Fixed coordinate ratio support is limitedposition_dodge2()- Advanced dodging features not available
- Performance Differences :
- Large dataset handling is generally slower in plotnine
- Some smoothing methods in
stat_smooth()have fewer options - Rendering of complex plots with many layers is less optimized
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tidyplots_python-0.1.0.tar.gz.
File metadata
- Download URL: tidyplots_python-0.1.0.tar.gz
- Upload date:
- Size: 28.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.0.1 CPython/3.11.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e3c286585e3a3e06903a4944bca6a011af0eb84c1f06f3875af50f2a2b11e834
|
|
| MD5 |
79f76fa2e847502b4fc72763fab6fda7
|
|
| BLAKE2b-256 |
151aa30bef2e030090523b492231bb518fa7f0f1b47851261234af0b6cadc913
|
File details
Details for the file tidyplots_python-0.1.0-py3-none-any.whl.
File metadata
- Download URL: tidyplots_python-0.1.0-py3-none-any.whl
- Upload date:
- Size: 23.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.0.1 CPython/3.11.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
42e6e3c0d50122f0acf48afed7f93993f7ae950e7295e0f26a78df742b358732
|
|
| MD5 |
72907a6e7ce843b439eadcb1b3f705dc
|
|
| BLAKE2b-256 |
be494200e2e174d43dbc123bf8120bb78f4effafa99d8bdd64002b59f58c6f02
|