Skip to main content

High-performance extension package for pandas with fast groupby operations

Project description

groupby-lib

A high-performance extension package for pandas that provides fast groupby operations using NumPy and Numba acceleration intended for large data sets. Performance improvements increase as data is scaled up (see benchmarking below). It also provides a more flexible API and convenience methods in the group-by space like adding margins and calculating ratios.

Overview

groupby-lib enhances Pandas groupby functionality with optimized implementations that leverage NumPy arrays and Numba's just-in-time compilation for significant performance improvements. The package is designed to work seamlessly with pandas DataFrames and Series while providing additional flexibility for various array types.

Faster GroupBy Operations

Optimized group-by operations, particularly with categorical data (uses the existing factorization) and with multi-threading on large datasets across both row and columns axes. NB: the benchmarking below is after all numba just-in-time compilations have been completed. alt text

Inline Filtering of Groupby Operations

Inline filtering such that Series/DataFrames do not have to be filtered before group-by operations. This saves time and memory directly, and also promotes re-use of GroupBy objects which boosts performance dramatically (the majority of run-time in most group-by ops is in the factorization step)

Pass a Boolean mask, slice or fancy-indexer to the group-by method call: alt text

  • Flexible Input Types: Support for NumPy arrays, pandas Series/DataFrames backed by NumPy or Arrow, and Polars Series/DataFrames. Here are some examples which are not exhaustive
from groupby_lib.groupby import GroupBy

arr = np.random.randint(0, 10, 10000)
keys = arr % 3 + 2

# Create GroupBy object
gb = GroupBy(keys)

# Perform aggregations
gb.agg(arr, agg_func=["sum", "mean", "count", "min", "max"])
     sum      mean  count  min  max
2  18054  4.489928   4021    0    9
3  12085  4.035058   2995    1    7
4  14749  4.942694   2984    2    8

# on dict of arrays / Series
gb.std(dict(x=arr, x_squared=arr ** 2, pd_series=pd.Series(arr)))
           x	x_squared	pd_series
2	3.351510	31.601683	3.351510
3	2.452965	20.108977	2.452965
4	2.429882	24.677406	2.429882

# build a GroupBy from a dict/list of keys or a DataFrame
key_list = [keys, pd.Series(["A", "B"]).repeat(5000)]
GroupBy(key_list).size()
3  A    1537
   B    1497
2  A    2027
   B    1990
4  A    1436
   B    1513
  • Pandas Compatible: Results are returned as pandas Series/DataFrames
  • Multi-threading: Automatic parallelization for large datasets across both row and column axes

Installation

# Install dependencies
conda install groupby-lib

The GroupBy Class

The core of groupby-lib is the GroupBy class located in groupby_lib.groupby.core. This class provides efficient groupby operations with a pandas-like API.

Basic Usage

Supported Aggregation Methods

The GroupBy class supports various aggregation/selection functions:

  • sum()/mean() - Sum/mean of values in each group
  • min()/max() - Min/Max value in each group
  • var()/std() - Variance/Std. Dev. of values in each group
  • count() - Count of non-null values in each group
  • size() - Total count of values in each group (including nulls)
  • median() - Median value in each group
  • head()/tail()/ - The top/bottom N rows/elements in each group
  • first()/last() - The first/last non-null elements in each group
  • nth() - The nth row/element in each group
  • shift() - Shift elements group-wise
  • diff() - Diff elements group-wise
  • rolling_[min()/max()/sum()/mean()] - rolling aggregations
  • cum[min()/max()/count()/sum()] - cumulative aggregations

Working with Pandas Data

# Using pandas Series
df = pd.DataFrame({
    'category': ['A', 'B', 'A', 'C', 'B', 'A'],
    'value1': [1, 2, 3, 4, 5, 6],
    'value2': [10, 20, 30, 40, 50, 60]
})

# Group by category
gb = GroupBy(df['category'])

# Aggregate single column
mean_value1 = gb.mean(df['value1'])
print(mean_value1)
# Output:
# A    3.333333
# B    3.500000
# C    4.000000
# Name: value1, dtype: float64

# Aggregate multiple columns
result = gb.sum(df[['value1', 'value2']])
print(result)
# Output:
#    value1  value2
# A      10     100
# B       7      70
# C       4      40

Multi-level Grouping

# Group by multiple keys
keys1 = ['A', 'A', 'B', 'B', 'A', 'B']
keys2 = ['X', 'Y', 'X', 'Y', 'X', 'Y']
values = [1, 2, 3, 4, 5, 6]

gb = GroupBy([keys1, keys2])
result = gb.sum(values)
print(result)
# Output:
# A  X    6
#    Y    2
# B  X    3
#    Y   10
# dtype: int64

Using Masks for Filtering

# Apply mask to filter data during aggregation
keys = ['A', 'B', 'A', 'C', 'B', 'A']
values = [1, 2, 3, 4, 5, 6]
mask = np.array([True, False, True, True, False, True])

gb = GroupBy(keys)
result = gb.sum(values, mask=mask)
print(result)
# Output: Only includes values where mask is True
# A    10  # (1 + 3 + 6)
# C     4  # (4)
# dtype: int64

Transform Operations

Transform operations return results with the same shape as the input:

keys = ['A', 'B', 'A', 'B', 'A']
values = [1, 2, 3, 4, 5]

gb = GroupBy(keys)
transformed = gb.sum(values, transform=True)
print(transformed)
# Output: Each position shows the group sum
# 0    9  # Sum of group A
# 1    6  # Sum of group B  
# 2    9  # Sum of group A
# 3    6  # Sum of group B
# 4    9  # Sum of group A
# dtype: int64

Generic Aggregation with agg()

# Single aggregation function
gb = GroupBy(['A', 'B', 'A', 'B'])
result = gb.agg([1, 2, 3, 4], 'mean')

# Multiple aggregation functions
result = gb.agg([1, 2, 3, 4], ['sum', 'mean', 'min', 'max'])
print(result)
# Output:
#    sum  mean  min  max
# A    4   2.0    1    3
# B    6   3.0    2    4

Ratio Calculations

# Calculate ratios between two value sets
numerator = [1, 2, 3, 4]
denominator = [10, 20, 30, 40]
keys = ['A', 'B', 'A', 'B']

gb = GroupBy(keys)
ratio = gb.ratio(numerator, denominator)
print(ratio)
# Output:
# A    0.133333  # (1+3)/(10+30)
# B    0.100000  # (2+4)/(20+40)
# dtype: float64

Head Operation

Get the first n rows from each group:

keys = ['A', 'A', 'B', 'B', 'A', 'B']
values = [1, 2, 3, 4, 5, 6]

gb = GroupBy(keys)
first_two = gb.head(values, n=2)
print(first_two)
# Output: First 2 values from each group
# A  0    1
#    1    2
# B  0    3
#    1    4
# dtype: int64

Pivot Tables

Create pivot tables using the crosstab function:

from groupby_lib.groupby import crosstab

# Sample data
index_keys = ['Jan', 'Feb', 'Jan', 'Feb']
column_keys = ['A', 'A', 'B', 'B']  
values = [10, 20, 30, 40]

result = crosstab(
    index=index_keys,
    columns=column_keys,
    values=values,
    agg_func='sum'
)
print(result)
# Output:
#      A   B
# Feb  20  40
# Jan  10  30

Performance Benefits

groupby-lib provides significant performance improvements for large datasets:

  • Numba Acceleration: JIT compilation for fast numerical operations
  • Multi-threading: Automatic parallelization for datasets > 1M rows
  • Memory Efficiency: Optimized memory usage patterns
  • Reduced Overhead: Minimal pandas overhead for core operations

Dependencies

  • Python 3.10+
  • NumPy
  • pandas
  • numba
  • pyarrow
  • polars
  • scipy (optional for GroupScatter support)

Development

See CLAUDE.md for development guidelines including:

  • Code style conventions
  • Testing requirements
  • Build and lint commands
  • Contribution guidelines

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

groupby_lib-0.1.0.tar.gz (902.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

groupby_lib-0.1.0-py3-none-any.whl (58.8 kB view details)

Uploaded Python 3

File details

Details for the file groupby_lib-0.1.0.tar.gz.

File metadata

  • Download URL: groupby_lib-0.1.0.tar.gz
  • Upload date:
  • Size: 902.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.10

File hashes

Hashes for groupby_lib-0.1.0.tar.gz
Algorithm Hash digest
SHA256 e4d3a9d610e7ab45396f33da10d07e838a1cfe1a4d5229f3fdbfef0c59ade21c
MD5 da7112d234abb241fa5e1932b8ae6f9e
BLAKE2b-256 fb532648b3779127aba259d1850a2279c8525326533839ac724d706d6b1d419d

See more details on using hashes here.

File details

Details for the file groupby_lib-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: groupby_lib-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 58.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.10

File hashes

Hashes for groupby_lib-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 034e0bbbe91efb46077fdf79b64c39f7cee9ab73d99bb0b72755f8100b05b64b
MD5 f7f711ade55b9ab7b42e0efa71590a9b
BLAKE2b-256 a5b8bc1c73c8aab9df4047247165c9792e8ffb2d0124ef6221414e322dc67b75

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page