histogrammar·PyPI

Composable histogram primitives for distributed data reduction

These details have not been verified by PyPI

Project links

repository

Project description

histogrammar is a Python package for creating histograms. histogrammar has multiple histogram types, supports numeric and categorical features, and works with Numpy arrays and Pandas and Spark dataframes. Once a histogram is filled, it’s easy to plot it, store it in JSON format (and retrieve it), or convert it to Numpy arrays for further analysis.

At its core histogrammar is a suite of data aggregation primitives designed for use in parallel processing. In the simplest case, you can use this to compute histograms, but the generality of the primitives allows much more.

Several common histogram types can be plotted in Matplotlib and Bokeh with a single method call. If Numpy or Pandas is available, histograms and other aggregators can be filled from arrays ten to a hundred times more quickly via Numpy commands, rather than Python for loops.

This Python implementation of histogrammar been tested to guarantee compatibility with its Scala implementation.

Latest Python release: v1.1.0 (Feb 2025). Latest update: Feb 2025.

References

Histogrammar is a core component of popmon, a package by ING bank that allows one to check the stability of a dataset. popmon works with both pandas and spark datasets, largely thanks to Histogrammar.

Announcements

Changes

See Changes log here.

Spark 3.X

With Spark 3.X, based on Scala 2.12 or 2.13, make sure to pick up the correct histogrammar jar files:

spark = SparkSession.builder.config("spark.jars.packages", "io.github.histogrammar:histogrammar_2.12:1.0.30,io.github.histogrammar:histogrammar-sparksql_2.12:1.0.30").getOrCreate()

For Scala 2.13, in the string above simply replace “2.12” with “2.13”.

December, 2023

Example notebooks

Tutorial	Colab link
Basic tutorial
Detailed example (featuring configuration, Apache Spark and more)
Exercises

Documentation

See histogrammar-docs for a complete introduction to histogrammar. (A bit old but still good.) There you can also find documentation about the Scala implementation of histogrammar.

Check it out

The historgrammar library requires Python 3.8+ and is pip friendly. To get started, simply do:

$ pip install histogrammar

or check out the code from our GitHub repository:

$ git clone https://github.com/histogrammar/histogrammar-python
$ pip install -e histogrammar-python

where in this example the code is installed in edit mode (option -e).

You can now use the package in Python with:

import histogrammar

Congratulations, you are now ready to use the histogrammar library!

Quick run

As a quick example, you can do:

import pandas as pd
import histogrammar as hg
from histogrammar import resources

# open synthetic data
df = pd.read_csv(resources.data('test.csv.gz'), parse_dates=['date'])
df.head()

# create a histogram, tell it to look for column 'age'
# fill the histogram with column 'age' and plot it
hist = hg.Histogram(num=100, low=0, high=100, quantity='age')
hist.fill.numpy(df)
hist.plot.matplotlib()

# generate histograms of all features in the dataframe using automatic binning
# (importing histogrammar automatically adds this functionality to a pandas or spark dataframe)
hists = df.hg_make_histograms()
print(hists.keys())

# multi-dimensional histograms are also supported. e.g. features longitude vs latitude
hists = df.hg_make_histograms(features=['longitude:latitude'])
ll = hists['longitude:latitude']
ll.plot.matplotlib()

# store histogram and retrieve it again
ll.toJsonFile('longitude_latitude.json')
ll2 = hg.Factory().fromJsonFile('longitude_latitude.json')

These examples also work with Spark dataframes (sdf):

from pyspark.sql.functions import col
hist = hg.Histogram(num=100, low=0, high=100, quantity=col('age'))
hist.fill.sparksql(sdf)

For more examples please see the example notebooks and tutorials.

Project contributors

This package was originally authored by DIANA-HEP and is now maintained by volunteers.

Contact and support

Issues & Ideas & Support: https://github.com/histogrammar/histogrammar-python/issues

Please note that histogrammar is supported only on a best-effort basis.

License

histogrammar is completely free, open-source and licensed under the Apache-2.0 license.

Project details

These details have not been verified by PyPI

Project links

repository

Release history Release notifications | RSS feed

This version

1.1.0

Feb 10, 2025

1.0.34

Dec 16, 2024

1.0.33

Dec 16, 2022

1.0.32

Sep 9, 2022

1.0.31

Aug 23, 2022

1.0.30

Jun 21, 2022

1.0.29

Jun 15, 2022

1.0.28

Jun 4, 2022

1.0.27

May 20, 2022

1.0.26

Apr 9, 2022

1.0.25

Apr 5, 2021

1.0.24

Apr 3, 2021

1.0.23

Mar 23, 2021

1.0.22

Mar 23, 2021

1.0.21

Mar 15, 2021

1.0.20

Feb 5, 2021

1.0.12

Apr 22, 2020

1.0.11

Nov 14, 2019

1.0.10

Sep 2, 2019

1.0.9

Aug 15, 2017

1.0.8

Mar 30, 2017

1.0.7

Mar 22, 2017

1.0.6

Nov 29, 2016

1.0.5

Nov 11, 2016

1.0.4

Sep 26, 2016

1.0.3

Sep 22, 2016

1.0.2

Sep 6, 2016

1.0.0

Sep 2, 2016

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

histogrammar-1.1.0.tar.gz (4.0 MB view details)

Uploaded Feb 10, 2025 Source

Built Distribution

histogrammar-1.1.0-py3-none-any.whl (200.7 kB view details)

Uploaded Feb 10, 2025 Python 3

File details

Details for the file histogrammar-1.1.0.tar.gz.

File metadata

Download URL: histogrammar-1.1.0.tar.gz
Upload date: Feb 10, 2025
Size: 4.0 MB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.1.0 CPython/3.11.10

File hashes

Hashes for histogrammar-1.1.0.tar.gz
Algorithm	Hash digest
SHA256	`3b47ed3a1336bbc6458d6b9680287621a2b3db4ac6607cb23e988e27cec0e8bc`
MD5	`76d9ed3c3ce878b394e3ae16463817ed`
BLAKE2b-256	`9770491990f1b95b1e14cacdd02bb29d64a2c2a916fab9d2338f16e15e7b84de`

See more details on using hashes here.

File details

Details for the file histogrammar-1.1.0-py3-none-any.whl.

File metadata

Download URL: histogrammar-1.1.0-py3-none-any.whl
Upload date: Feb 10, 2025
Size: 200.7 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.1.0 CPython/3.11.10

File hashes

Hashes for histogrammar-1.1.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`8170d4b128abc0dfe056ddc92ead6731cdbf8db8f4ce38120513f2f538e12831`
MD5	`ad80bec993eaff7e3798ea40276baa6e`
BLAKE2b-256	`098f3515ded54283da5613f5fc25929a68886e5ae0cdd63cabb8c9239eb53e5e`

See more details on using hashes here.

histogrammar 1.1.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

References

Announcements

Changes

Spark 3.X

Example notebooks

Documentation

Check it out

Quick run

Project contributors

Contact and support

License

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes