Skip to main content

Statistical tools package

Project description

pstatstools

License: MIT Python Version

A comprehensive Python package for statistical analysis, data visualization, and hypothesis testing with an intuitive API.

Overview

pstatstools is a user-friendly statistical toolkit that offers:

  1. Simple but powerful Sample class for descriptive statistics, hypothesis testing, and visualization
  2. Extensive probability distributions support through a flexible distribution API
  3. Advanced categorical data analysis with contingency tables and chi-square tests
  4. Non-parametric tests for when assumptions of parametric tests aren't met
  5. Error propagation utilities for measurements with uncertainties
  6. Beautiful visualizations that work seamlessly in Jupyter notebooks
  7. ANOVA and inferential statistics for experimental design analysis

Quick Start

import pstatstools as pst
import matplotlib.pyplot as plt

# Create a sample with a single line of code
sample_data = [1, 4, 3, 5, 7, 3, 3, 2, 1, 5]
s = pst.sample(sample_data)

# Get beautiful HTML-formatted descriptive statistics in Jupyter notebooks
s.describe()  # Rich output with tables for central tendency, dispersion, etc.

# Visualize your data
s.histogram(bins=8, title="My Data Distribution")
plt.show()

# Run hypothesis tests
result = s.t_test(popmean=2.5)
print(f"T-test p-value: {result['p_value']:.4f}")

# Work with distributions
norm_dist = pst.distribution('normal', mu=0, sigma=1)
norm_dist.plot_pdf()
plt.show()

# Analyze categorical data
contingency_table = pst.categorical.ContingencyTable([[20, 15], [10, 25]])
chi2_result = contingency_table.chi_square_test()
print(f"Chi-square p-value: {chi2_result['p_value']:.4f}")

# Calculate with error propagation
result, uncertainty = pst.propagate_error_product(10.5, 0.2, 3.2, 0.1)
print(f"Result: {pst.format_with_uncertainty(result, uncertainty)}")

Installation

pip install pstatstools

The Amazing Sample Class

The Sample class is the heart of pstatstools - it wraps your data into an object with powerful statistical methods:

import pstatstools as pst
import numpy as np

# Create a sample from any iterable
data = [23, 25, 21, 24, 29, 22, 20, 25, 24, 27]
s = pst.sample(data)

# Basic statistics with intuitive methods
print(f"Mean: {s.mean():.2f}")         # 24.00
print(f"Median: {s.median():.2f}")     # 24.00
print(f"Std Dev: {s.std():.2f}")       # 2.62
print(f"95% CI: {s.ci()}")             # (22.17, 25.83)

# Beautiful formatted output in Jupyter notebooks
s.describe()  # Creates an elegant HTML table with all key statistics

Jupyter Notebook Integration

The describe() method generates rich, interactive output in Jupyter notebooks:

  • Organized tables for central tendency, dispersion, quartiles, and shape statistics
  • Automatic detection of notebook environment for optimal display
  • Confidence intervals and interpretative guidelines
  • Customizable precision with the round_to parameter

Visualization Made Simple

Creating statistical plots is effortless:

import matplotlib.pyplot as plt

# Histogram with normal curve overlay
s.histogram(bins=10, density=True, show_norm=True)

# Box plot
s.boxplot()

# Q-Q plot for normality check
s.qq_plot()

# Show all plots
plt.show()

Hypothesis Testing

Run common statistical tests with a single method call:

# One-sample t-test
result = s.t_test(popmean=22)
print(f"t-statistic: {result['t_statistic']:.2f}, p-value: {result['p_value']:.4f}")

# Compare with another sample
s2 = pst.sample([19, 18, 22, 20, 23, 21, 20, 19, 21, 20])
result = s.welch_test(s2)
print(f"Welch's t-test p-value: {result['p_value']:.4f}")

# Test for normality
normality = s.normality_test()
print(f"Is normal? {normality['is_normal']}")

Distribution Framework

Create and work with probability distributions using the intuitive distribution factory function:

# Create a normal distribution
norm = pst.distribution('normal', mu=100, sigma=15)

# Calculate probabilities
p = norm.cdf(115) - norm.cdf(85)  # P(85 < X < 115)
print(f"Probability between 85 and 115: {p:.4f}")

# Plot the distribution
norm.plot_pdf(shade_area=(85, 115))
plt.show()

# Compare with data
data = norm.random(1000)  # Generate 1000 random samples
norm.compare_with_data(data)
plt.show()

Categorical Data Analysis

import pstatstools as pst

# Create a contingency table
table = pst.categorical.ContingencyTable([
    [30, 10, 15],
    [25, 20, 10]
])

# Run chi-square test
result = table.chi_square_test()
print(f"Chi-square: {result['chi2_statistic']:.2f}, p-value: {result['p_value']:.4f}")

# Plot the table
table.plot(kind='heatmap')
plt.show()

Non-Parametric Tests

# Run Mann-Whitney U test (non-parametric t-test alternative)
result = pst.nonparametric.mann_whitney_test([1, 2, 3, 4, 5], [2, 3, 4, 5, 6])
print(f"Mann-Whitney U p-value: {result['p_value']:.4f}")

# Run Wilcoxon signed-rank test (non-parametric paired t-test alternative)
result = pst.nonparametric.wilcoxon_test([1, 2, 3, 4, 5], [1.1, 2.2, 3.3, 4.4, 5.5])
print(f"Wilcoxon p-value: {result['p_value']:.4f}")

Error Propagation

# Simple error propagation for product
result, uncertainty = pst.propagate_error_product(10.2, 0.1, 5.3, 0.2)
print(f"10.2 ± 0.1 × 5.3 ± 0.2 = {result:.2f} ± {uncertainty:.2f}")

# Monte Carlo error propagation for complex functions
def my_function(x, y, z):
    return (x**2 + y) / z

result, uncertainty = pst.monte_carlo_error_propagation(
    my_function,
    {'x': (5.0, 0.1), 'y': (10.0, 0.2), 'z': (2.0, 0.05)}
)
print(f"Result: {pst.format_with_uncertainty(result, uncertainty)}")

ANOVA and Experimental Design

# One-way ANOVA from summary data
result = pst.anova_from_summary_table(
    ss_total=500, 
    ss_between=200, 
    n_groups=3, 
    n_total=30
)
print(f"F-statistic: {result['f_statistic']:.2f}, p-value: {result['p_value']:.4f}")

# Power analysis for ANOVA
power = pst.anova_power_analysis(groups=3, n_per_group=10, effect_size=0.4)
print(f"Statistical power: {power:.3f}")

# Sample size calculation
sample_size = pst.calculate_sample_size_anova(groups=3, effect_size=0.3, power=0.8)
print(f"Required sample size: {sample_size}")

Probability Calculations

# Calculate binomial probability
p = pst.probability.binomial_probability(k=3, n=10, p=0.2)
print(f"P(X = 3) = {p:.4f}")

# Calculate poisson probability 
p = pst.probability.poisson_probability(k=5, lambda_=3, cumulative=True)
print(f"P(X ≤ 5) = {p:.4f}")

# Calculate waiting time in a Poisson process
stats = pst.probability.calculate_waiting_time(arrival_rate=2, n_arrivals=3)
print(f"Expected waiting time: {stats['mean']:.2f} units")
print(f"95% CI: ({stats['ci_95'][0]:.2f}, {stats['ci_95'][1]:.2f})")

Dependencies

  • Python >= 3.6
  • NumPy
  • SciPy
  • Matplotlib
  • Pandas
  • Statsmodels

Complete Documentation

For complete API documentation and more examples, visit our documentation site.

Contributing

Contributions are welcome! Please feel free to submit a pull request or open an issue.

License

Distributed under the terms of the MIT License.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pstatstools-0.0.11.tar.gz (60.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pstatstools-0.0.11-py3-none-any.whl (58.8 kB view details)

Uploaded Python 3

File details

Details for the file pstatstools-0.0.11.tar.gz.

File metadata

  • Download URL: pstatstools-0.0.11.tar.gz
  • Upload date:
  • Size: 60.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for pstatstools-0.0.11.tar.gz
Algorithm Hash digest
SHA256 0d773dfdd3319b3af01de1064080956836e2e197fcf601ad091bccc9f6fb26fe
MD5 1d81c88040d3cd2ce029a41d09851d8e
BLAKE2b-256 24cd33d1ebfc061b1ecdd3ac5308325cdd891fab5ab9d486bf7edb6e46b035cf

See more details on using hashes here.

File details

Details for the file pstatstools-0.0.11-py3-none-any.whl.

File metadata

  • Download URL: pstatstools-0.0.11-py3-none-any.whl
  • Upload date:
  • Size: 58.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for pstatstools-0.0.11-py3-none-any.whl
Algorithm Hash digest
SHA256 e678c7133234a2a953c8569ac3b5bb90805a028697e90811a4c4b12072175fcc
MD5 832a7f495fe5f42c48b6088cb902021e
BLAKE2b-256 d023838bf83f3e2609381031044f2fab202d37f64410532c548388d059991f0e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page