Skip to main content

PySuricata

Build Status PyPI version Python versions License: MIT codecov Documentation Downloads LinkedIn

PySuricata Logo

Exploratory Data Analysis for Python, Built on Streaming Algorithms

Quick StartDocumentationExamples


What It Does

PySuricata generates self-contained HTML reports from pandas or polars DataFrames. Reports include per-column statistics, histograms, correlation chips, missing value analysis, and outlier detection.

Data is processed in chunks using streaming algorithms, so memory usage stays bounded in the number of rows — a million rows costs no more than twenty thousand. It is not bounded in the number of columns: each column keeps its own sketches for the whole run and gets its own card in the report, so both memory and report size grow linearly with the width of the frame. Measured at 20,000 rows: ~1.3 MB of RSS and ~59 KB of report per column, so a 600-column frame needs roughly 850 MB. See #207.

It also does two things a profiler usually does not: summarize() returns the same numbers as a versioned JSON payload with no HTML in the way, and pysuricata check compares a dataset against a stored baseline and exits non-zero when a threshold is crossed — so the same single pass can run in a notebook and in CI.

Quick Start

Installation

# using uv (recommended)
uv add pysuricata

# or using pip
pip install pysuricata

With polars support (optional):

uv add pysuricata[polars]
# or: pip install pysuricata[polars]

Generate a Report

import pandas as pd
from pysuricata import profile

url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)

report = profile(df)
report.save_html("titanic_report.html")

▶ See a live example report →

A PySuricata report: the dataset summary, the five columns that need a look, and a numeric column card with its histogram and bin controls

The examples below assume a df in scope. The Quick Start frame works, or anything of your own:

import numpy as np
import pandas as pd

rng = np.random.default_rng(0)
df = pd.DataFrame(
    {
        "age": rng.normal(30, 12, 800).round(1),
        "fare": rng.gamma(2, 20, 800).round(2),
        "sex": rng.choice(["male", "female"], 800),
        "booked": pd.date_range("2024-01-01", periods=800, freq="h"),
    }
)

Features

  • Streaming architecture — Data is processed in configurable chunks, keeping memory bounded in rows (not in columns — see above). Useful for datasets with more rows than fit in RAM.
  • Pandas and Polars — Works natively with pandas.DataFrame, polars.DataFrame and polars.LazyFrame, plus Parquet files, DuckDB relations and Arrow batches.
  • Self-contained HTML — Single file with inline CSS, JS, and SVG charts. No external assets needed.
  • Configurable — Control chunk size, sample size, correlations and more with keyword options, a preset=, or a ProfileConfig.
  • Reproducible — Seeded random sampling produces deterministic results across runs.
  • Typed — Ships py.typed; summarize() returns a payload carrying a schema_version.
  • CLI toolprofile, summarize and check from the command line.

How It Works

PySuricata uses well-known streaming algorithms from the academic literature:

Algorithm Purpose Time Space
Welford/Pébay Exact mean, variance, skewness, kurtosis O(1) per value O(1)
KMV sketch Distinct count estimation (~2.2% error) O(log k) per value O(k)
Misra-Gries Top-k frequent values O(1) amortized O(k)
Reservoir sampling Uniform random sample for quantiles O(1) per value O(s)

k = sketch size (max_uniques, default 2048), s = sample size (numeric_sample_size, default 20 000)

KMV's relative standard error is 1/sqrt(k - 2), which is where the ~2.2% comes from. Approximate values are labelled approximate in the report and carry their error bound rather than being printed as exact integers.

All statistics are computed in a single pass over the data.

What's in a Report

Each column is analyzed based on its type:

  • Numeric — Mean, variance, skewness, kurtosis, quantiles, histogram, outlier detection (IQR, MAD, z-score), correlations
  • Categorical — Top values, distinct count, entropy, Gini impurity, string length statistics
  • DateTime — Temporal range, hour/day/month distributions, monotonicity detection
  • Boolean — True/false ratios, entropy, balance score

Plus dataset-level metrics: row/column counts, memory usage, missing value percentages, and duplicate row estimates.

Streaming Large Datasets

Process datasets larger than RAM by passing a generator:

import pandas as pd
from pysuricata import profile

def read_in_chunks():
    for i in range(100):
        yield pd.read_parquet(f"data/part-{i}.parquet")

report = profile(read_in_chunks())
report.save_html("large_report.html")

A Parquet path, a DuckDB relation or an Arrow source can be streamed directly, without loading the whole thing:

import duckdb
from pysuricata import profile
from pysuricata.sources import stream_duckdb, stream_parquet

report = profile(stream_parquet("data/events.parquet"))

relation = duckdb.connect("warehouse.db").sql("SELECT * FROM events")
report = profile(stream_duckdb(relation))

Statistics Only (No HTML)

Use summarize() for CI/CD quality checks. The payload carries a schema_version and is treated as a contract:

from pysuricata import summarize

stats = summarize(df)

assert stats["schema_version"] == 1
assert stats["dataset"]["missing_cells_pct"] < 5.0
assert stats["dataset"]["duplicate_rows_pct_est"] < 1.0

print(f"Mean age: {stats['columns']['age']['mean']:.1f}")

Comparing Two Datasets

compare() runs both through the same single pass and reports what moved:

from pysuricata import compare

last_week, this_week = df.iloc[:3], df.iloc[3:]
diff = compare(last_week, this_week).to_dict()

Configuration

Pass keyword options for the common cases:

from pysuricata import profile

report = profile(
    df,
    chunk_size=250_000,   # default 50_000
    sample=20_000,
    seed=42,
    correlations=True,
    title="My Analysis",
)

Or start from a preset — "fast" or "thorough":

from pysuricata import profile

report = profile(df, preset="fast")

For everything else, build a ProfileConfig. Keyword options and config= are mutually exclusive:

from pysuricata import profile, ProfileConfig

config = ProfileConfig()
config.compute.chunk_size = 250_000
config.compute.random_seed = 42
config.compute.corr_threshold = 0.5
config.render.title = "My Analysis"

report = profile(df, config=config)

See the Configuration Guide for all options.

CLI

# Generate an HTML report
pysuricata profile data.csv --output report.html

# Get JSON statistics
pysuricata summarize data.csv

# Compare against a stored baseline; exit non-zero when a threshold is crossed
pysuricata check data.csv --write-baseline baseline.json
pysuricata check data.csv --baseline baseline.json --max-missing-pct 5

check exits 0 on pass, 1 when a threshold is crossed, and 2 when the check could not run — so it drops into CI without a wrapper.

Documentation

Contributing

Contributions are welcome. See the Contributing Guide.

git clone https://github.com/alvarodiez20/pysuricata.git
cd pysuricata
uv sync --dev
uv run pytest

License

MIT License. See LICENSE for details.

Acknowledgments

Built using algorithms from:

  • Welford, B.P. (1962) — Streaming moments
  • Pébay, P. (2008) — Parallel merging of moments
  • Bar-Yossef, Z. et al. (2002) — KMV distinct count estimation
  • Misra, J. & Gries, D. (1982) — Streaming heavy hitters

Named after suricatas (meerkats) — small, vigilant animals that work cooperatively and thrive in harsh environments with limited resources.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pysuricata-0.1.2.tar.gz (893.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pysuricata-0.1.2-py3-none-any.whl (669.4 kB view details)

Uploaded Python 3

File details

Details for the file pysuricata-0.1.2.tar.gz.

File metadata

  • Download URL: pysuricata-0.1.2.tar.gz
  • Upload date:
  • Size: 893.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pysuricata-0.1.2.tar.gz
Algorithm Hash digest
SHA256 81aeee6efa9a916bf2891ab8e39840f9701ee49a7432a4d432814947d2751b7d
MD5 9df13c942cf4f4a2ac83dfbdfdda6dcc
BLAKE2b-256 b2dd0166286c1a76a152533616754b1d53bffdb5170ad3445f4dbf8b8a136fd0

See more details on using hashes here.

Provenance

The following attestation bundles were made for pysuricata-0.1.2.tar.gz:

Publisher: cd.yml on alvarodiez20/pysuricata

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pysuricata-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: pysuricata-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 669.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pysuricata-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 44d6e292ac60f4364658e5fbac6135789433b6a5c1b44cbb003e423fc56d5937
MD5 ee94109e3878069e1b147702d1f2c7e9
BLAKE2b-256 ca0085c8b18e0914a10a03cf37843562a165a8f5e31d199f86ddb3a0341b3e62

See more details on using hashes here.

Provenance

The following attestation bundles were made for pysuricata-0.1.2-py3-none-any.whl:

Publisher: cd.yml on alvarodiez20/pysuricata

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.0

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

This release

0.1.2 This release

2 files

0.1.1

2 files

0.1.0

2 files

0.0.73

2 files

0.0.72

2 files

0.0.71

2 files

0.0.70

2 files

0.0.68

2 files

0.0.66

2 files

0.0.65

2 files

0.0.64

2 files

0.0.63

2 files

0.0.62

2 files

0.0.61

2 files

0.0.60

2 files

0.0.59

2 files

0.0.58

2 files

0.0.57

2 files

0.0.56

2 files

0.0.55

2 files

0.0.54

2 files

0.0.53

2 files

0.0.52

2 files

0.0.51

2 files

0.0.50

2 files

0.0.49

2 files

0.0.48

2 files

0.0.47

2 files

0.0.46

2 files

0.0.45

2 files

0.0.44

2 files

0.0.43

2 files

0.0.42

2 files

0.0.41

2 files

0.0.40

2 files

0.0.39

2 files

0.0.38

2 files

0.0.37

2 files

0.0.36

2 files

0.0.35

2 files

0.0.34

2 files

0.0.33

2 files

0.0.32

2 files

0.0.31

2 files

0.0.30

2 files

0.0.29

2 files

0.0.28

2 files

0.0.27

2 files

0.0.26

2 files

0.0.25

2 files

0.0.24

2 files

0.0.23

2 files

0.0.22

2 files

0.0.21

2 files

0.0.20

2 files

0.0.19

2 files

0.0.18

2 files

0.0.17

2 files

0.0.16

2 files

0.0.15

2 files

0.0.14

2 files

0.0.13

2 files

0.0.12

2 files

0.0.11

2 files

0.0.10

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page