Skip to main content

StatGuardian

A Rust-native data quality engine with a declarative contract DSL: schema validation, drift detection, and anomaly detection for Pandas and Polars.

Tests PyPI Python 3.8+

Stop data quality issues from reaching production. StatGuardian validates data at runtime against a versionable contract, catching schema violations, statistical drift, and anomalies before they reach downstream consumers.

30-Second Start

import polars as pl
import statguardian

contract = statguardian.DataContract.from_dsl("""
dataset orders {
    schema {
        order_id: string, not_null, unique
        amount:   float,  positive
        status:   string, not_null, enum=["pending","paid","cancelled"]
    }
    quality {
        completeness(order_id) > 0.999
    }
}
""")

df = pl.read_parquet("orders.parquet")
report = statguardian.execute(contract, df)
print(report.summary())
print(f"Passed: {report.passed}")

Why StatGuardian?

  • Contracts are declarative and versionable (.sg files), not scattered assertions in application code
  • Rust-native execution — schema, quality, drift, and anomaly checks run in the compiled engine, not a Python loop
  • One contract, multiple frameworks: the same .sg file validates Pandas and Polars DataFrames, Delta Lake tables, and Apache Iceberg tables
  • Drift and anomaly detection are first-class DSL constructs, not a separate library

A reproducible benchmark comparing StatGuardian against other validation libraries is tracked in docs/bench/benchmark.py — run it against your own workload rather than relying on any library's marketing numbers, including ours.

Real-World Use Cases

E-commerce order validation

contract = statguardian.DataContract.from_dsl("""
dataset orders {
    schema {
        order_id: string, not_null, unique
        amount:   float,  positive
        status:   string, not_null, enum=["pending","shipped","delivered"]
    }
}
""")
report = statguardian.execute(contract, orders_df)

Drift monitoring between two batches

report = statguardian.execute(contract, incoming_df, reference=baseline_df)
for d in report.drift_results():
    if not d["passed"]:
        print(f"Drift detected in {d['column']}: PSI={d.get('psi', 0):.4f}")

Key Capabilities

  • Declarative contract DSL: schema, quality rules, statistical drift thresholds, and anomaly checks in one file
  • Type checking with detailed, structured violation messages
  • Statistical drift detection (PSI, KS test) between a dataset and a reference baseline
  • Built-in anomaly detection (outliers, duplicates)
  • Supports Pandas and Polars DataFrames, Delta Lake, and Apache Iceberg tables with the same contract
  • Rust-native execution core

Features

Core Validation

  • Type validation (int, float, str, bool, datetime, etc.)
  • Min/max constraints for numeric types
  • Enum validation for categorical data
  • Null/not-null constraints
  • Pattern matching for strings (regex)
  • Custom validation functions
  • Composite constraints (multiple rules per field)

Data Quality Analysis

  • Automatic drift detection (schema changes)
  • Anomaly detection (outliers, unexpected values)
  • Statistical profiling (mean, std, quartiles)
  • Missing value reporting
  • Duplicate detection

Framework Support

  • Pandas DataFrames (convert with pl.from_pandas(df) before calling execute() — see Known Issues)
  • Polars DataFrames (native)
  • Delta Lake tables (time-travel validation)
  • Apache Iceberg tables (snapshot validation)
  • Unified contract across all frameworks

Requirements

  • Python: 3.8+
  • Core: Rust-native validation engine (precompiled wheel, no local Rust toolchain needed)
  • Data Frameworks: polars (required), pandas (optional, via pip install statguardian[pandas])

Examples

See examples/ for complete, runnable scripts, including python_quickstart.py (schema validation, drift detection, anomaly detection, JSON/Prometheus output) and .sg contract files.

Schema validation

contract = statguardian.DataContract.from_dsl("""
dataset users {
    schema {
        id:    int,    not_null, unique, primary_key
        email: string, regex="^[^@]+@[^@]+\\.[^@]+$"
        age:   int,    between(0, 120)
    }
    quality {
        completeness(id) > 0.99
    }
}
""")

report = statguardian.execute(contract, df)
print(report.summary())
for v in report.violations():
    print(v["severity"], v["column"], v["message"])

Anomaly detection

contract = statguardian.DataContract.from_dsl("""
dataset events {
    schema { id: int, not_null }
    anomalies {
        detect_outliers(id, method="iqr")
        @blocking: detect_duplicates(id)
    }
}
""")
report = statguardian.execute(contract, df)

Custom Python validators + merging with a contract report

@statguardian.validator(column="amount", severity="blocking")
def amount_is_sane(values):
    bad_rows = [i for i, v in enumerate(values) if v > 1_000_000]
    return (bad_rows, "amount over 1,000,000") if bad_rows else None

report = statguardian.execute(contract, df)
extra = statguardian.run_custom_validators(df)
merged = statguardian.merge_violations(report, extra)
print(merged.summary())

API Reference

Core

  • DataContract.from_dsl(dsl_string) / DataContract.from_file(path) — compile a contract
  • execute(contract, df, reference=None) -> ValidationReport — validate a Pandas/Polars DataFrame
  • execute_file(contract, path, reference_path=None) — validate Parquet/CSV/JSON/Avro/ORC/Arrow IPC files
  • execute_delta(contract, path, ...), execute_iceberg(contract, path, ...) — lakehouse table validation
  • execute_sql, execute_spark, execute_cloud — SQL, PySpark, and object-storage sources

ValidationReport

  • .passed, .health_score, .grade, .violation_count
  • .violations(), .drift_results(), .column_profiles()
  • .summary(), .to_json(), .to_prometheus()

Custom validators

  • validator(column=...) — register a Python function as a custom check
  • run_custom_validators(df) — run registered validators, returns violation dicts
  • merge_violations(report, extra_violations) -> MergedReport — combine a ValidationReport with custom-validator violations into one pass/fail result

Full CLI usage: docs/CLI.md. DSL syntax: see examples/*.sg.

dbt Integration

Run StatGuard contracts against your dbt models as part of dbt build, and surface pass/fail as a native dbt test — see integrations/dbt-statguardian and docs/DBT_INTEGRATION.md.

pip install "statguardian[dbt]"
dbt build
statguardian dbt validate --project-dir . --write-results
dbt test

Installation

pip install statguardian

For development:

git clone https://github.com/Mullassery/StatGuardian
cd StatGuardian
pip install -e ".[dev]"
pytest

Documentation

Known Issues

  • execute() accepts a Polars DataFrame, not a raw pandas DataFrame. Passing a pandas DataFrame directly raises an unhelpful AttributeError (verified against the current build) rather than converting automatically — call pl.from_pandas(df) first. The pandas extra is used by the SQL/Spark/GPU connectors internally, which already do this conversion for you.
  • Performance numbers are not yet published as a reproducible, checked-in benchmark result — docs/bench/benchmark.py exists but its output has never been committed. Treat any speed claims (including from this project) as unverified until you've run the benchmark yourself.
  • docs/ROADMAP.md, docs/ROADMAP_HONEST.md, and docs/ROADMAP_INTEGRATED.md currently overlap and are not kept in sync — some content in ROADMAP_HONEST.md predates features (e.g. Iceberg support) that have since shipped. Treat docs/SECURITY_AUDIT.md as the current source of truth for security status; the roadmap docs need consolidation.
  • SQL connector extras (connectorx, psycopg2-binary, cloud warehouse drivers) use floating minimum versions rather than pinned versions — see docs/SECURITY_AUDIT.md for the rationale and tradeoffs.

License

Proprietary — free to use with attribution. See LICENSE.

Metadata

Release files for statguardian 2.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for statguardian 2.6.0
File Size Uploaded
statguardian-2.6.0.tar.gz 126.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for statguardian 2.6.0
File Interpreter ABI Platform
statguardian-2.6.0-cp38-abi3-macosx_11_0_arm64.whl CPython 3.8 abi3 macOS 11.0+ ARM64 Details

Total release size: 9.5 MB

Release files / statguardian-2.6.0.tar.gz

Download URL statguardian-2.6.0.tar.gz
Size 126.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a45b9588b60457cf5a9f8c439e11116a487153842fa3b054410b553df81e2cdf
BLAKE2b-256 checksum
How to use checksums
26a3ebc5f8824268077e7a9db41510bb73cbe28a1fab4430ffebb5c4aea0aeff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / statguardian-2.6.0-cp38-abi3-macosx_11_0_arm64.whl

Download URL statguardian-2.6.0-cp38-abi3-macosx_11_0_arm64.whl
Size 9.4 MB
Tags CPython 3.8 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
903a0125b7085217390c3e3abdc481e92bfaff0ab64640dd0c98f05f7d72ab6d
BLAKE2b-256 checksum
How to use checksums
5c40fd1bd32bc2411f9d4739690f10c483ef53ce21ef036d76b7b129a77e1230
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

2.6.1

1 release file

This release

2.6.0 This release

2 release files

2.5.0

2 release files

2.4.0

2 release files

2.3.2

2 release files

2.3.1

1 release file

2.3.0

2 release files

2.2.1

1 release file

2.2.0

1 release file

2.1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page