Skip to main content

DataGuard AI

CI Python 3.10+ License: Apache-2.0 Status: Public Beta

Open-source, AI-assisted data quality and governance for modern data platforms.

Scan your data. Find quality and governance risks. Understand why they matter. Generate the fix.

DataGuard AI is a developer-first toolkit for data quality, data governance, PII discovery, schema drift, data contracts, and CI/CD-friendly validation. Detection is deterministic and does not require an LLM. Optional AI assistance can explain structured findings and suggest remediation without making quality detection dependent on a model.

Current version: v0.2.0 public beta.

Why DataGuard AI?

Data teams often manage quality rules, contracts, PII checks, schema drift, metadata, and AI assistants in separate workflows. DataGuard AI provides a lightweight layer developers can run locally or in CI to surface these risks through one interface.

What v0.2 can do

  • Profile CSV and JSON data, with optional Parquet support
  • Scan DuckDB and PostgreSQL tables
  • Run reusable YAML data-quality rules
  • Detect nulls, duplicate rows, uniqueness issues, range violations, and freshness risks
  • Identify likely PII using value and column-name signals
  • Compute transparent quality and governance scores
  • Generate standalone HTML and machine-readable JSON reports
  • Snapshot schemas and detect schema drift
  • Generate starter dbt tests and YAML data contracts
  • Export portable Great Expectations expectation configuration
  • Run in GitHub Actions
  • Explain findings locally, with optional AI-assisted remediation guidance

Quick start

Install from source during the public beta

Until DataGuard AI is published to PyPI, clone the repository and install it locally:

git clone https://github.com/adiranjan25/dataguard-ai.git
cd dataguard-ai
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .

Run the synthetic retail demo:

dataguard demo --rows 100

Or scan your own dataset:

dataguard scan data/customers.csv

Generate JSON and HTML reports:

dataguard scan data/customers.csv --json-out report.json --html-out report.html

Coming next: after the PyPI release, installation will become simply pip install dataguard-ai.

Detection philosophy

AI assists; deterministic and statistical checks detect and verify.

The default scanning path does not require an LLM. This keeps findings reproducible and allows teams to use DataGuard AI without sending raw production datasets to an external model.

YAML rule engine

Supported v0.2 custom rule types:

  • not_null
  • unique
  • accepted_values
  • between
  • regex
  • max_null_pct
  • row_count_between

Example:

quality:
  max_null_pct: 5
governance:
  owner: data-platform@example.com
rules:
  - type: not_null
    column: customer_id
    severity: CRITICAL
  - type: unique
    column: customer_id
  - type: accepted_values
    column: state
    values: [TX, CA, NY]
  - type: between
    column: amount
    min: 0
    max: 100000
dataguard scan customers.csv --config dataguard.yml --html-out report.html

Database scanning

DuckDB

python -m pip install -e ".[duckdb]"
dataguard scan-duckdb analytics.duckdb --table customers --html-out report.html

PostgreSQL

python -m pip install -e ".[postgres]"
export DATAGUARD_POSTGRES_URL='postgresql+psycopg://user:password@host/database'
dataguard scan-postgres --table public.customers --html-out report.html

Current limitation: database scans load the selected table/result into memory. Warehouse-scale pushdown profiling is a roadmap item.

Reports

dataguard scan customers.csv --html-out report.html
dataguard scan customers.csv --json-out report.json

The HTML report includes quality/governance scores, findings, severity, affected columns, suggested remediation, column profiles, and detected PII.

Data contracts and dbt

dataguard contract data/customers.csv --out contract.yml
dataguard generate-dbt data/customers.csv --out schema.yml

These outputs are intended as reviewable starting points rather than replacements for domain-specific contract design.

Great Expectations

dataguard generate-gx data/customers.csv --out gx-expectations.json

The exporter deliberately produces reviewable configuration rather than modifying an existing Great Expectations project.

Schema drift

dataguard snapshot data/customers.csv --out baseline.json
dataguard drift data/customers_v2.csv --baseline baseline.json

GitHub Actions

The repository includes an example workflow at .github/workflows/dataguard.yml that demonstrates scanning sample data and uploading JSON/HTML reports as workflow artifacts.

Project CI separately runs linting and automated tests against Python 3.10, 3.11, and 3.12.

Optional AI explanations

Local deterministic explanations are available without an external model.

python -m pip install -e ".[ai]"
export OPENAI_API_KEY=...
dataguard explain report.json --provider openai

The included provider sends structured findings rather than raw dataset rows. Always review your organization's security, privacy, and data-handling requirements before enabling an external provider.

Architecture

                    Data sources
                         |
       +-----------------+-----------------+
       |                 |                 |
   CSV / JSON         DuckDB          PostgreSQL
   / Parquet             |                 |
       +-----------------+-----------------+
                         |
                         v
                  DataGuard scanner
                         |
        +----------------+----------------+
        |                |                |
     Profiling       Rule engine     PII detection
        |                |                |
        +----------------+----------------+
                         |
                         v
              Quality + governance
                         |
       +-----------+-----+------+-----------+
       |           |            |           |
       v           v            v           v
      CLI         JSON         HTML    Contracts / dbt / GX

Detection and optional AI explanation are intentionally separated.

Commands

Command Purpose
dataguard scan PATH Profile file data and report quality/governance findings
dataguard demo Generate and scan synthetic retail datasets
dataguard scan-duckdb DATABASE --table TABLE Scan a DuckDB table
dataguard scan-postgres --table TABLE Scan a PostgreSQL table
dataguard contract PATH Generate a starter data contract
dataguard generate-dbt PATH Generate starter dbt tests
dataguard generate-gx PATH Generate portable GX expectation configuration
dataguard snapshot PATH Save a schema/profile baseline
dataguard drift PATH --baseline FILE Compare current schema with a baseline
dataguard explain REPORT.json Explain findings locally or with optional AI

Run dataguard --help for current CLI options.

Python API

from dataguard.scanner import scan_path

report = scan_path("customers.csv")
print(report.quality_score)

for finding in report.findings:
    print(finding.severity, finding.column, finding.message)

Retail demo

The bundled synthetic retail demo creates customer, order, and inventory datasets with intentionally injected quality/governance problems so developers can explore DataGuard AI without providing proprietary data.

dataguard demo --rows 5000

Project status

DataGuard AI is currently a public beta (v0.2.0). The API, configuration schema, scoring model, and command behavior may evolve before v1.0.

The project is suitable for experimentation, development workflows, demos, and community feedback. Evaluate it against your own requirements before using it as a production control.

Roadmap

v0.3 — Metadata and context

  • OpenMetadata integration
  • DataHub integration
  • dbt artifact ingestion
  • Dataset-to-dataset referential checks
  • Richer statistical drift and anomaly detection

v0.4 — Agent access

  • MCP server
  • Agent-accessible quality, contract, and governance tools
  • Ollama/local-model provider
  • Governance context for AI agents
  • Assisted remediation workflows

Toward v1.0

  • PyPI distribution and automated release workflow
  • Warehouse-scale profiling/pushdown
  • Broader integration tests
  • Benchmark datasets and reproducible evaluation
  • Stable configuration and CLI contracts

Contributing

Contributions are welcome. See CONTRIBUTING.md for development setup and contribution guidance.

Useful first contributions include additional PII detectors, report/export formats, documentation improvements, tests, database adapters, and integration examples.

Security and privacy

  • Never submit secrets, credentials, or proprietary production datasets in GitHub issues.
  • Use environment variables or an appropriate secrets manager for database and model credentials.
  • Treat detected PII findings as sensitive operational metadata.
  • Review organizational security/privacy requirements before using an external AI provider.
  • See SECURITY.md for vulnerability-reporting guidance.

License

DataGuard AI is licensed under the Apache License 2.0. See LICENSE.

Metadata

Release files for dataguard-ai 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dataguard-ai 0.2.0
File Size Uploaded
dataguard_ai-0.2.0.tar.gz 21.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dataguard-ai 0.2.0
File Interpreter ABI Platform
dataguard_ai-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 42.6 kB

Release files / dataguard_ai-0.2.0.tar.gz

Download URL dataguard_ai-0.2.0.tar.gz
Size 21.8 kB
Tags Source
SHA-256 checksum
How to use checksums
2f01a505bc94c9003786eaae7fb716ac45af4f17667b4e72fa6e7e4540bc003f
BLAKE2b-256 checksum
How to use checksums
20b44579741d5c89c3c4e7f9a0b67aa1cc268873ab9a110e93d106e07d2bc68c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / dataguard_ai-0.2.0-py3-none-any.whl

Download URL dataguard_ai-0.2.0-py3-none-any.whl
Size 20.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
28c05d69616e3fd26094520ea0393b028c318d144661b50276f3f885ed37c864
BLAKE2b-256 checksum
How to use checksums
ee178c48e0bd327a3e2eaa4476e63ba52bf9ffdefef292cd214de1f0374ad0c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page