DataGuard AI
Open-source, AI-assisted data quality and governance for modern data platforms.
Scan your data. Find quality and governance risks. Understand why they matter. Generate the fix.
DataGuard AI is a developer-first toolkit for data quality, data governance, PII discovery, schema drift, data contracts, and CI/CD-friendly validation. Detection is deterministic and does not require an LLM. Optional AI assistance can explain structured findings and suggest remediation without making quality detection dependent on a model.
Current version: v0.2.1 public beta.
Why DataGuard AI?
Data teams often manage quality rules, contracts, PII checks, schema drift, metadata, and AI assistants in separate workflows. DataGuard AI provides a lightweight layer developers can run locally or in CI to surface these risks through one interface.
What v0.2 can do
- Profile CSV and JSON data, with optional Parquet support
- Scan DuckDB and PostgreSQL tables
- Run reusable YAML data-quality rules
- Detect nulls, duplicate rows, uniqueness issues, range violations, and freshness risks
- Identify likely PII using value and column-name signals
- Compute transparent quality and governance scores
- Generate standalone HTML and machine-readable JSON reports
- Snapshot schemas and detect schema drift
- Generate starter dbt tests and YAML data contracts
- Export portable Great Expectations expectation configuration
- Run in GitHub Actions
- Explain findings locally, with optional AI-assisted remediation guidance
Quick start
Install from PyPI
pip install dataguard-ai
Run the synthetic retail demo:
dataguard demo --rows 100
Or scan your own dataset:
dataguard scan data/customers.csv
Generate JSON and HTML reports:
dataguard scan data/customers.csv --json-out report.json --html-out report.html
Install from source
For development or contributing:
git clone https://github.com/adiranjan25/dataguard-ai.git
cd dataguard-ai
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"
Detection philosophy
AI assists; deterministic and statistical checks detect and verify.
The default scanning path does not require an LLM. This keeps findings reproducible and allows teams to use DataGuard AI without sending raw production datasets to an external model.
YAML rule engine
Supported v0.2 custom rule types:
not_nulluniqueaccepted_valuesbetweenregexmax_null_pctrow_count_between
Example:
quality:
max_null_pct: 5
governance:
owner: data-platform@example.com
rules:
- type: not_null
column: customer_id
severity: CRITICAL
- type: unique
column: customer_id
- type: accepted_values
column: state
values: [TX, CA, NY]
- type: between
column: amount
min: 0
max: 100000
dataguard scan customers.csv --config dataguard.yml --html-out report.html
Explicit uniqueness rules
DataGuard does not assume that every column ending in _id is a primary key. Columns such as customer_id or store_id often represent foreign keys and may legitimately contain repeated values.
When uniqueness is part of the dataset contract, declare it explicitly:
rules:
- type: unique
column: customer_id
severity: HIGH
A generic column named id is still treated as a primary-key-like identifier by the built-in heuristic.
Database scanning
DuckDB
pip install "dataguard-ai[duckdb]"
dataguard scan-duckdb analytics.duckdb --table customers --html-out report.html
PostgreSQL
pip install "dataguard-ai[postgres]"
export DATAGUARD_POSTGRES_URL='postgresql+psycopg://user:password@host/database'
dataguard scan-postgres --table public.customers --html-out report.html
Current limitation: database scans load the selected table/result into memory. Warehouse-scale pushdown profiling is a roadmap item.
Reports
dataguard scan customers.csv --html-out report.html
dataguard scan customers.csv --json-out report.json
The HTML report includes quality/governance scores, findings, severity, affected columns, suggested remediation, column profiles, and detected PII.
Data contracts and dbt
dataguard contract data/customers.csv --out contract.yml
dataguard generate-dbt data/customers.csv --out schema.yml
These outputs are intended as reviewable starting points rather than replacements for domain-specific contract design.
Great Expectations
dataguard generate-gx data/customers.csv --out gx-expectations.json
The exporter deliberately produces reviewable configuration rather than modifying an existing Great Expectations project.
Schema drift
dataguard snapshot data/customers.csv --out baseline.json
dataguard drift data/customers_v2.csv --baseline baseline.json
GitHub Actions
The repository includes an example workflow at .github/workflows/dataguard.yml that demonstrates scanning sample data and uploading JSON/HTML reports as workflow artifacts.
Project CI separately runs linting and automated tests against Python 3.10, 3.11, and 3.12.
Optional AI explanations
Local deterministic explanations are available without an external model.
pip install "dataguard-ai[ai]"
export OPENAI_API_KEY=...
dataguard explain report.json --provider openai
The included provider sends structured findings rather than raw dataset rows. Always review your organization's security, privacy, and data-handling requirements before enabling an external provider.
Architecture
Data sources
|
+-----------------+-----------------+
| | |
CSV / JSON DuckDB PostgreSQL
/ Parquet | |
+-----------------+-----------------+
|
v
DataGuard scanner
|
+----------------+----------------+
| | |
Profiling Rule engine PII detection
| | |
+----------------+----------------+
|
v
Quality + governance
|
+-----------+-----+------+-----------+
| | | |
v v v v
CLI JSON HTML Contracts / dbt / GX
Detection and optional AI explanation are intentionally separated.
Commands
| Command | Purpose |
|---|---|
dataguard scan PATH |
Profile file data and report quality/governance findings |
dataguard demo |
Generate and scan synthetic retail datasets |
dataguard scan-duckdb DATABASE --table TABLE |
Scan a DuckDB table |
dataguard scan-postgres --table TABLE |
Scan a PostgreSQL table |
dataguard contract PATH |
Generate a starter data contract |
dataguard generate-dbt PATH |
Generate starter dbt tests |
dataguard generate-gx PATH |
Generate portable GX expectation configuration |
dataguard snapshot PATH |
Save a schema/profile baseline |
dataguard drift PATH --baseline FILE |
Compare current schema with a baseline |
dataguard explain REPORT.json |
Explain findings locally or with optional AI |
Run dataguard --help for current CLI options.
Python API
from dataguard.scanner import scan_path
report = scan_path("customers.csv")
print(report.quality_score)
for finding in report.findings:
print(finding.severity, finding.column, finding.message)
Retail demo
The bundled synthetic retail demo creates customer, order, and inventory datasets with intentionally injected quality/governance problems so developers can explore DataGuard AI without providing proprietary data.
dataguard demo --rows 5000
Project status
DataGuard AI is currently a public beta (v0.2.1). The API, configuration schema, scoring model, and command behavior may evolve before v1.0.
The project is suitable for experimentation, development workflows, demos, and community feedback. Evaluate it against your own requirements before using it as a production control.
Roadmap
v0.3 — Metadata and context
- OpenMetadata integration
- DataHub integration
- dbt artifact ingestion
- Dataset-to-dataset referential checks
- Richer statistical drift and anomaly detection
v0.4 — Agent access
- MCP server
- Agent-accessible quality, contract, and governance tools
- Ollama/local-model provider
- Governance context for AI agents
- Assisted remediation workflows
Toward v1.0
- PyPI distribution and automated release workflow
- Warehouse-scale profiling/pushdown
- Broader integration tests
- Benchmark datasets and reproducible evaluation
- Stable configuration and CLI contracts
Contributing
Contributions are welcome. See CONTRIBUTING.md for development setup and contribution guidance.
Useful first contributions include additional PII detectors, report/export formats, documentation improvements, tests, database adapters, and integration examples.
Security and privacy
- Never submit secrets, credentials, or proprietary production datasets in GitHub issues.
- Use environment variables or an appropriate secrets manager for database and model credentials.
- Treat detected PII findings as sensitive operational metadata.
- Review organizational security/privacy requirements before using an external AI provider.
- See SECURITY.md for vulnerability-reporting guidance.
License
DataGuard AI is licensed under the Apache License 2.0. See LICENSE.
Metadata
Release files for dataguard-ai 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dataguard_ai-0.2.1.tar.gz | 22.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dataguard_ai-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 43.2 kB
Release files / dataguard_ai-0.2.1.tar.gz
| Download URL | dataguard_ai-0.2.1.tar.gz |
|---|---|
| Size | 22.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
94e808385f60f9b3a46f2fe2d5b6a852cf855323c383a495d773e42ca9e113f4
|
|
BLAKE2b-256 checksum How to use checksums |
b76112ff74ea2ff11f310f22fa9e21eb4b31529cd8cb42c1df4f0f41426932ac
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / dataguard_ai-0.2.1-py3-none-any.whl
| Download URL | dataguard_ai-0.2.1-py3-none-any.whl |
|---|---|
| Size | 20.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f856ab0eaa03bcafa6ac56f33d8a66acd1130027f15a67e74e57dc44a58e19bc
|
|
BLAKE2b-256 checksum How to use checksums |
e723e3b6b45126cf7311a550a7b098d0a953e90b0ce5ed8c7b5e0be910bca17a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log