Skip to main content

DataSentry

DataSentry

Automatically find, explain, and safely fix bad data.
A local-first data quality copilot with deterministic detection, evidence-backed issues, and human-approved reversible repairs.

Live demo · Quick start · MCP setup · Why DataSentry · Roadmap · Contribute

Release PyPI Python CI Coverage License

DataSentry: scan data, inspect evidence-backed issues, and start the repair loop

中文导读:DataSentry 自动发现数据质量问题,给出样本、比例和置信度等证据,再生成可预览、可验证、可回滚的修复方案。检测核心不依赖 LLM;AI 只负责辅助规则与修复建议,是否执行由人决定。数据可完全留在本机。

The data quality loop

Detect  →  Explain  →  Repair  →  Verify
  │           │           │          │
39 detectors  evidence     preview    re-scan
              samples      copy       regression gate
              ratios       rollback
              confidence

Most data-quality failures are not hard because a check cannot be written. They are hard because you first have to discover the problem, understand whether it is real, fix it without damaging the source, and prove the fix did not introduce a regression.

DataSentry is built around that complete loop.

Quick start

pip install datasentry-ai

datasentry scan orders.csv
datasentry issues list --severity high

Start the interactive terminal UI with:

datasentry

Or launch the Web UI and REST API:

datasentry-server
# http://localhost:8000/ui/

Why DataSentry?

1. Discover problems before you know which rules to write

DataSentry ships with 39 evidence-driven detectors for missingness, invalid values, dates, encodings, duplicates, cross-field rules, foreign keys, and statistical outliers. A scan produces a six-dimension quality score across completeness, validity, uniqueness, consistency, integrity, and timeliness.

2. Evidence first, AI second

Detection and scoring are deterministic. Every issue can carry samples, ratios, and confidence so you can inspect why it was raised. LLMs are optional and are never the authority deciding whether data is valid.

3. Repair without gambling on the source file

The repair workflow is deliberately conservative:

propose → preview → apply to a copy → verify → rollback if needed

The original file is never overwritten. Repair runs are fingerprinted and auditable, and verification re-scans the repaired copy to surface fixed, persistent, and newly introduced problems.

4. Keep sensitive data local

DuckDB executes the core scan locally. LLM assistance is optional and can use a local Ollama provider. When an LLM is used, DataSentry includes PII redaction, an encrypted vault, and audit records for calls.

What it can scan

  • CSV, Parquet, JSONL, XLSX
  • DuckDB and SQLite
  • PostgreSQL and MySQL
  • Cloud objects on s3://, gs://, and az://
  • Single files or batch paths/globs

Safe repair workflow

# 1) inspect issues
datasentry issues list --severity high

# 2) generate a proposal (read-only)
datasentry repair propose <issue_id> --file orders.csv

# 3) preview the proposed change
datasentry repair preview <issue_id> --file orders.csv

# 4) apply to a repaired copy; the command returns a repair run_id
datasentry repair apply <issue_id> --file orders.csv

# 5) prove the repair did not introduce new issue types
datasentry repair verify <run_id>

# 6) inspect the row-level change or roll back
datasentry repair diff <run_id>
datasentry repair rollback <run_id>

Batch repair commands are also available for scan-wide workflows.

Quality gates for CI

Use DataSentry as a release gate instead of only as an interactive profiler:

datasentry scan orders.csv --fail-on high

Reports can be exported as JSON, Markdown, HTML, JUnit, and SARIF, making the same evidence usable by humans, CI systems, and code-scanning surfaces.

Drift and history

Persisted scans let DataSentry compare datasets over time:

datasentry drift latest orders
datasentry score

Tracked signals include schema changes, row-count changes, quality-score movement, and issue-distribution drift.

Use DataSentry with AI agents

DataSentry includes an MCP stdio server that exposes the same underlying SDK used by the CLI and REST API.

datasentry mcp --project /path/to/project

This lets MCP-capable agents scan files, inspect issues, read quality scores and trends, compare drift, validate contracts, manage scheduled jobs, and invoke other DataSentry tools without bypassing the project's safety invariants.

For copy-paste setup recipes for VS Code and Claude Desktop, path troubleshooting, and guidance on read-only vs state-changing agent workflows, see DataSentry MCP setup.

The important boundary remains the same: AI may propose; humans approve state-changing repairs. Keep client-side confirmation enabled for state-changing MCP tools.

Interfaces

Surface Best for
CLI / TUI local exploration, scripts, repair workflows
REST API application and service integration
Web UI issue review, reports, trends, comparisons, repair workbench
MCP stdio AI-agent workflows
Exporters JSON, Markdown, HTML, JUnit, SARIF

Reproducible performance benchmark

Performance claims should be reproducible, not just quoted. The repository includes a benchmark that generates synthetic dirty data and measures profiling, full detection/fusion/scoring, numeric-outlier detection, JSONL reading, sampling, score drift, and memory high-water marks.

uv sync
uv run python benchmarks/bench_scan.py 1000000 42
# optional sampling benchmark
uv run python benchmarks/bench_scan.py 1000000 42 --sampling-size 200000

The acceptance and optimization budgets are encoded directly in benchmarks/bench_scan.py, so results can be compared on your own hardware instead of relying on an unspecified machine. See docs/BENCHMARKS.md for the reporting format and benchmark policy.

Architecture

flowchart LR
    Sources[Files / DBs / cloud objects] --> DuckDB[Local execution]
    DuckDB --> Detect[39 detectors]
    Detect --> Evidence[Evidence fusion]
    Evidence --> Score[6-dimension score]
    Score --> Gate[Quality gate]
    Gate --> Reports[Reports / history]
    Reports --> CLI[CLI / TUI]
    Reports --> Web[Web / REST]
    Reports --> MCP[MCP]
    Evidence --> Proposal[Repair proposal]
    Proposal --> Preview[Preview]
    Preview --> Apply[Apply to copy]
    Apply --> Verify[Verify by re-scan]
    Verify --> Rollback[Rollback artifact]
    LLM[Optional OpenAI / Ollama] -. proposes .-> Proposal

Core invariants

  • Local-first — the deterministic core runs locally; cloud LLMs are optional.
  • Deterministic detection — AI is not used to decide whether the base detectors fire.
  • Evidence-backed issues — findings are designed to be inspectable, not opaque scores.
  • Human-in-the-loop repair — state-changing AI suggestions require approval.
  • Reversible changes — repair applies to a copy and creates rollback artifacts.

Engineering

uv sync
make check          # lint + mypy --strict + pytest/coverage gate
make demo           # reproducible demo
make bench          # performance benchmark
make build          # build distributions

The project uses strict typing, automated CI, regression tests, reproducible benchmark gates, and architecture decision records for load-bearing design choices.

Where DataSentry fits

DataSentry is a good fit when you want automatic issue discovery plus a controlled remediation loop in one local-first tool.

If your only requirement is enforcing a small set of already-known contracts, a dedicated rule-first validator may be simpler. If you need a full enterprise metadata catalog, DataSentry is intentionally not trying to replace one. The project focuses on finding bad data, explaining the evidence, and closing the repair loop safely.

Roadmap

The public roadmap is maintained in ROADMAP.md. Near-term priorities emphasize adoption and integration as much as new detectors: broader Python compatibility evaluation, GitHub Actions, reproducible benchmark reporting, dbt/Airflow examples, connectors, and community plugins.

Contributing

Contributions do not need to start with core engine code. Useful first contributions include:

  • new detectors with focused tests;
  • connectors and integration examples;
  • reproducible benchmark cases;
  • documentation and translations;
  • bug reproductions using minimal synthetic data;
  • usability improvements to CLI/TUI/Web workflows.

See CONTRIBUTING.md, CODE_OF_CONDUCT.md, and SECURITY.md.

If you are looking for a small first contribution, check issues labeled good first issue or help wanted.

Documentation

License

Apache-2.0 — see LICENSE.


If DataSentry helped you catch bad data before it reached production, consider giving the repository a ⭐.
It helps other data engineers discover the project.

Release files for datasentry-ai 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datasentry-ai 1.0.1
File Size Uploaded
datasentry_ai-1.0.1.tar.gz 1.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for datasentry-ai 1.0.1
File Interpreter ABI Platform
datasentry_ai-1.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 1.6 MB

Release files / datasentry_ai-1.0.1.tar.gz

Download URL datasentry_ai-1.0.1.tar.gz
Size 1.5 MB
Tags Source
SHA-256 checksum
How to use checksums
f82b0a745e86d690d44a86002324e11951ca94a61019315ecd40d334a7764557
BLAKE2b-256 checksum
How to use checksums
014b5cf4b029495965e1caece22a6c0df4dcac0d7de9761bd562840f818e9469
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.

Transparency log

Release files / datasentry_ai-1.0.1-py3-none-any.whl

Download URL datasentry_ai-1.0.1-py3-none-any.whl
Size 130.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
64cafe30e448a9066ef579de6375e93a0ace6f20fbc3ab96521f18a1642076cb
BLAKE2b-256 checksum
How to use checksums
e5338b4ee8049a35dce5a6b96f1d237d22fd775be230b2ec0d99a6eea0749e50
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.

Transparency log

Release history Release notifications | RSS feed

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

This release

1.0.1 This release

2 release files

1.0.0

2 release files

0.56.0

2 release files

0.55.0

2 release files

0.54.0

2 release files

0.53.0

2 release files

0.52.0

2 release files

0.51.0

2 release files

0.50.0

2 release files

0.49.0

2 release files

0.48.0

2 release files

0.47.0

2 release files

0.46.0

2 release files

0.45.0

2 release files

0.44.0

2 release files

0.43.0

2 release files

0.42.0

2 release files

0.41.0

2 release files

0.40.0

2 release files

0.39.0

2 release files

0.38.0

2 release files

0.37.0

2 release files

0.36.0

2 release files

0.35.0

2 release files

0.34.0

2 release files

0.33.0

2 release files

0.32.0

2 release files

0.31.0

2 release files

0.30.0

2 release files

0.29.0

2 release files

0.28.0

2 release files

0.27.1

2 release files

0.27.0

2 release files

0.26.2

2 release files

0.26.1

2 release files

0.26.0

2 release files

0.25.0

2 release files

0.24.0

2 release files

0.23.0

2 release files

0.22.0

2 release files

0.21.0

2 release files

0.20.0

2 release files

0.19.0

2 release files

0.18.0

2 release files

0.17.0

2 release files

0.16.0

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page