Skip to main content

DataSentry

DataSentry

Evidence-driven, local-first AI copilot for data quality.
Detect · Explain · Validate · Repair — with statistical evidence, AI assistance, and human approval.

Release PyPI Python License Tests Coverage GitHub Pages


中文导读:DataSentry 是一个以统计证据为基础、以 AI 为辅助、以人工审批为保障的本地优先数据质量平台。 一次扫描生成六维质量评分,每个问题带证据链;自然语言即可提出规则与修复方案,但只有人工批准才生效。 数据不出机器(LLM 可接本地 Ollama),DuckDB 执行引擎,百万行 10 秒级。 调度体系已就绪:cron 任务队列 → 分布式执行节点 → 多 worker 容错路由 → 并行派发。

What is DataSentry?

DataSentry scans your data (CSV / Parquet / JSONL / XLSX / DuckDB / SQLite / PostgreSQL / MySQL / cloud objects on s3:// gs:// az://) and produces:

  • 39 evidence-driven detectors — missingness, dates, encodings, cross-field rules, cross-table foreign keys, duplicates (exact + fuzzy), outlier models (Isolation Forest / LOF), and more. Every issue carries a statistical evidence chain: samples, ratios, confidence.
  • Six-dimension quality score — completeness, validity, uniqueness, consistency, integrity, timeliness — with explainable weights and per-dimension contributions.
  • Repair loop with human approval — propose → preview (rule re-run before/after) → apply (fingerprinted copy + rollback artifact) → rollback. AI suggests; you decide.
  • Drift engine — compare historical scans: schema, row-count, score and issue-distribution drift.
  • Quality gates in CIscan --fail-on blocks releases by severity or score; export reports as JSON / Markdown / HTML / JUnit / SARIF.
  • LLM assistance, safely — natural language → rule candidates with preflight simulation; PII redacted before any prompt into an encrypted vault with key rotation (llm restore / rotate-key); every call audited (llm status).
  • Cron scheduling — persistent SQLite job queue: cron jobs, manual triggers, run history, webhooks, per-job quality gates and change-aware skip (no re-scan when the source is unchanged).
  • Distributed execution — any instance runs as a worker (datasentry worker); a worker pool gives round-robin routing, failover, cooldown and optional health checks, plus parallel dispatch (DATASENTRY_MAX_WORKERS).
  • Plugin ecosystemplugin.yaml metadata, install/uninstall lifecycle, and SHA-256 integrity locks (tamper-resistant loading, plugin test sandbox).
  • Multiple surfaces — CLI, REST API, server-rendered Web UI with cross-scan trends, and an MCP stdio server (15 tools) so LLM agents can use the tools directly.

Sample quality report

Live demo reportorders-report.html (200 rows with 15 injected quality issues)

Quick start

pip install datasentry-ai     # or: uv sync (source checkout)

datasentry scan orders.csv               # detect → fuse → score → persist, one step
datasentry issues list                   # issues by severity / dimension
datasentry score <run_id>                # six-dimension quality score
datasentry repair propose <issue_id> --file orders.csv   # fix proposal
datasentry drift latest orders           # drift between the two latest scans
datasentry-server                       # Web UI + REST API at http://localhost:8000

Scheduled jobs on remote workers (multi-worker pool with failover; jobs stay in the scheduler's SQLite queue, execution is delegated to datasentry worker nodes):

DATASENTRY_WORKER_TOKEN=<secret> datasentry worker --host 0.0.0.0 --port 8001   # execution node (any instance)
DATASENTRY_WORKERS="http://worker-a:8001:secret;http://worker-b:8001:secret" datasentry-server
# scheduler round-robins jobs across workers; a failing/unreachable worker is
# cooled down (60s) and the next worker takes over; unset DATASENTRY_WORKERS
# to keep running everything locally (zero migration).

# Parallel execution: default is synchronous (one job at a time);
# set a worker count to dispatch due jobs concurrently on a thread pool.
DATASENTRY_MAX_WORKERS=4 datasentry-server

Scan a DuckDB file (optional — any CSV/Parquet/JSONL/XLSX/SQLite works):

datasentry scan analytics.duckdb --table payments
datasentry scan analytics.db --table payments     # SQLite

Scan a MySQL table (via DuckDB mysql extension, no client library; --table required) or a cloud file (CSV/Parquet/ JSONL over s3:// gs:// az://, credentials from process env / secrets):

datasentry scan "mysql://user:pass@localhost:3306/analytics" --table payments
datasentry scan s3://bucket/orders.csv            # AWS credentials from env

Scan a PostgreSQL table (DSN is passed on the command line / via DATASENTRY_PG_DSN and is never persisted or logged):

datasentry scan "postgresql://user:pass@localhost:5432/analytics" --table payments
DATASENTRY_PG_DSN="postgresql://user:pass@localhost:5432/analytics" \
  datasentry scan postgresql:// --table payments --schema public

Credentials

connection_ref resolution chain: process environment variable, then ~/.config/datasentry/secrets.env (overridable via DATASENTRY_CONFIG_HOME or XDG_CONFIG_HOME), then DataSourceNotFoundError:

datasentry secrets set DATASENTRY_PG_DSN      # interactive, no echo, chmod 600
datasentry secrets list                       # key names only (audit-safe)
datasentry secrets get DATASENTRY_PG_DSN
datasentry secrets rm DATASENTRY_PG_DSN

The secrets file uses KEY=VALUE lines (env-var-shaped keys, source-able); the directory is 0700 and the file 0600 — both enforced on read and write. Credentials never enter scan runs, logs, reports, or webhook payloads; all connector errors are redacted (postgresql://*** / passwd=***).

Contract-driven scanning (optional):

datasentry contract validate contract.yaml
datasentry contract export contract.yaml --as pandera   # or --as ge
datasentry scan orders.csv --contract contract.yaml     # gate + rules bound

Architecture

flowchart LR
    subgraph Sources
        CSV[CSV / Parquet / JSONL / XLSX] --> Exec[DuckDB SQL executor]
        DDB[(.duckdb / .db files)] --> Exec
        PG[(PostgreSQL / SQLite / MySQL / cloud)] --> Exec
    end
    Exec --> Dets[39 detectors]
    Dets --> Fuse[Evidence fusion]
    Fuse --> Score[Six-dimension scoring]
    Score --> Gate[Quality gate]
    Gate --> Report[JSON / MD / HTML / JUnit / SARIF]
    Report --> UI[Web UI + trends]
    Report --> MCP[MCP stdio server]
    Report --> CLI[CLI / REST]
    subgraph Scheduling
        Q[(SQLite job queue)] --> Sched[Scheduler + worker thread]
        Sched -->|dispatch| Pool[Worker pool: round-robin + failover + parallel]
        Pool --> W1[Worker A: /rpc/execute]
        Pool --> W2[Worker B: /rpc/execute]
    end
    subgraph AI
        LLM[LLM provider: OpenAI / Ollama]
        Red[PII redaction + encrypted vault]
        Audit[llm_cache + audit]
        LLM --> Red
        Red --> Rules[NL → rule candidates]
        Rules --> Repair[AI repair candidates]
        Audit -.->|every call| Rules
    end
    Repair --> RepairEngine[Repair engine: propose → preview → apply → rollback]
  • Local-first: DuckDB executes everything; LLM is optional (auto-degrades when unconfigured) and can run on local Ollama so data never leaves the machine.
  • Deterministic core: detectors, scoring and repair are pure statistics — no AI guesswork in detection.
  • Human in the loop: rules and repairs are proposals until you approve them; every repair is fingerprinted and rollback-able.

Features

Area What you get
Detection 39 detectors across 6 dimensions; SQL-pushdown single-table; plugin API (plugins/ auto-load, SHA-256 integrity locks)
Scoring 0–100 six-dimension score, severity normalization, contract criticality
Contracts YAML contract DSL → validation + gate + Pandera / Great Expectations export
Repair trim / normalize case / replace missing token / set null / clip values; preview re-runs rules
Drift schema / row-count / score / issue-distribution signals between historical scans
AI NL→rules with preflight + approval gate; AI repair candidates with locked operation surface; PII vault + key rotation
Scheduling cron jobs, manual triggers, run history (pruned), webhooks, quality gates, change-aware skip — CLI / REST / MCP 三面同语义
Distributed datasentry worker nodes; pool routing with failover + cooldown + health checks; parallel dispatch (DATASENTRY_MAX_WORKERS)
Plugins plugin.yaml metadata, install/uninstall, integrity locks, test sandbox (three-state exit codes)
Interfaces CLI · REST API · Web UI (/ui, /ui/trends) · MCP stdio (15 tools)
Engineering 11-stage CI, wheel build + isolated install smoke, 1e6-row benchmark gate

Documentation

Doc Content
docs/DEVELOPMENT.md Full development notes, per-step decisions and conventions
docs/00-设计裁决记录-ADR.md 93 architecture decision records (design rationale)
docs/01-一致性检查.md Spec consistency checks
docs/03-MVP-V1-划分.md MVP vs V1 feature scoping

Development

uv sync
make check          # ruff + mypy --strict + pytest with 85% coverage gate
make demo           # demo script
make bench          # 1e6-row benchmark (60s gate)
make build          # build both wheels (datasentry + datasentry_core)

Requirements: Python ≥ 3.12, uv. CI validates lint, types, coverage, demo, benchmark, API/UI smoke and wheel installability on every push.

Contributing

  • Report issues with the exact data shape (or a minimal CSV) and the command you ran.
  • Code: add a detector → register it in build_initial_detectors → cover it in tests/make check.
  • Every change should reference its ADR decision; see docs/DEVELOPMENT.md for conventions.
  • Please keep the human-in-the-loop invariant: anything AI proposes must remain a proposal until a human approves it.

License

Apache-2.0 — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datasentry_ai-0.21.0.tar.gz (1.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datasentry_ai-0.21.0-py3-none-any.whl (97.7 kB view details)

Uploaded Python 3

File details

Details for the file datasentry_ai-0.21.0.tar.gz.

File metadata

  • Download URL: datasentry_ai-0.21.0.tar.gz
  • Upload date:
  • Size: 1.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for datasentry_ai-0.21.0.tar.gz
Algorithm Hash digest
SHA256 3c85f66ceb51389fe18d7d89bfba24daf02b071f3a7334cdc576340d11a39e73
MD5 9cd2f96cbafbcd3752b3cb01d925a7db
BLAKE2b-256 f9c1ccff4d955b3e6f349dd4b36ed4d764e82e911a6b6e1ad25a78700976d69f

See more details on using hashes here.

Provenance

The following attestation bundles were made for datasentry_ai-0.21.0.tar.gz:

Publisher: publish.yml on Jackxiaozhiren/datasentry

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file datasentry_ai-0.21.0-py3-none-any.whl.

File metadata

  • Download URL: datasentry_ai-0.21.0-py3-none-any.whl
  • Upload date:
  • Size: 97.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for datasentry_ai-0.21.0-py3-none-any.whl
Algorithm Hash digest
SHA256 98807747b03629658a483477dfbc794deabe235e66035950ebd631a7d2b2bcec
MD5 f30e4054094eb50f43ae91ca1fa22149
BLAKE2b-256 de71d05c146adac2663d2e290c99d72b9dc30229e6d020e65bfb6f5fbf35d8da

See more details on using hashes here.

Provenance

The following attestation bundles were made for datasentry_ai-0.21.0-py3-none-any.whl:

Publisher: publish.yml on Jackxiaozhiren/datasentry

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page