DVI — Data Versioning Intelligence
Catch the data incidents that pass every green check.
Quick start · How it works · Signatures · Docs · Contributing · Live demo
DVI is a self-hostable intelligence layer that detects semantic data incidents — the ones where the pipeline runs green but a business number goes silently wrong. It sits on top of your existing stack (dbt, your warehouse, Git); it does not replace any of it. Detection and root-cause ranking are deterministic and explainable — no LLM sits in the decision path.
The problem DVI exists to solve
Data pipelines fail differently from software. A pipeline can execute successfully, keep its schema, meet its freshness SLA, and pass every null/volume check — and still make the business number silently wrong.
-- The whole pipeline stays green. The revenue dashboard is now wrong.
- WHERE status = 'completed'
+ WHERE status = 'COMPLETE'
Before After
country = "UK" country = "United Kingdom"
Schema unchanged. Row count unchanged. Freshness normal. Yet downstream logic
that expects "UK" now under-counts a region, and executive revenue drifts by
double digits.
Structural/volumetric observability tools do not catch this class of failure, because nothing structural changed. DVI is built specifically for it.
Quick start
Requires Python 3.11+. DVI isn't on PyPI yet (#13), so install from source:
git clone https://github.com/anuran-de/dvi.git
cd dvi
python -m venv .venv
source .venv/bin/activate # Windows (Git Bash): source .venv/Scripts/activate
pip install -e ".[dev]"
See it catch a real incident (a deploy silently renames "UK" → "United Kingdom"):
python scripts/demo.py
Structural checks: schema OK | row_count OK | freshness OK | nulls OK
DATA INCIDENT
Severity : HIGH
Confidence : 95% (calibrated, out-of-fold ECE 0.05)
Suspected data incident from change 'deploy #482 (country normalization)'.
Value 'UK' (20.0% of the distribution) appears replaced by 'United Kingdom'
(20.0%) on model.shop.fact_orders; 3 downstream asset(s) affected.
Evidence:
* Change 'deploy #482 ...' was deployed 2 min before the first symptom.
* 'deploy #482 ...' directly changed model.shop.fact_orders, where DVI
observed: Value 'UK' (20.0%) appears replaced by 'United Kingdom' (20.0%).
* No corresponding change was observed upstream of the targeted asset(s).
Use it on your own data
DVI runs from a single dvi.toml describing one asset, its before/after
snapshots, its dbt lineage, and the changes to attribute against:
asset = "model.shop.fct_orders"
columns = ["revenue", "country"] # optional; omit = all shared columns
[source] # "file" or "warehouse"
kind = "file"
before = "artifacts/fct_orders.main.parquet"
after = "artifacts/fct_orders.pr.parquet"
[lineage]
manifest = "target/manifest.json" # dbt manifest → models + exposures
[[changes]] # one or more; RCA attributes to these
id = "pr-1234"
label = "Refactor revenue rollup"
targets = ["model.shop.stg_orders"]
timestamp = 2026-08-30T12:00:00Z
[gate]
fail_on = "high" # low | medium | high | critical
dvi analyze --config dvi.toml --output-dir .dvi
It writes .dvi/dvi-report.md + .dvi/dvi-report.json and sets the exit code:
0 (clean or below gate), 1 (gate tripped), 2 (could not run). Full
reference: docs/cli.md.
In CI (GitHub Action)
DVI ships a composite Action that runs on a pull request and posts the report as a sticky comment, failing the check when the gate trips:
name: DVI
on: pull_request
permissions:
contents: read
pull-requests: write
jobs:
dvi:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: anuran-de/dvi@main
with:
config: dvi.toml
A ready-to-copy workflow ships at .github/workflows/dvi-example.yml.
How it works
DVI turns two snapshots of one asset into an evidence-backed incident (or a clean bill of health):
- Profiles each column over time — including value distributions, not just counts.
- Detects change signatures — deterministic tests that recognise the statistical fingerprint of a semantic change (e.g. a dominant category's mass collapsing while a new value absorbs it) without needing to "understand" meaning.
- Corroborates a detected change against deployments/commits and downstream lineage — an isolated blip stays a symptom; only a change that correlates with a deploy and propagates downstream becomes an incident.
- Ranks likely root causes with evidence — every claim is backed by observable facts, never a fabricated confidence number.
- Names the business-level impact — dbt exposures downstream (dashboards,
ML features, applications) are named by type, and a business-critical
consumer can escalate severity into a
criticaltier — only for a material change.
Two capabilities make it production-ready at scale:
- Warehouse pushdown profiling — step 1 can run inside the warehouse: a
SqlDialect(DuckDB, Snowflake) computes the profile in SQL and only the compact result comes back, so profiling a billion-row table moves a handful of aggregates, not the table. See docs/warehouse-pushdown.md. - Incident history — detection is stateless by default, but an optional
local store gives incidents a stable identity so recurrences dedupe and an
asset's history is queryable over time. Opt in with
[store]indvi.toml. See docs/incident-store.md.
Design principles
- Integrate, don't replace. Reuse dbt's lineage, your warehouse, DuckDB. Build only the intelligence layer from scratch.
- Deterministic first. Detection and ranking are deterministic and explainable. An LLM, if used at all, only narrates evidence — it never decides whether something changed.
- Evidence before explanation. Every root-cause claim carries the observable facts that support it.
- Honest confidence. No hand-tuned "92%". Confidence is either omitted (rank
- evidence) or calibrated and measured on held-out data — a logistic model whose out-of-fold reliability is reported (ECE ≈ 0.05), not asserted.
- Symptom ≠ incident. Corroboration (time × deployment × downstream propagation) is required before anything pages a human.
The signatures
Each is a deterministic test over two column profiles. More specific signatures
suppress more general ones on the same column (a rigid ×100 re-encoding reports
as unit/scale shift, not a generic distribution shift).
| # | Signature | Catches |
|---|---|---|
| 1 | Value substitution | A category renamed/replaced ("UK" → "United Kingdom") |
| 2 | Case/format normalization | Same categories re-spelled ("active" → "ACTIVE", stray whitespace) |
| 3 | Category split/merge | One category fans into many, or many collapse into one |
| 4 | Numeric distribution shift | A behavioral change in shape/location of a numeric column |
| 5 | Unit/scale shift | A rigid re-encoding (dollars → cents, timezone offset) |
Commodity signatures (null-explosion, cardinality, volume, duplicate-rate, schema/type) are slotted in where cheap.
Why you can trust it
DVI is measured against decoys, not just positives — the hard test for a change detector is staying silent when nothing changed.
Benchmark — 100% recall at 0% false positives
python scripts/benchmark.py
The credibility comes from the decoys: legitimate changes that superficially look like incidents and must stay silent (a 2× jump in volume, a new market at 1.5% share, sub-threshold numeric drift). The runner sweeps the one continuous knob (the distribution-shift threshold) to trace the operating curve:
At the shipped default threshold (0.10):
recall : 100%
false-positive rate : 0%
Operating curve (distribution-shift threshold sweep):
threshold recall fp_rate false positives
0.01 100% 43% 3 numeric decoys/negatives fire
0.08 100% 0% -
0.43 80% 0% distribution-shift positive missed
The safe band [0.08, 0.43) gives full recall at zero false positives — a
trade-off that is measured, not asserted. The same runner scores root-cause
ranking under concurrent distractor deploys: 100% top-1 accuracy.
Validated on real data (diamonds, 53,940 rows)
A synthetic benchmark can flatter its own detector. So DVI is validated against a real public dataset — split into two disjoint samples of the same distribution where every fired symptom is, by construction, a false positive.
Validation on real data (diamonds, 53,940 rows)
Real-vs-real false positives: 0/210 column-checks fire (0%) across 30 disjoint splits
Injected-rename recall : 30/30 (100%)
This exposed and fixed a real robustness gap: the first run false-fired on nearly every split, because a share moving 3 points is a real event at 250k rows and pure sampling noise at 250. The fix is a sample-size-aware significance guard. Residual false positives exist only at very small samples (~1% at n=250) and vanish by n=1000.
Calibrated confidence, proven out-of-fold
When the model says 0.7, about 70% of such symptoms are real — and that's proven
on held-out data, not hand-tuned. The model is a small pure-Python logistic
regression (no numpy/sklearn) over three features: magnitude,
significance_margin, and log10 of the sample size.
Calibrated confidence (per-symptom, k-fold cross-validated)
Dataset: 58 fired symptoms, 34 real (59% positive)
Out-of-fold ECE: 0.0466 MCE: 0.2165 Brier: 0.0051
Every row is scored by a model that never trained on it; the shipped model is refit on all data and frozen to JSON, so inference needs no training data.
Documentation
| Doc | What's in it |
|---|---|
| docs/architecture.md | System design and the module map |
| docs/cli.md | dvi.toml reference, dvi analyze, exit codes, the GitHub Action |
| docs/warehouse-pushdown.md | In-warehouse profiling, the executor contract, DuckDB/Snowflake |
| docs/incident-store.md | Persisting incident history across runs |
| docs/frontend.md | The web UI (landing + operator dashboard) and how to deploy it |
| CHANGELOG.md | Full per-milestone history (M1 → M6) |
Live demo: the operator UI is deployed at dvintelligence.vercel.app — an editorial landing page plus an incident dashboard, detail timeline, and blast-radius graph, all rendered from real pipeline output.
Roadmap
DVI is built as a walking skeleton — the riskiest, most novel part (does semantic detection + causal ranking actually work?) is proven first; UI and connectors come last. Every milestone below is complete and green in CI; see the CHANGELOG for the detail.
| Milestone | Adds | Proves |
|---|---|---|
| M1 ✅ | Value-substitution signature end-to-end on synthetic data | The core hypothesis is alive |
| M2 ✅ | Signatures 2–5 + negatives/decoys benchmark + real-data validation | Full recall; 0 false positives on real same-distribution data |
| M3 ✅ | Calibrated logistic confidence + out-of-fold reliability table | Honest, measured confidence (ECE ≈ 0.05) |
| M3.1 ✅ | Review-driven hardening: import-cycle, non-finite, noise floors, determinism | Correctness & honesty under scrutiny |
| M4 ✅ | Blast-radius + external-asset lineage (dashboards/ML/APIs) | Business-level impact |
| M5a ✅ | Warehouse pushdown profiling (DuckDB + Snowflake dialect), detection-equivalence | Pushdown is detection-equivalent to local profiling |
| M5b ✅ | CLI (dvi analyze) + composite GitHub Action posting sticky PR reports |
Real-user adoption path |
| M6 ✅ | Editorial landing page + operator UI, static-exported to Vercel | Operator experience |
Explicitly not built yet
- PyPI package (#13) — install from source until then.
- Automatic BI/ML lineage discovery (Tableau/Looker/feature stores) — downstream assets register via dbt exposures until then.
- Warehouses beyond DuckDB (executed in CI) and Snowflake (dialect +
SQL-gen tests, not CI-executed) — another warehouse needs a new
SqlDialect. - Auto-derived change events —
[[changes]]is declared explicitly indvi.toml; DVI does not yet infer changes from commit/deploy history. - Multi-asset runs — one
dvi analyzerun covers one asset. - Forges beyond GitHub and any autonomous remediation.
Contributing
Contributions are welcome — DVI is early and the open issues are the best place to start.
pip install -e ".[dev]"
pytest # the full suite (also runs the demo + benchmark end to end)
ruff check . # lint
A few conventions this project holds to:
- Test-driven. New behavior lands as a failing test first, then the code to pass it. CI runs the suite twice under different hash seeds as a determinism guard, so avoid depending on set/dict iteration order.
- Deterministic and explainable. Detection and ranking must not depend on an LLM or on unseeded randomness. Confidence numbers are either omitted or measured — never hand-tuned.
- Few dependencies on purpose. The engine uses only polars, duckdb,
networkx, and pydantic. New runtime dependencies need a strong reason (this is
why, e.g., Snowflake's
pyarrow-pulling driver isn't in CI).
Open an issue to discuss anything larger before you build it.
License
MIT — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dvi-0.1.0.tar.gz.
File metadata
- Download URL: dvi-0.1.0.tar.gz
- Upload date:
- Size: 99.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6a082d57964f79afcaac708af106e1960bd72d767503d8482afc51558134910e
|
|
| MD5 |
25b430a30c01f24f244fc9c97bfd5959
|
|
| BLAKE2b-256 |
cd5d1b543fe4e444186d28483962a1c09101522f83c0aac614eb15cdaec7f88b
|
Provenance
The following attestation bundles were made for dvi-0.1.0.tar.gz:
Publisher:
release.yml on anuran-de/dvi
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
dvi-0.1.0.tar.gz -
Subject digest:
6a082d57964f79afcaac708af106e1960bd72d767503d8482afc51558134910e - Sigstore transparency entry: 2702692568
- Sigstore integration time:
-
Permalink:
anuran-de/dvi@4ead21f5164ec04b29f57cc76717aa6fff05f168 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/anuran-de
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@4ead21f5164ec04b29f57cc76717aa6fff05f168 -
Trigger Event:
push
-
Statement type:
File details
Details for the file dvi-0.1.0-py3-none-any.whl.
File metadata
- Download URL: dvi-0.1.0-py3-none-any.whl
- Upload date:
- Size: 77.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ece2b1e78483e806de9e0e927fb71abcd7b5b61b92eb337267614c29dcb442f2
|
|
| MD5 |
232470de94342fdaa5266eb703a3002c
|
|
| BLAKE2b-256 |
95924cf59fbb413448b0f703f7410f58b0ced2a0544223fb3b891442583ab509
|
Provenance
The following attestation bundles were made for dvi-0.1.0-py3-none-any.whl:
Publisher:
release.yml on anuran-de/dvi
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
dvi-0.1.0-py3-none-any.whl -
Subject digest:
ece2b1e78483e806de9e0e927fb71abcd7b5b61b92eb337267614c29dcb442f2 - Sigstore transparency entry: 2702692577
- Sigstore integration time:
-
Permalink:
anuran-de/dvi@4ead21f5164ec04b29f57cc76717aa6fff05f168 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/anuran-de
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@4ead21f5164ec04b29f57cc76717aa6fff05f168 -
Trigger Event:
push
-
Statement type: