Turbine
Contract-driven data quality for data products.
Turbine turns a YAML data contract into running quality checks. You declare a table's schema, ownership, freshness expectations, and validity rules in one file; Turbine validates the YAML offline, checks the live database matches it, runs every quality check against the data, scores the result, and exposes everything over a REST API and dashboard.
It uses ODCS v3.1.0, so contracts round-trip with the rest of the data ecosystem.
Example
A contract is a single YAML file:
kind: DataContract
apiVersion: v3.1.0
id: orders
name: Orders
domain: sales
version: "1.0.0"
status: active
description:
purpose: Order records used for billing and fulfilment.
usage: Join to customers on customer_id.
slaProperties:
- property: latency
element: orders.created_at
value: 24
unit: hour
schema:
- name: orders
description: One row per order.
properties:
- name: order_id
logicalType: integer
required: true
primaryKey: true
- name: amount
logicalType: number
required: true
- name: status
logicalType: string
required: true
- name: created_at
logicalType: timestamp
required: true
servers:
- server: default
type: duckdb
database: data/orders.duckdb
schema: main
team:
members:
- username: sales-owner
name: Sales Data Team
role: Data Product Owner
support:
- channel: Data support
tool: slack
scope: issues
url: https://example.test/data-support
customProperties:
- property: availabilityClassification
value: B1
- property: integrityClassification
value: I2
- property: confidentialityClassification
value: V2
- property: privacyClassification
value: P1
Run every check on it:
turbine check contracts/orders.yml --datasource default
You get a per-check verdict, a quality score per dimension (completeness, accuracy, consistency, timeliness, validity), and the failing rows persisted for follow-up.
Installation
Start a new Python project, then add the backend you use:
uv init --python 3.13
uv add "turbine-data[duckdb]" # local files, zero credentials
# uv add "turbine-data[postgres]" # PostgreSQL
# uv add "turbine-data[snowflake]" # Snowflake
# uv add "turbine-data[all]" # every driver + dashboard
In an existing Python project, skip uv init and run the matching uv add command.
Requires Python 3.12 or newer. The PyPI package is turbine-data; the CLI is turbine.
Quick start
# 1. Scaffold a project
uv run turbine init --database duckdb
# 2. Edit the generated contract and DuckDB path, then validate and run
uv run turbine lint src/<project>/contracts/example-contract.yml
uv run turbine check src/<project>/contracts/example-contract.yml --datasource default
The starter contract deliberately leaves team, support, and customProperties
for you to fill in; lint reports those three fields until you add them. DuckDB needs
no credentials file. PostgreSQL and Snowflake initialization create .env.example;
copy it to the ignored .env file and fill its values for those backends.
For a populated multi-backend project, see the manual demo.
Features
- YAML contracts in ODCS v3.1.0 — schema, ownership, SLAs, quality checks in one file
- Quite a few check types — missing, duplicate, invalid values, freshness, row count, custom SQL, Python, group, and window checks (z-score, spike, flatline)
- Schema drift detection — compare your contract to the live database before running a single check
- Dimension-aware scoring — every check is weighted by its quality dimension (completeness, accuracy, consistency, timeliness, validity)
- Row-level flagging — failing rows are persisted in a per-cell bitmap matrix; query which rows failed which checks across runs
- Management API + dashboard —
turbine serveexposes runs, results, scores, and flagged rows over HTTP - Code generation — scaffold SQLModel models and FastAPI routers from contracts
- IDE support — full LSP with VS Code and soon JetBrains extensions
Management API
turbine serve --datasource default --port 8000
Endpoints under /api/v1/manage/: /contracts, /checks/run, /runs/{id}, /runs/{id}/results, /flagged-rows/{table}. Browse /api/v1/manage/docs for the interactive OpenAPI page.
turbine dashboard --port 5173
Renders the same data as charts, run history, and a flagged-rows explorer.
CLI
turbine lint <contract.yml> # validate the YAML offline
turbine validate <contract.yml> --datasource <name> # compare to live database schema
turbine check <contract.yml> --datasource <name> # lint + validate + run every check
turbine status # project health, flagged-row counts
turbine bump # update contract versions
turbine new contract|datasource|check # scaffold a new resource
turbine generate # SQLModel + FastAPI from contracts
Ecosystem
Orchestrators. Run Turbine Check Runs as native steps in your workflow tool:
dagster-turbine— Check Definitions become DagsterAssetCheckSpecs. Partitioned assets scope each Check Run to their partition window.airflow-turbine—TurbineOperatorwith a deferred trigger; tasks wait on a Check Run without blocking a worker.turbine-client— sync + async Python client. Use directly when you need glue beyond the integrations above.
Editors.
- VS Code extension — diagnostics, autocomplete, quick fixes, run-from-editor. Install Turbine from the Marketplace.
- JetBrains plugin — same surface for IntelliJ, PyCharm, DataGrip.
Documentation
Start at the documentation index, then choose the section that matches your task:
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file turbine_data-0.8.3.tar.gz.
File metadata
- Download URL: turbine_data-0.8.3.tar.gz
- Upload date:
- Size: 5.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.2 {"installer":{"name":"uv","version":"0.11.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b56b2878b43f747589ab7fcfbb44b3068ba9619436520af37ed1e81b65a69690
|
|
| MD5 |
107f8b5ad96a5ce2ebc08c8a148c81de
|
|
| BLAKE2b-256 |
9bf6da225d4d10b0916f3b1b37bead5e2dfd7650a5bb4b3f6bf0181dde785578
|
File details
Details for the file turbine_data-0.8.3-py3-none-any.whl.
File metadata
- Download URL: turbine_data-0.8.3-py3-none-any.whl
- Upload date:
- Size: 1.1 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.2 {"installer":{"name":"uv","version":"0.11.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1262223808e6ab50a39943cdd34c89b65e41851708eb4244fc4b92e87f5d9da7
|
|
| MD5 |
80db50a9025b8148ff5894a2bd889469
|
|
| BLAKE2b-256 |
6c8938b5a3cfbacd089da29e78bb197e894204542196889739619c051215920e
|