dscribe-dq
Run dScribe data quality rules against your Databricks or MSSQL databases and write the results back to dScribe.
For library internals, architecture, and contributing, see DEVELOPMENT.md.
Prerequisites
- A dScribe account with at least one asset that has data quality rules defined in its ODCS spec
- Your dScribe API key (Settings → API keys in the dScribe UI)
- The asset UUID you want to validate
- Access to the database the rules target (Databricks or MSSQL)
Installation
pip install dscribe-dq
How it works
run_validation fetches your rules from dScribe, runs them with Great Expectations, and returns a DQContext. You then pipe that context through one or more post-processors — small steps that each do one thing (write back to dScribe, upload CSVs, generate reports). This keeps validation and reporting cleanly separated.
run_validation() → DQContext → [step1, step2, ...] → done
The only post-processor included in the SDK is write_back_to_dscribe, which posts pass/fail results back to dScribe so the asset's quality status updates in the UI. Additional post-processors (blob uploads, HTML reports) are available in the separate postprocessors package used by the dScribe runner.
Quickstart
1. Find your asset ID and API key
In the dScribe UI, open the asset you want to validate. The asset ID is the UUID in the URL:
https://app.dscribe.cloud/catalog/assets/337eaa9e-47ed-4b37-a124-050d4932a520
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Your API key is under Settings → API keys.
2. Run validation and write back to dScribe
from dscribe_dq import run_validation, DScribeClient, write_back_to_dscribe
ctx = run_validation(
dscribe_key="<your-api-key>",
asset_id="337eaa9e-47ed-4b37-a124-050d4932a520",
source_configs={
# key must match the server id in the ODCS servers block
"09bcc0f9-9d21-460d-9cb9-942b00e360bf": {
"type": "databricks",
"host": "adb-858283489583940.0.azuredatabricks.net",
"client_id": "<client-id>",
"client_secret": "<client-secret>",
"tenant_id": "<tenant-id>",
"http_path": "/sql/1.0/warehouses/abc123def456",
"catalog": "hive_metastore", # optional
"schema": "default", # optional
}
},
)
client = DScribeClient(api_key="<your-api-key>", base_url="<your-base-url>")
pipeline = [write_back_to_dscribe(client)]
for step in pipeline:
step(ctx)
Each rule in dScribe gets a lastCheckStatus (passed or failed), lastCheckTimestamp, and failure metrics added to its customProperties after the pipeline runs.
Databricks notebook
%pip install dscribe-dq
from dscribe_dq import run_validation, DScribeClient, write_back_to_dscribe
DSCRIBE_API_KEY = dbutils.secrets.get(scope="dscribe-dq", key="DSCRIBE_API_KEY")
DSCRIBE_BASE_URL = "<your-base-url>"
ASSET_ID = "<your-asset-uuid>"
ctx = run_validation(
dscribe_key=DSCRIBE_API_KEY,
base_url=DSCRIBE_BASE_URL,
asset_id=ASSET_ID,
source_configs={
"<server-id>": {
"type": "databricks",
"host": spark.conf.get("spark.databricks.workspaceUrl"),
"client_id": dbutils.secrets.get(scope="dscribe-dq", key="CLIENT_ID"),
"client_secret": dbutils.secrets.get(scope="dscribe-dq", key="CLIENT_SECRET"),
"tenant_id": dbutils.secrets.get(scope="dscribe-dq", key="TENANT_ID"),
"http_path": "/sql/1.0/warehouses/<warehouse-id>",
}
},
)
client = DScribeClient(api_key=DSCRIBE_API_KEY, base_url=DSCRIBE_BASE_URL)
pipeline = [write_back_to_dscribe(client)]
for step in pipeline:
step(ctx)
The
http_pathcan be found in the Databricks UI under SQL Warehouses → your warehouse → Connection details.
Connecting to MSSQL
SQL Server authentication:
source_configs={
"<server-id>": {
"type": "sqlserver",
"host": "your-server.database.windows.net",
"database": "your-db",
"schema": "SalesLT",
"authentication": "SQL Server",
"username": "your-user",
"password": "your-password",
}
}
Entra ID (service principal) authentication:
source_configs={
"<server-id>": {
"type": "sqlserver",
"host": "your-server.database.windows.net",
"database": "your-db",
"schema": "SalesLT",
"authentication": "Entra ID",
"tenant_id": "<tenant-id>",
"client_id": "<client-id>",
"client_secret": "<client-secret>",
}
}
Multiple sources in one call
If your asset has rules targeting both Databricks and MSSQL, pass both in source_configs. Rules are automatically grouped by source and run independently:
source_configs={
"<mssql-server-id>": { "type": "sqlserver", ... },
"<databricks-server-id>": { "type": "databricks", ... },
}
Options
| Parameter | Type | Default | Description |
|---|---|---|---|
dscribe_key |
str | — | dScribe API key (or set DSCRIBE_API_KEY env var) |
asset_id |
str | — | Asset UUID to validate (or set DSCRIBE_ASSET_ID env var) |
base_url |
str | — | dScribe API base URL (or set DSCRIBE_BASE_URL env var) |
source_configs |
dict | {} |
Per-source connection settings keyed by server ID from the ODCS spec |
connector_config |
dict | {} |
Default connection settings used when no per-source config is found |
collect_failed_rows |
bool | True |
Fetch the actual failing rows for each failed rule |
enable_profiling |
bool | False |
Compute descriptive statistics (row count, null counts, distributions) |
log_level |
str | "INFO" |
Logging verbosity: DEBUG, INFO, RESULT, WARNING, ERROR |
Supported ODCS metrics
ODCS metric |
What it checks |
|---|---|
rowCount |
Row count within expected bounds |
nullValues |
No NULL values in a column |
missingValues |
No missing/empty values in a column |
duplicateValues |
All values in a column (or column set) are unique |
invalidValues |
Values match an allowed list or regex pattern |
Environment variable reference
| Variable | Description |
|---|---|
DSCRIBE_API_KEY |
dScribe API key |
DSCRIBE_ASSET_ID |
Asset UUID to validate |
DSCRIBE_BASE_URL |
dScribe API base URL |
DATABRICKS_HOST |
Databricks workspace hostname |
DATABRICKS_CLIENT_ID |
Azure AD service principal client ID |
DATABRICKS_CLIENT_SECRET |
Azure AD service principal client secret |
DATABRICKS_TENANT_ID |
Azure AD tenant ID |
DATABRICKS_HTTP_PATH |
SQL warehouse HTTP path |
DATABRICKS_WAREHOUSE_ID |
SQL warehouse ID (alternative to HTTP path) |
DATABRICKS_CATALOG |
Default Unity Catalog catalog name |
DATABRICKS_SCHEMA |
Default schema name |
MSSQL_HOST |
MSSQL server hostname |
MSSQL_DATABASE |
MSSQL database name |
MSSQL_USER |
SQL Server username |
MSSQL_PASSWORD |
SQL Server password |
MSSQL_AUTH |
SQL Server or Entra ID |
MSSQL_TENANT_ID |
Azure tenant ID (Entra ID auth only) |
MSSQL_CLIENT_ID |
Azure client ID (Entra ID auth only) |
MSSQL_CLIENT_SECRET |
Azure client secret (Entra ID auth only) |
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dscribe_dq-0.0.6.tar.gz.
File metadata
- Download URL: dscribe_dq-0.0.6.tar.gz
- Upload date:
- Size: 21.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
poetry/1.6.1 CPython/3.11.14 Darwin/24.6.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
69a46055e79ffd77aad987edf39d7eb741e9d931b60b3afeaa8da1fc5f172aaf
|
|
| MD5 |
eab56fa08ee89f9b306d9b4205150854
|
|
| BLAKE2b-256 |
f7ad5f5096c9947bc528d1f9a57b290bd4cd955918a56c300ee079a6b8043295
|
File details
Details for the file dscribe_dq-0.0.6-py3-none-any.whl.
File metadata
- Download URL: dscribe_dq-0.0.6-py3-none-any.whl
- Upload date:
- Size: 22.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
poetry/1.6.1 CPython/3.11.14 Darwin/24.6.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9d2f1ed62f6b08952784e12ff49a72291e65596aeddcd1a24a737d228181adc7
|
|
| MD5 |
e18d3950abd896aa30ef53da64d0cdc1
|
|
| BLAKE2b-256 |
9323a94bd547404eebdbc7e6ed4808e621aaa56127cd7d1fa7b226c2281c655c
|