Skip to main content

dscribe-dq

Run dScribe data quality rules against your Databricks or MSSQL databases and write the results back to dScribe.

For library internals, architecture, and contributing, see DEVELOPMENT.md.

Prerequisites

  • A dScribe account with at least one asset that has data quality rules defined in its ODCS spec
  • Your dScribe API key (Settings → API keys in the dScribe UI)
  • The asset UUID you want to validate
  • Access to the database the rules target (Databricks or MSSQL)

Installation

pip install dscribe-dq

How it works

run_validation fetches your rules from dScribe, runs them with Great Expectations, and returns a DQContext. You then pipe that context through one or more post-processors — small steps that each do one thing (write back to dScribe, upload CSVs, generate reports). This keeps validation and reporting cleanly separated.

run_validation() → DQContext → [step1, step2, ...] → done

The only post-processor included in the SDK is write_back_to_dscribe, which posts pass/fail results back to dScribe so the asset's quality status updates in the UI. Additional post-processors (blob uploads, HTML reports) are available in the separate postprocessors package used by the dScribe runner.

Quickstart

1. Find your asset ID and API key

In the dScribe UI, open the asset you want to validate. The asset ID is the UUID in the URL:

https://app.dscribe.cloud/catalog/assets/337eaa9e-47ed-4b37-a124-050d4932a520
                                                  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^

Your API key is under Settings → API keys.

2. Run validation and write back to dScribe

from dscribe_dq import run_validation, DScribeClient, write_back_to_dscribe

ctx = run_validation(
    dscribe_key="<your-api-key>",
    asset_id="337eaa9e-47ed-4b37-a124-050d4932a520",
    source_configs={
        # key must match the server id in the ODCS servers block
        "09bcc0f9-9d21-460d-9cb9-942b00e360bf": {
            "type": "databricks",
            "host": "adb-858283489583940.0.azuredatabricks.net",
            "client_id": "<client-id>",
            "client_secret": "<client-secret>",
            "tenant_id": "<tenant-id>",
            "http_path": "/sql/1.0/warehouses/abc123def456",
            "catalog": "hive_metastore",   # optional
            "schema": "default",           # optional
        }
    },
)

client = DScribeClient(api_key="<your-api-key>", base_url="<your-base-url>")

pipeline = [write_back_to_dscribe(client)]

for step in pipeline:
    step(ctx)

Each rule in dScribe gets a lastCheckStatus (passed or failed), lastCheckTimestamp, and failure metrics added to its customProperties after the pipeline runs.

Databricks notebook

%pip install dscribe-dq
from dscribe_dq import run_validation, DScribeClient, write_back_to_dscribe

DSCRIBE_API_KEY = dbutils.secrets.get(scope="dscribe-dq", key="DSCRIBE_API_KEY")
DSCRIBE_BASE_URL = "<your-base-url>"
ASSET_ID = "<your-asset-uuid>"

ctx = run_validation(
    dscribe_key=DSCRIBE_API_KEY,
    base_url=DSCRIBE_BASE_URL,
    asset_id=ASSET_ID,
    source_configs={
        "<server-id>": {
            "type": "databricks",
            "host": spark.conf.get("spark.databricks.workspaceUrl"),
            "client_id": dbutils.secrets.get(scope="dscribe-dq", key="CLIENT_ID"),
            "client_secret": dbutils.secrets.get(scope="dscribe-dq", key="CLIENT_SECRET"),
            "tenant_id": dbutils.secrets.get(scope="dscribe-dq", key="TENANT_ID"),
            "http_path": "/sql/1.0/warehouses/<warehouse-id>",
        }
    },
)

client = DScribeClient(api_key=DSCRIBE_API_KEY, base_url=DSCRIBE_BASE_URL)

pipeline = [write_back_to_dscribe(client)]

for step in pipeline:
    step(ctx)

The http_path can be found in the Databricks UI under SQL Warehouses → your warehouse → Connection details.

Connecting to MSSQL

SQL Server authentication:

source_configs={
    "<server-id>": {
        "type": "sqlserver",
        "host": "your-server.database.windows.net",
        "database": "your-db",
        "schema": "SalesLT",
        "authentication": "SQL Server",
        "username": "your-user",
        "password": "your-password",
    }
}

Entra ID (service principal) authentication:

source_configs={
    "<server-id>": {
        "type": "sqlserver",
        "host": "your-server.database.windows.net",
        "database": "your-db",
        "schema": "SalesLT",
        "authentication": "Entra ID",
        "tenant_id": "<tenant-id>",
        "client_id": "<client-id>",
        "client_secret": "<client-secret>",
    }
}

Multiple sources in one call

If your asset has rules targeting both Databricks and MSSQL, pass both in source_configs. Rules are automatically grouped by source and run independently:

source_configs={
    "<mssql-server-id>": { "type": "sqlserver", ... },
    "<databricks-server-id>": { "type": "databricks", ... },
}

Options

Parameter Type Default Description
dscribe_key str dScribe API key (or set DSCRIBE_API_KEY env var)
asset_id str Asset UUID to validate (or set DSCRIBE_ASSET_ID env var)
base_url str dScribe API base URL (or set DSCRIBE_BASE_URL env var)
source_configs dict {} Per-source connection settings keyed by server ID from the ODCS spec
connector_config dict {} Default connection settings used when no per-source config is found
collect_failed_rows bool True Fetch the actual failing rows for each failed rule
enable_profiling bool False Compute descriptive statistics (row count, null counts, distributions)
log_level str "INFO" Logging verbosity: DEBUG, INFO, RESULT, WARNING, ERROR

Supported ODCS metrics

ODCS metric What it checks
rowCount Row count within expected bounds
nullValues No NULL values in a column
missingValues No missing/empty values in a column
duplicateValues All values in a column (or column set) are unique
invalidValues Values match an allowed list or regex pattern

Environment variable reference

Variable Description
DSCRIBE_API_KEY dScribe API key
DSCRIBE_ASSET_ID Asset UUID to validate
DSCRIBE_BASE_URL dScribe API base URL
DATABRICKS_HOST Databricks workspace hostname
DATABRICKS_CLIENT_ID Azure AD service principal client ID
DATABRICKS_CLIENT_SECRET Azure AD service principal client secret
DATABRICKS_TENANT_ID Azure AD tenant ID
DATABRICKS_HTTP_PATH SQL warehouse HTTP path
DATABRICKS_WAREHOUSE_ID SQL warehouse ID (alternative to HTTP path)
DATABRICKS_CATALOG Default Unity Catalog catalog name
DATABRICKS_SCHEMA Default schema name
MSSQL_HOST MSSQL server hostname
MSSQL_DATABASE MSSQL database name
MSSQL_USER SQL Server username
MSSQL_PASSWORD SQL Server password
MSSQL_AUTH SQL Server or Entra ID
MSSQL_TENANT_ID Azure tenant ID (Entra ID auth only)
MSSQL_CLIENT_ID Azure client ID (Entra ID auth only)
MSSQL_CLIENT_SECRET Azure client secret (Entra ID auth only)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dscribe_dq-0.0.6.tar.gz (21.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dscribe_dq-0.0.6-py3-none-any.whl (22.9 kB view details)

Uploaded Python 3

File details

Details for the file dscribe_dq-0.0.6.tar.gz.

File metadata

  • Download URL: dscribe_dq-0.0.6.tar.gz
  • Upload date:
  • Size: 21.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.6.1 CPython/3.11.14 Darwin/24.6.0

File hashes

Hashes for dscribe_dq-0.0.6.tar.gz
Algorithm Hash digest
SHA256 69a46055e79ffd77aad987edf39d7eb741e9d931b60b3afeaa8da1fc5f172aaf
MD5 eab56fa08ee89f9b306d9b4205150854
BLAKE2b-256 f7ad5f5096c9947bc528d1f9a57b290bd4cd955918a56c300ee079a6b8043295

See more details on using hashes here.

File details

Details for the file dscribe_dq-0.0.6-py3-none-any.whl.

File metadata

  • Download URL: dscribe_dq-0.0.6-py3-none-any.whl
  • Upload date:
  • Size: 22.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.6.1 CPython/3.11.14 Darwin/24.6.0

File hashes

Hashes for dscribe_dq-0.0.6-py3-none-any.whl
Algorithm Hash digest
SHA256 9d2f1ed62f6b08952784e12ff49a72291e65596aeddcd1a24a737d228181adc7
MD5 e18d3950abd896aa30ef53da64d0cdc1
BLAKE2b-256 9323a94bd547404eebdbc7e6ed4808e621aaa56127cd7d1fa7b226c2281c655c

See more details on using hashes here.

Release history Release notifications | RSS feed

1.0.8

2 files

1.0.7

2 files

1.0.6

2 files

1.0.5

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

This release

0.0.6 This release

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page