Skip to main content

databuck-spark-sdk

Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.

Install

python -m pip install databuck-spark-sdk
python -c "import databuck"

Importing databuck downloads databuck-spark-sdk.jar from the DataBuck S3 URL to databuck/jars/ in the installed Python package, with progress output. The downloaded path is set in DATABUCK_SPARK_SDK_JAR for the current Python process. A later import reuses the existing JAR. The JAR is not included in the wheel because it exceeds PyPI's default per-file upload limit. An internet connection and write access to the installed package directory are required for the first download.

python -m databuck is an explicit download command that also prints the local path. Set DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0 before import to defer the download when offline. DataBuck.jar_path() returns the path after download. Downloading the JAR does not install it as a Databricks compute library.

To use an existing JAR or another writable destination, set DATABUCK_SPARK_SDK_JAR to that file path. If the download URL changes, set DATABUCK_SPARK_SDK_JAR_URL to the new HTTPS URL before importing or running python -m databuck. The download is checked against the published JAR's SHA-256 hash. If the JAR content changes, publish a new Python package version with its new hash, or set DATABUCK_SPARK_SDK_JAR_SHA256 to the expected hash.

For a Databricks notebook, install the Python distribution as a library and make the JAR available on the cluster. If the package directory is read-only, configure DATABUCK_SPARK_SDK_JAR to a writable driver path before importing.

Usage

from pyspark.sql import SparkSession
import os

from databuck import DataBuck

spark = (
    SparkSession.builder
    .config("spark.jars", DataBuck.jar_path())
    .getOrCreate()
)

df = spark.read.csv("customers.csv", header=True)

print(DataBuck.count(df))

Discover and export rules for Databricks

json_path = DataBuck.discover_and_export(
    df,
    "/Volumes/catalog/schema/volume/expectations.json",
    context={
        "business_context": (
            "Missing customer email should be monitored. "
            "Orders without an order ID must be discarded. "
            "An unknown currency must stop publication."
        ),
        "gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
    },
)
print(json_path)

The single call runs profile discovery and context-aware BuckGPT discovery, then writes both sets of row expectations into one JSON file. Context-aware expectations are named BuckGPT_Rule_001, BuckGPT_Rule_002, etc. Their invalid-row SQL queries are converted to Lakeflow pass conditions. Only audited, passed BuckGPT rules in the required SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition> shape can be exported; an unsupported query fails the export. business_context is optional when context contains a Gemini key and any reference documents. The existing pdf_paths field accepts .pdf and .docx files. PDF files are uploaded to Gemini; text from Word documents is extracted locally and included in the generation and audit prompts. Images or scanned pages in a Word document are not read.

To export only automatic profiling rules, omit context:

json_path = DataBuck.discover_and_export(
    df, "/Volumes/catalog/schema/volume/expectations.json"
)

Without context, BuckGPT is not called and all exported rules use warn.

The JSON contains three dictionaries: warn, drop, and fail. Gemini classifies every exported expectation using the rule expression, DataFrame schema, up to five sample rows, and any optional business context. PDF files are used for BuckGPT rule generation, not passed again to action classification. This sends the schema and sample to Gemini. If an LLM response is incomplete or invalid, export fails instead of assigning a destructive action. In a Lakeflow pipeline, use the dictionaries with dp.expect_all, dp.expect_all_or_drop, and dp.expect_all_or_fail, respectively.

For example, in the Lakeflow pipeline source file:

import json
from pyspark import pipelines as dp

with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
    expectations = json.load(stream)

@dp.table
@dp.expect_all(expectations["warn"])
@dp.expect_all_or_drop(expectations["drop"])
@dp.expect_all_or_fail(expectations["fail"])
def customers_checked():
    return spark.read.table("catalog.schema.customers_source")

The pipeline must read a DataFrame with the columns used by the exported expectations. Regenerate the JSON when the source schema or business policy changes, then refresh the pipeline.

{
  "warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
  "drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
  "fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
}

The example shows the file shape; the actual actions are chosen by Gemini. DataBuck.discover_rules(df) and DataBuck.discover(df, context) are still available separately. The new DataBuck.discover_and_export(...) call combines them.

Each action dictionary maps expectation names to Spark SQL pass conditions, such as {"not_null_subscriber_id": "subscriber_id IS NOT NULL"}. Null, pattern, and length rules are converted from their profiling metadata, including rules from older SDK JARs whose expression field is empty. Dataset-level rules and catalog rules without a row condition remain in rules but are omitted from the file. Null rules are exported only when their threshold is 0%. The export does not encode aggregate failure thresholds. Profiling pattern A maps to [A-Z], and # maps to [0-9].

Metadata

Release files for databuck-spark-sdk 0.5.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for databuck-spark-sdk 0.5.3
File Size Uploaded
databuck_spark_sdk-0.5.3.tar.gz 24.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for databuck-spark-sdk 0.5.3
File Interpreter ABI Platform
databuck_spark_sdk-0.5.3-py3-none-any.whl Python 3 none any Details

Total release size: 44.5 kB

Release files / databuck_spark_sdk-0.5.3.tar.gz

Download URL databuck_spark_sdk-0.5.3.tar.gz
Size 24.4 kB
Tags Source
SHA-256 checksum
How to use checksums
a6d8c2f1d1cabd88758c13b41e8812439bf5079a2eca973c3452cadab676ff7a
BLAKE2b-256 checksum
How to use checksums
e74c7d1889c1eee0acdb797aa451c80f6388f6c10b18c9f3e8184b3c8875d2ec
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / databuck_spark_sdk-0.5.3-py3-none-any.whl

Download URL databuck_spark_sdk-0.5.3-py3-none-any.whl
Size 20.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f2e41e0a224a16f33989da9ae813b810243e65b05a65cbdc6e04697b0252fe70
BLAKE2b-256 checksum
How to use checksums
44f496eb8e9ef0027799350560347a6b9baae918d19cf722e9d4318b5159b3e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.4

2 release files

This release

0.5.3 This release

2 release files

0.5.2

2 release files

0.5.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page