databuck-spark-sdk
Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.
Install
python -m pip install databuck-spark-sdk
python -c "import databuck"
Importing databuck downloads databuck-spark-sdk.jar from the DataBuck S3
URL to databuck/jars/ in the installed Python package, with progress output.
The downloaded path is set in DATABUCK_SPARK_SDK_JAR for the current Python
process. A later import reuses the existing JAR.
The JAR is not included in the wheel because it exceeds PyPI's
default per-file upload limit. An internet connection and write access to the
installed package directory are required for the first download.
python -m databuck is an explicit download command that also prints the
local path. Set DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0 before import to defer the
download when offline. DataBuck.jar_path() returns the path after download.
Downloading the JAR does not install it as a Databricks compute library.
To use an existing JAR or another writable destination, set
DATABUCK_SPARK_SDK_JAR to that file path. If the download URL changes, set
DATABUCK_SPARK_SDK_JAR_URL to the new HTTPS URL before importing or running
python -m databuck. The download is checked against
the published JAR's SHA-256 hash. If the JAR content changes, publish a new
Python package version with its new hash, or set
DATABUCK_SPARK_SDK_JAR_SHA256 to the expected hash.
For a Databricks notebook, install the Python distribution as a library and
make the JAR available on the cluster. If the package directory is read-only,
configure DATABUCK_SPARK_SDK_JAR to a writable driver path before importing.
Usage
from pyspark.sql import SparkSession
import os
from databuck import DataBuck
spark = (
SparkSession.builder
.config("spark.jars", DataBuck.jar_path())
.getOrCreate()
)
df = spark.read.csv("customers.csv", header=True)
print(DataBuck.count(df))
Discover and export rules for Databricks
json_path = DataBuck.discover_and_export(
df,
"/Volumes/catalog/schema/volume/expectations.json",
context={
"business_context": (
"Missing customer email should be monitored. "
"Orders without an order ID must be discarded. "
"An unknown currency must stop publication."
),
"gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
},
)
print(json_path)
The single call runs profile discovery and context-aware BuckGPT discovery,
then writes both sets of row expectations into one JSON file. Context-aware
expectations are named BuckGPT_Rule_001, BuckGPT_Rule_002, etc. Their
invalid-row SQL queries are converted to Lakeflow pass conditions. Only
audited, passed BuckGPT rules in the required
SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition> shape can be
exported; an unsupported query fails the export.
business_context is optional when context contains a Gemini key and any
reference PDFs. Paths in pdf_paths must point to existing .pdf files;
convert .docx documents to PDF before using them here.
To export only automatic profiling rules, omit context:
json_path = DataBuck.discover_and_export(
df, "/Volumes/catalog/schema/volume/expectations.json"
)
Without context, BuckGPT is not called and all exported rules use warn.
The JSON contains three dictionaries: warn, drop, and fail. Gemini
classifies every exported expectation using the rule expression, DataFrame
schema, up to five sample rows, and any optional business context. PDF files
are used for BuckGPT rule generation, not passed again to action classification.
This sends the schema and sample to Gemini. If an
LLM response is incomplete or invalid, export fails instead of assigning a
destructive action.
In a Lakeflow pipeline, use the dictionaries with dp.expect_all,
dp.expect_all_or_drop, and dp.expect_all_or_fail, respectively.
For example, in the Lakeflow pipeline source file:
import json
from pyspark import pipelines as dp
with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
expectations = json.load(stream)
@dp.table
@dp.expect_all(expectations["warn"])
@dp.expect_all_or_drop(expectations["drop"])
@dp.expect_all_or_fail(expectations["fail"])
def customers_checked():
return spark.read.table("catalog.schema.customers_source")
The pipeline must read a DataFrame with the columns used by the exported expectations. Regenerate the JSON when the source schema or business policy changes, then refresh the pipeline.
{
"warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
"drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
"fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
}
The example shows the file shape; the actual actions are chosen by Gemini.
DataBuck.discover_rules(df) and
DataBuck.discover(df, context) are still available separately. The new
DataBuck.discover_and_export(...) call combines them.
Each action dictionary maps expectation names to Spark SQL pass conditions,
such as {"not_null_subscriber_id": "subscriber_id IS NOT NULL"}. Null, pattern,
and length rules are converted from their profiling metadata, including rules
from older SDK JARs whose expression field is empty. Dataset-level rules
and catalog rules without a row condition remain in rules but are omitted
from the file. Null rules are exported only when their threshold is 0%.
The export does not encode aggregate failure thresholds. Profiling pattern
A maps to [A-Z], and # maps to [0-9].
Metadata
Release files for databuck-spark-sdk 0.5.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| databuck_spark_sdk-0.5.2.tar.gz | 23.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| databuck_spark_sdk-0.5.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 42.9 kB
Release files / databuck_spark_sdk-0.5.2.tar.gz
| Download URL | databuck_spark_sdk-0.5.2.tar.gz |
|---|---|
| Size | 23.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e8ce57b61ca3152e213c1cb57443b4a10cf3ab392f462ab876e1cd5162016ef9
|
|
BLAKE2b-256 checksum How to use checksums |
20f37b250b375a61713f2ff92211488122cd1c0e85d706edba03abac0a356f60
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / databuck_spark_sdk-0.5.2-py3-none-any.whl
| Download URL | databuck_spark_sdk-0.5.2-py3-none-any.whl |
|---|---|
| Size | 19.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0c7251d4f56c2088161cd1b8bec55eb87c4a25600e79ecd12dba15e4a369da6b
|
|
BLAKE2b-256 checksum How to use checksums |
960c7a4782e83985cdb3be8dec244e023d63e3a321c2955eeb827643b7ec0488
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log