Skip to main content

databuck-spark-sdk

Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.

Install

python -m pip install databuck-spark-sdk
python -c "import databuck"

Importing databuck downloads databuck-spark-sdk.jar from the DataBuck S3 URL to databuck/jars/ in the installed Python package, with progress output. The downloaded path is set in DATABUCK_SPARK_SDK_JAR for the current Python process. A later import reuses the existing JAR. The JAR is not included in the wheel because it exceeds PyPI's default per-file upload limit. An internet connection and write access to the installed package directory are required for the first download.

python -m databuck is an explicit download command that also prints the local path. Set DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0 before import to defer the download when offline. DataBuck.jar_path() returns the path after download. Downloading the JAR does not install it as a Databricks compute library.

To use an existing JAR or another writable destination, set DATABUCK_SPARK_SDK_JAR to that file path. If the download URL changes, set DATABUCK_SPARK_SDK_JAR_URL to the new HTTPS URL before importing or running python -m databuck. The download is checked against the published JAR's SHA-256 hash. If the JAR content changes, publish a new Python package version with its new hash, or set DATABUCK_SPARK_SDK_JAR_SHA256 to the expected hash.

For a Databricks notebook, install the Python distribution as a library and make the JAR available on the cluster. If the package directory is read-only, configure DATABUCK_SPARK_SDK_JAR to a writable driver path before importing.

Usage

from pyspark.sql import SparkSession
import os

from databuck import DataBuck

spark = (
    SparkSession.builder
    .config("spark.jars", DataBuck.jar_path())
    .getOrCreate()
)

df = spark.read.csv("customers.csv", header=True)

print(DataBuck.count(df))

Discover and export rules for Databricks

json_path = DataBuck.discover_and_export(
    df,
    "/Volumes/catalog/schema/volume/expectations.json",
    context={
        "business_context": (
            "Missing customer email should be monitored. "
            "Orders without an order ID must be discarded. "
            "An unknown currency must stop publication."
        ),
        "gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
    },
)
print(json_path)

The single call runs profile discovery and context-aware BuckGPT discovery, then writes both sets of row expectations into one JSON file. Context-aware expectations are named BuckGPT_Rule_001, BuckGPT_Rule_002, etc. Their invalid-row SQL queries are converted to Lakeflow pass conditions. Only audited, passed BuckGPT rules in the required SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition> shape can be exported; an unsupported query fails the export. business_context is optional when context contains a Gemini key and any reference PDFs. Paths in pdf_paths must point to existing .pdf files; convert .docx documents to PDF before using them here.

To export only automatic profiling rules, omit context:

json_path = DataBuck.discover_and_export(
    df, "/Volumes/catalog/schema/volume/expectations.json"
)

Without context, BuckGPT is not called and all exported rules use warn.

The JSON contains three dictionaries: warn, drop, and fail. Gemini classifies every exported expectation using the rule expression, DataFrame schema, up to five sample rows, and any optional business context. PDF files are used for BuckGPT rule generation, not passed again to action classification. This sends the schema and sample to Gemini. If an LLM response is incomplete or invalid, export fails instead of assigning a destructive action. In a Lakeflow pipeline, use the dictionaries with dp.expect_all, dp.expect_all_or_drop, and dp.expect_all_or_fail, respectively.

For example, in the Lakeflow pipeline source file:

import json
from pyspark import pipelines as dp

with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
    expectations = json.load(stream)

@dp.table
@dp.expect_all(expectations["warn"])
@dp.expect_all_or_drop(expectations["drop"])
@dp.expect_all_or_fail(expectations["fail"])
def customers_checked():
    return spark.read.table("catalog.schema.customers_source")

The pipeline must read a DataFrame with the columns used by the exported expectations. Regenerate the JSON when the source schema or business policy changes, then refresh the pipeline.

{
  "warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
  "drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
  "fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
}

The example shows the file shape; the actual actions are chosen by Gemini. DataBuck.discover_rules(df) and DataBuck.discover(df, context) are still available separately. The new DataBuck.discover_and_export(...) call combines them.

Each action dictionary maps expectation names to Spark SQL pass conditions, such as {"not_null_subscriber_id": "subscriber_id IS NOT NULL"}. Null, pattern, and length rules are converted from their profiling metadata, including rules from older SDK JARs whose expression field is empty. Dataset-level rules and catalog rules without a row condition remain in rules but are omitted from the file. Null rules are exported only when their threshold is 0%. The export does not encode aggregate failure thresholds. Profiling pattern A maps to [A-Z], and # maps to [0-9].

Metadata

Release files for databuck-spark-sdk 0.5.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for databuck-spark-sdk 0.5.2
File Size Uploaded
databuck_spark_sdk-0.5.2.tar.gz 23.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for databuck-spark-sdk 0.5.2
File Interpreter ABI Platform
databuck_spark_sdk-0.5.2-py3-none-any.whl Python 3 none any Details

Total release size: 42.9 kB

Release files / databuck_spark_sdk-0.5.2.tar.gz

Download URL databuck_spark_sdk-0.5.2.tar.gz
Size 23.3 kB
Tags Source
SHA-256 checksum
How to use checksums
e8ce57b61ca3152e213c1cb57443b4a10cf3ab392f462ab876e1cd5162016ef9
BLAKE2b-256 checksum
How to use checksums
20f37b250b375a61713f2ff92211488122cd1c0e85d706edba03abac0a356f60
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / databuck_spark_sdk-0.5.2-py3-none-any.whl

Download URL databuck_spark_sdk-0.5.2-py3-none-any.whl
Size 19.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0c7251d4f56c2088161cd1b8bec55eb87c4a25600e79ecd12dba15e4a369da6b
BLAKE2b-256 checksum
How to use checksums
960c7a4782e83985cdb3be8dec244e023d63e3a321c2955eeb827643b7ec0488
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.4

2 release files

0.5.3

2 release files

This release

0.5.2 This release

2 release files

0.5.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page