Skip to main content

databuck-spark-sdk

Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.

Install

python -m pip install databuck-spark-sdk
python -c "import databuck"

Importing databuck downloads databuck-spark-sdk.jar from the DataBuck S3 URL to databuck/jars/ in the installed Python package, with progress output. The downloaded path is set in DATABUCK_SPARK_SDK_JAR for the current Python process. A later import reuses the existing JAR. The JAR is not included in the wheel because it exceeds PyPI's default per-file upload limit. An internet connection and write access to the installed package directory are required for the first download.

python -m databuck is an explicit download command that also prints the local path. Set DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0 before import to defer the download when offline. DataBuck.jar_path() returns the path after download. Downloading the JAR does not install it as a Databricks compute library.

To use an existing JAR or another writable destination, set DATABUCK_SPARK_SDK_JAR to that file path. If the download URL changes, set DATABUCK_SPARK_SDK_JAR_URL to the new HTTPS URL before importing or running python -m databuck. The download is checked against the published JAR's SHA-256 hash. If the JAR content changes, publish a new Python package version with its new hash, or set DATABUCK_SPARK_SDK_JAR_SHA256 to the expected hash.

For a Databricks notebook, install the Python distribution as a library and make the JAR available on the cluster. If the package directory is read-only, configure DATABUCK_SPARK_SDK_JAR to a writable driver path before importing.

Usage

from pyspark.sql import SparkSession
import os

from databuck import DataBuck

spark = (
    SparkSession.builder
    .config("spark.jars", DataBuck.jar_path())
    .getOrCreate()
)

df = spark.read.csv("customers.csv", header=True)

print(DataBuck.count(df))

Discover and export rules for Databricks

json_path = DataBuck.discover_and_export(
    df,
    "/Volumes/catalog/schema/volume/expectations.json",
    context={
        "business_context": (
            "Missing customer email should be monitored. "
            "Orders without an order ID must be discarded. "
            "An unknown currency must stop publication."
        ),
        "gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
    },
)
print(json_path)

The single call runs profile discovery and context-aware BuckGPT discovery, then writes both sets of row expectations into one JSON file. Context-aware expectations are named BuckGPT_Rule_001, BuckGPT_Rule_002, etc. Their invalid-row SQL queries are converted to Lakeflow pass conditions. Only audited, passed BuckGPT rules in the required SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition> shape can be exported; an unsupported query fails the export. The call prints the auto-discovered rules and the BuckGPT rules before it exports them. Some auto-discovered rules may not be exportable as row-level Lakeflow expectations; the JSON contains only those that can be converted. business_context is optional when context contains a Gemini key and any reference documents. The existing pdf_paths field accepts .pdf and .docx files. PDF files are uploaded to Gemini; text from Word documents is extracted locally and included in the generation and audit prompts. Images or scanned pages in a Word document are not read.

To export only automatic profiling rules, omit context:

json_path = DataBuck.discover_and_export(
    df, "/Volumes/catalog/schema/volume/expectations.json"
)

Without context, BuckGPT is not called and all exported rules use warn.

The JSON contains three dictionaries: warn, drop, and fail. Gemini classifies every exported expectation using the rule expression, DataFrame schema, up to five sample rows, and any optional business context. PDF files are used for BuckGPT rule generation, not passed again to action classification. This sends the schema and sample to Gemini. If an LLM response is incomplete or invalid, export fails instead of assigning a destructive action. In a Lakeflow pipeline, use the dictionaries with dp.expect_all, dp.expect_all_or_drop, and dp.expect_all_or_fail, respectively.

For example, in the Lakeflow pipeline source file:

import json
from pyspark import pipelines as dp

with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
    expectations = json.load(stream)

@dp.table
@dp.expect_all(expectations["warn"])
@dp.expect_all_or_drop(expectations["drop"])
@dp.expect_all_or_fail(expectations["fail"])
def customers_checked():
    return spark.read.table("catalog.schema.customers_source")

The pipeline must read a DataFrame with the columns used by the exported expectations. Regenerate the JSON when the source schema or business policy changes, then refresh the pipeline.

{
  "warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
  "drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
  "fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
}

The example shows the file shape; the actual actions are chosen by Gemini. DataBuck.discover_rules(df) and DataBuck.discover(df, context) are still available separately. The new DataBuck.discover_and_export(...) call combines them.

Each action dictionary maps expectation names to Spark SQL pass conditions, such as {"not_null_subscriber_id": "subscriber_id IS NOT NULL"}. Null, pattern, and length rules are converted from their profiling metadata, including rules from older SDK JARs whose expression field is empty. Dataset-level rules and catalog rules without a row condition remain in rules but are omitted from the file. Null rules are exported only when their threshold is 0%. The export does not encode aggregate failure thresholds. Profiling pattern A maps to [A-Z], and # maps to [0-9].

Metadata

Release files for databuck-spark-sdk 0.5.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for databuck-spark-sdk 0.5.4
File Size Uploaded
databuck_spark_sdk-0.5.4.tar.gz 24.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for databuck-spark-sdk 0.5.4
File Interpreter ABI Platform
databuck_spark_sdk-0.5.4-py3-none-any.whl Python 3 none any Details

Total release size: 44.9 kB

Release files / databuck_spark_sdk-0.5.4.tar.gz

Download URL databuck_spark_sdk-0.5.4.tar.gz
Size 24.7 kB
Tags Source
SHA-256 checksum
How to use checksums
a5dab822f369768211db66a1ac0de6e8faf27beb106302c042e46a722e80f549
BLAKE2b-256 checksum
How to use checksums
aef805496f130e0a8e9d4fb6f2585b8d2ccce477a6dcc322cb420d4f4dfa7805
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / databuck_spark_sdk-0.5.4-py3-none-any.whl

Download URL databuck_spark_sdk-0.5.4-py3-none-any.whl
Size 20.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9d568f1e3b5551905cd9bc53b43bf8afbfefcaea05b8dfccf6e90ff2e6292386
BLAKE2b-256 checksum
How to use checksums
b0a8a651f22394a98fbe07e18db1660b69ce347fab01bdc3ff830d7d13ed465e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.4 This release

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page