Skip to main content

A library that wraps the databricks-sql-connector for running queries on Databricks

Project description

Databricks Warehouse

A Python library that wraps the Databricks SQL connector (with optional fallbacks to Databricks Connect) and provides a standard interface for querying SQL warehouses from outside of Databricks.

Overview

This library offers a simplified interface for running SQL queries on a Databricks cluster or warehouse and handling authentication using OAuth.

The goal of this library is to make querying Databricks "just work" in a standard way across projects for Data Scientists. This library is ideal for projects like external apps that sit outside of Databricks, run relatively small queries, and don't operate on PySpark DataFrames. Note that for many projects, especially those that leverage PySpark, the best way to interact with Databricks might be with Databricks Connect, not this library.

Installation

SQL Connector Only

To include databricks_warehouse in your project, add it as a dependency in your project's pyproject.toml:

[tool.poetry.dependencies]
...
databricks_warehouse = { path = "../../libraries/databricks_warehouse", develop = True }

Databricks Connect Fallback

The Databricks SQL Connector can only be used from compute that's running outside of Databricks. To support queries from Databricks jobs as well, the preferred method is to use Databricks Connect. All query methods have built-in fallbacks to Databricks Connect which will be run if we detect we're in a Databricks environment. To include this optional feature, make sure to include the connect extra (excluded by default to limit dependencies):

pip install databricks_warehouse[connect]
poetry add databricks_warehouse --extras connect

Primary Methods

The following methods can be imported from databricks_warehouse directly and are the primary ways that users should interact with this library:

  1. read_databricks -- Run a SQL query against an endpoint and return the results as a Pandas DataFrame.
  2. read_databricks_pl -- Run a SQL query against an endpoint and return the results as a Polars DataFrame.
  3. execute_databricks -- Execute a SQL statement and don't return results. This is useful for non-SELECT statements such as INSERT, CREATE, etc.

Configuration

Databricks Client Unified Authentication

This project uses the Databricks SDK for configuring credentials and cluster information. As such, it adheres to the Databricks Client Unified Authentication protocol: https://docs.databricks.com/en/dev-tools/auth/unified-auth.html. This means that any setup you already have for configuring Databricks connections (namely environment variables and your ~/.databrickscfg file) should directly apply here.

Due to security issues with Personal Access Tokens, this library only supports User-to-Machine (U2M) and Machine-to-Machine (M2M) OAuth.

Environment Variables

Set the following environment variables to configure the connection defaults:

Authentication

Always Required:

  • DATABRICKS_HOST: The URL of the Databricks workspace.

Optional:

  • DATABRICKS_CLIENT_ID and DATABRICKS_CLIENT_SECRET: For Machine-to-Machine OAuth using a service principal. Set this when running from automated jobs as a service principal.

If DATABRICKS_CLIENT_ID and DATABRICKS_CLIENT_SECRET are not provided, we fall back to User-to-Machine OAuth using the Databricks CLI instead. See this documentation for more information.

Cluster Specification

Either:

  • DATABRICKS_WAREHOUSE_ID: The ID of the SQL warehouse to run queries on. This should be the primary way the library is used, as SQL warehouses are highly optimized for queries. Cannot be set with DATABRICKS_CLUSTER_ID.

OR

  • DATABRICKS_CLUSTER_ID: The ID of the cluster you wish to connect to. Cannot be set with DATABRICKS_WAREHOUSE_ID.

Configuration Priority Ordering

As this library leverages the Databricks SDK for configuration, we have the same ability to read settings either from environment variables or from a Databricks Configuration File (~/.databrickscfg). In all cases, environment variables take precedence over configuration profiles. See Authentication order of evaluation for more information.

Examples

from databricks_warehouse import read_databricks, read_databricks_pl, execute_databricks

# By default, the target is determined from environment variables or ~/.databrickscfg settings.
# This will run on whatever $DATABRICKS_CLUSTER_ID / $DATABRICKS_WAREHOUSE_ID specify.
# If running remotely from outside of Databricks, this will use the Databricks SQL Connector.
# If running from a Databricks job, this will use Databricks Spark instead.
df = read_databricks("SELECT * FROM your_table LIMIT 10")

# Settings can also be overridden if you require more flexibility than environment variables allow
df = read_databricks(
    "SELECT * FROM your_table LIMIT 10",
    host="https://your-workspace.cloud.databricks.com",
    warehouse_id="your-warehouse-id",
)

# Same, but returns results as Polars DataFrame instead of Pandas
df_polars = read_databricks_pl("SELECT * FROM your_table LIMIT 10")

# For non-SELECT statements (CREATE, DELETE, INSERT, etc.), use `execute_databricks` instead
execute_databricks("CREATE TABLE dev.my_table (...)")

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

databricks_warehouse-1.0.0.tar.gz (5.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

databricks_warehouse-1.0.0-py3-none-any.whl (5.7 kB view details)

Uploaded Python 3

File details

Details for the file databricks_warehouse-1.0.0.tar.gz.

File metadata

  • Download URL: databricks_warehouse-1.0.0.tar.gz
  • Upload date:
  • Size: 5.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.8.3 CPython/3.12.5 Darwin/24.4.0

File hashes

Hashes for databricks_warehouse-1.0.0.tar.gz
Algorithm Hash digest
SHA256 cca6f9b71b0b30ae78ecf7789acffab6340fc9e1f2934f3ed1a2c74bd92bd737
MD5 2608c668bda1539398c614264fabacfe
BLAKE2b-256 bc480545a9ea69e3542a5a81bc61cf231b719f46f9a51dbc449615986a6e3a51

See more details on using hashes here.

File details

Details for the file databricks_warehouse-1.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for databricks_warehouse-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ba3e821622342e52efd90f678516489a6bac7dc56481c7b95805697dad6c3f7f
MD5 7542a0ed0ffcdff88a1d0cf96f4d2b27
BLAKE2b-256 27b6948e1b2d1aa48a2e196240c69d08bd2abf68795b9f848055f9ae371eeadc

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page