Skip to main content

pypi pyversions

A thin wrapper around pyhive and pyodbc for creating a DBAPI connection to Databricks Workspace and SQL Analytics clusters. SQL Analytics clusters require the Simba ODBC driver.

Also provides SQLAlchemy Dialects using pyhive and pyodbc for Databricks clusters. Databricks SQL Analytics clusters only support the pyodbc-driven dialect.

Installation

Install using pip. You must specify at least one of the extras {hive or odbc}. For odbc the Simba driver is required:

pip install databricks-dbapi[hive,odbc]

For SQLAlchemy support install with:

pip install databricks-dbapi[hive,odbc,sqlalchemy]

Usage

PyHive

The connect() function returns a pyhive Hive connection object, which internally wraps a thrift connection.

Connecting with http_path, host, and a token:

import os

from databricks_dbapi import hive


token = os.environ["DATABRICKS_TOKEN"]
host = os.environ["DATABRICKS_HOST"]
http_path = os.environ["DATABRICKS_HTTP_PATH"]


connection = hive.connect(
    host=host,
    http_path=http_path,
    token=token,
)
cursor = connection.cursor()

cursor.execute("SELECT * FROM some_table LIMIT 100")

print(cursor.fetchone())
print(cursor.fetchall())

The pyhive connection also provides async functionality:

import os

from databricks_dbapi import hive
from TCLIService.ttypes import TOperationState


token = os.environ["DATABRICKS_TOKEN"]
host = os.environ["DATABRICKS_HOST"]
cluster = os.environ["DATABRICKS_CLUSTER"]


connection = hive.connect(
    host=host,
    cluster=cluster,
    token=token,
)
cursor = connection.cursor()

cursor.execute("SELECT * FROM some_table LIMIT 100", async_=True)

status = cursor.poll().operationState
while status in (TOperationState.INITIALIZED_STATE, TOperationState.RUNNING_STATE):
    logs = cursor.fetch_logs()
    for message in logs:
        print(message)

    # If needed, an asynchronous query can be cancelled at any time with:
    # cursor.cancel()

    status = cursor.poll().operationState

print(cursor.fetchall())

ODBC

The ODBC DBAPI requires the Simba ODBC driver.

Connecting with http_path, host, and a token:

import os

from databricks_dbapi import odbc


token = os.environ["DATABRICKS_TOKEN"]
host = os.environ["DATABRICKS_HOST"]
http_path = os.environ["DATABRICKS_HTTP_PATH"]


connection = odbc.connect(
    host=host,
    http_path=http_path,
    token=token,
    driver_path="/path/to/simba/driver",
)
cursor = connection.cursor()

cursor.execute("SELECT * FROM some_table LIMIT 100")

print(cursor.fetchone())
print(cursor.fetchall())

SQLAlchemy Dialects

databricks+pyhive

Installing registers the databricks+pyhive dialect/driver with SQLAlchemy. Fill in the required information when passing the engine URL.

from sqlalchemy import *
from sqlalchemy.engine import create_engine
from sqlalchemy.schema import *


engine = create_engine(
    "databricks+pyhive://token:<databricks_token>@<host>:<port>/<database>",
    connect_args={"http_path": "<cluster_http_path>"}
)

logs = Table("my_table", MetaData(bind=engine), autoload=True)
print(select([func.count("*")], from_obj=logs).scalar())

databricks+pyodbc

Installing registers the databricks+pyodbc dialect/driver with SQLAlchemy. Fill in the required information when passing the engine URL.

from sqlalchemy import *
from sqlalchemy.engine import create_engine
from sqlalchemy.schema import *


engine = create_engine(
    "databricks+pyodbc://token:<databricks_token>@<host>:<port>/<database>",
    connect_args={"http_path": "<cluster_http_path>", "driver_path": "/path/to/simba/driver"}
)

logs = Table("my_table", MetaData(bind=engine), autoload=True)
print(select([func.count("*")], from_obj=logs).scalar())

Refer to the following documentation for more details on hostname, cluster name, and http path:

Metadata

Release files for databricks-dbapi 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for databricks-dbapi 0.6.0
File Size Uploaded
databricks_dbapi-0.6.0.tar.gz 9.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for databricks-dbapi 0.6.0
File Interpreter ABI Platform
databricks_dbapi-0.6.0-py2.py3-none-any.whl Python 2, Python 3 none any Details

Total release size: 19.1 kB

Release files / databricks_dbapi-0.6.0.tar.gz

Download URL databricks_dbapi-0.6.0.tar.gz
Size 9.2 kB
Tags Source
SHA-256 checksum
How to use checksums
8789f375bc866f04c82a720e28f125f895cdc86c7d146eb4b05cec807b1313d1
BLAKE2b-256 checksum
How to use checksums
f71bd2ce1c8f8c83cd7f9ebb4f98d9bccdc080edf4325686b3e29582f6f4eb91
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.1.6 CPython/3.7.4 Darwin/20.1.0

Release files / databricks_dbapi-0.6.0-py2.py3-none-any.whl

Download URL databricks_dbapi-0.6.0-py2.py3-none-any.whl
Size 9.9 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
b0bf0dc008c58aa0b60fa10f23b448388e4eebb54128e8a70c842fa7e8338053
BLAKE2b-256 checksum
How to use checksums
c040362b5058f657bbeff0b4b991105023e4c4a6285a3c57f9cd647673f15b4a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.1.6 CPython/3.7.4 Darwin/20.1.0

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page