Skip to main content

arrowbricks

Runs SQL against a Databricks SQL warehouse via the Statement Execution API and hands you the result as Arrow -- a Cursor shaped like databricks-sql-python's (execute, fetchone/fetchmany/fetchall, fetchall_arrow/fetchmany_arrow), or stream_query_json for streaming NDJSON. One Arrow engine (arro3), no DuckDB, no pandas/pyarrow.

  • Single responsibility: Databricks to Arrow via arro3. No embedded query engine -- that's duckbricks, built on top of this.
  • Bring-your-own-auth -- a static token or your own token-refresh callable. No cloud-SDK dependency baked in.
  • Result-order preserved even though chunks can complete out of order over the network.
  • Chunks are fetched lazily as fetchone/fetchmany/fetchall actually need them, not all upfront.
  • Heartbeats between slow chunks (execute_streamed/stream_query_json), so a caller streaming this over e.g. SSE never goes silent during a cold warehouse start.

Install

pip install arrowbricks

Dependencies: httpx + arro3-core + arro3-io. That's the whole tree.

Quickstart

import asyncio
from arrowbricks import connect


async def main():
    conn = connect(
        host="adb-1234567890.1.azuredatabricks.net",
        warehouse_id="abcd1234efgh5678",
        token="dapi...",  # or token_provider=... -- see Auth below
    )
    cursor = conn.cursor()

    await cursor.execute("SELECT * FROM my_catalog.my_schema.my_table LIMIT 100")
    async for row in cursor:
        print(row)

    await cursor.execute("SELECT * FROM my_catalog.my_schema.my_table LIMIT 100")
    table = await cursor.fetchall_arrow()  # an arro3 Table


asyncio.run(main())

For streaming NDJSON (e.g. a FastAPI SSE endpoint, first row out as soon as its chunk arrives):

from arrowbricks import HEARTBEAT, DatabricksClient, stream_query_json

client = DatabricksClient(host=..., warehouse_id=..., token=...)

async for item in stream_query_json(client, "SELECT * FROM my_catalog.my_schema.big_table"):
    if item is HEARTBEAT:
        continue  # forward as an SSE keep-alive comment, e.g.
    print(item)  # one ready-to-send JSON string per row

See examples/basic.py for a runnable version, examples/cursor_paging.py for paging a large result with fetchmany/fetchmany_arrow without buffering it all upfront, examples/fastapi_sse.py for streaming a query to a client as Server-Sent Events, examples/fastapi_sse_pivot.py for the same over a buffered Cursor.fetchall_streamed result with one combined heartbeat/timeout budget across both the wait and the download, or examples/azure_auth.py for a caching token_provider built on Azure AD (DefaultAzureCredential).

Why not databricks-sql-connector?

The official driver is the right choice if you need full DB-API 2.0 compatibility over Databricks' Thrift/ODBC-style protocol. If you just want a query result as Arrow/JSON in your own async app, it drags in a lot for that: pandas, thrift, openpyxl, pybreaker, pyjwt, oauthlib, lz4, requests, urllib3 as hard dependencies. arrowbricks talks to the plain REST Statement Execution API instead, and its whole dependency tree is httpx + arro3-core + arro3-io. The Cursor API is deliberately shaped like the official driver's so switching between them is mostly a constructor change, but arrowbricks is async throughout (execute, fetchone, etc. are all coroutines) -- there's no sync escape hatch.

Why not duckbricks?

duckbricks does the same Databricks-to-Arrow work, then goes further: it uses a real embedded DuckDB engine to materialize results into your own DuckDB connection/table (feed_select_to_duckdb_table), or push a DuckDB query's result up to Databricks (feed_duckdb_table_to_databricks). If you need that -- a real local SQL engine sitting on top, not just "run this query, get Arrow/JSON back" -- use duckbricks; it depends on arrowbricks for the Databricks/Arrow half. If you don't need DuckDB at all, arrowbricks alone is the smaller, single-responsibility half.

Auth

connect/DatabricksClient take either:

  • token: str -- a static personal access token or pre-issued OAuth token, or
  • token_provider -- a callable (sync or async) returning a token string, called on every request.

arrowbricks has no opinion on how you get a token and no cloud-SDK dependency of its own. If your provider is expensive to call, cache/refresh inside it -- arrowbricks does no caching on your behalf.

conn = connect(host=..., warehouse_id=..., token_provider=my_token_provider)

API

  • connect(host, warehouse_id, *, token=None, token_provider=None, ...) -> Connection
  • Connection.cursor() -> Cursor
  • Connection.client -> DatabricksClient -- the same client cursor() uses, for lower-level access (e.g. stream_query_json, execute_json_statement, upload_volume_file).
  • Cursor.execute(sql, parameters=None, *, row_limit=None, offset=None, catalog=None, schema=None, total_timeout_s=None) -> Cursor -- submits and waits for the statement, like a real DB-API cursor. parameters, if given, is Databricks' own named-parameter format -- [{"name": ..., "value": ..., "type": ...}] bound against :name markers in sql.
  • Cursor.execute_streamed(...) -- same args, but an async generator yielding HEARTBEAT while waiting on a slow cold start, then the ready Cursor -- for bridging e.g. an SSE connection. Its timeout/heartbeats stop the moment the statement is ready, before any chunk has been downloaded -- see fetchall_streamed below for the download phase itself.
  • Cursor.fetchone() -> tuple | None, Cursor.fetchmany(size) -> list[tuple], Cursor.fetchall() -> list[tuple]
  • Cursor.fetchmany_arrow(size) -> arro3.core.Table, Cursor.fetchall_arrow() -> arro3.core.Table
  • Cursor.fetchall_streamed(*, total_timeout_s=None) / Cursor.fetchall_arrow_streamed(*, total_timeout_s=None) -- like fetchall()/fetchall_arrow(), but yield HEARTBEAT while pulling chunks instead of blocking silently, then the final rows/Table -- for a caller downloading a large result over SSE who needs heartbeats (and a timeout) through the download, not just the initial wait. Compose with execute_streamed and a shared deadline if you want one combined budget across both phases (see examples/fastapi_sse_pivot.py).
  • Cursor is an async iterator, yielding one row (tuple) at a time.
  • Cursor.description -- DB-API-style [(name, type_name, None, None, None, None, None), ...] after execute().
  • stream_query_json(client, sql, **kwargs) -- yields HEARTBEAT, then each row as a JSON string, as soon as its chunk arrives. Timestamps come out as full ISO-8601, every column key is always present ("col":null for a null value, never an omitted key).
  • DatabricksClient(host, warehouse_id, *, token=None, token_provider=None, ...) -- the lower-level client Connection wraps. client.execute_json_statement(sql, ...) for plain JSON rows with no Arrow parse at all; client.upload_volume_file(volume_path, data)/client.delete_volume_file(volume_path) for the Files API.
  • write_ipc_stream(table_or_chunk, buf) -- thin wrapper around arro3.io.write_ipc_stream that always writes uncompressed bodies (see below).

Cursor.execute/execute_streamed/stream_query_json all accept catalog, schema, row_limit, offset, and total_timeout_s.

A note on Arrow IPC compression

write_ipc_stream (and everything in this package that serializes Arrow-IPC bytes) always writes uncompressed bodies. arro3's own default (compression="LZ4") is transparently decompressed by DuckDB's Arrow reader, but not necessarily by every other Arrow IPC reader -- notably, duckdb-wasm's browser-side decoder silently fails to parse LZ4-compressed bodies. If you're producing bytes that might be consumed by something other than a Python DuckDB connection, this default matters.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

arrowbricks-0.2.0.tar.gz (50.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

arrowbricks-0.2.0-py3-none-any.whl (20.0 kB view details)

Uploaded Python 3

File details

Details for the file arrowbricks-0.2.0.tar.gz.

File metadata

  • Download URL: arrowbricks-0.2.0.tar.gz
  • Upload date:
  • Size: 50.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for arrowbricks-0.2.0.tar.gz
Algorithm Hash digest
SHA256 12ce989687b198e2a7ded53f78e1be6d11e547ea189708c8b6368346b9e56faf
MD5 1a1750c78d564683dd3d280bcee0cbd7
BLAKE2b-256 3b23d8b8dccd365da7b095afbd7ae77c3fb7919c00462f4dc7c40975acb932e2

See more details on using hashes here.

Provenance

The following attestation bundles were made for arrowbricks-0.2.0.tar.gz:

Publisher: release.yml on bmsuisse/arrowbricks

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arrowbricks-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: arrowbricks-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 20.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for arrowbricks-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 36620f4958ba0d99d4c8d2fef2440d1494b7cae8d2d8f5d68683a66fe7905ed9
MD5 5c7b41ee6fcff132812db1742db491d8
BLAKE2b-256 41355d6cfbd8ac19edfd030a8f7d3fcd218755ce1192877f4603b1018ee3a29b

See more details on using hashes here.

Provenance

The following attestation bundles were made for arrowbricks-0.2.0-py3-none-any.whl:

Publisher: release.yml on bmsuisse/arrowbricks

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

3.1.2

6 files

3.1.1

6 files

3.1.0

6 files

3.0.4

6 files

3.0.3

6 files

3.0.2

6 files

3.0.1

6 files

3.0.0

6 files

2.0.0

6 files

1.5.0

6 files

1.4.1

6 files

1.4.0

6 files

1.3.3

6 files

1.3.2

6 files

1.3.1

6 files

1.3.0

6 files

1.2.0

6 files

1.1.1

6 files

1.1.0

6 files

1.0.1

6 files

1.0.0

6 files

0.3.0

2 files

This release

0.2.0 This release

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page