Skip to main content

hotdata

Official Python client for the Hotdata HTTP API: workspaces, connections, datasets, SQL queries, results, secrets, uploads, indexes, jobs, embedding providers, and workspace context.

Requirements

Python 3.9+

Install

pip install hotdata

For an unreleased revision:

pip install "git+https://github.com/hotdata-dev/sdk-python.git"

From a local checkout (editable):

pip install -e .

Authentication

The API uses an API key sent as Authorization: Bearer <key>, plus an X-Workspace-Id header on requests scoped to a workspace.

import hotdata

configuration = hotdata.Configuration(
    api_key="YOUR_API_KEY",
    workspace_id="YOUR_WORKSPACE_ID",
)

host defaults to https://api.hotdata.dev. Override it if you target another environment.

Usage

import hotdata
from hotdata.rest import ApiException

configuration = hotdata.Configuration(
    api_key="YOUR_API_KEY",
    workspace_id="YOUR_WORKSPACE_ID",
)

with hotdata.ApiClient(configuration) as api_client:
    workspaces = hotdata.WorkspacesApi(api_client)
    try:
        response = workspaces.list_workspaces()
    except ApiException as e:
        print(f"API error: {e.status} {e.reason}\n{e.body}")

Each Api class groups endpoints by resource. Construct the client, then call the typed methods you need.

Arrow results

Query results can be fetched as an Apache Arrow IPC stream instead of JSON, which is faster and far more memory-efficient for large result sets. Install the optional extra:

pip install 'hotdata[arrow]'

Use hotdata.arrow.ResultsApi (a drop-in subclass of ResultsApi that adds Arrow methods):

from hotdata import ApiClient, Configuration
from hotdata.arrow import ResultsApi

with ApiClient(Configuration(api_key="...", workspace_id="...")) as client:
    results = ResultsApi(client)

    # Results are scoped to a database via the required `X-Database-Id` header,
    # so pass the id of the database the query ran in.
    database_id = "your_database_id"

    # Buffered: returns a pyarrow.Table.
    table = results.get_result_arrow(result_id, database_id)

    # Streaming: yields a pyarrow.RecordBatchStreamReader without
    # materializing the full table in memory.
    with results.stream_result_arrow(result_id, database_id) as reader:
        for batch in reader:
            ...

Both methods accept offset and limit for pagination. They raise hotdata.arrow.ResultNotReadyError if the result is still pending or processing — poll results.get_result(result_id, database_id) until status == "ready" first.

File uploads

hotdata.uploads.UploadsApi (also the default hotdata.UploadsApi) adds upload_file, which uploads a local file directly to object storage and finalizes it in one call. It opens an upload session, PUTs the bytes straight to storage — a single PUT for a small file, concurrent part PUTs for a large one — then finalizes. The bytes never round-trip through the API.

from hotdata import ApiClient, Configuration, UploadsApi

with ApiClient(Configuration(api_key="...", workspace_id="...")) as client:
    uploads = UploadsApi(client)

    finalized = uploads.upload_file(
        "data.parquet",
        content_type="application/parquet",
        progress=lambda done, total: print(f"{done}/{total} bytes"),
    )

    # Pass finalized.upload_id to the managed-table load endpoint.
    print(finalized.upload_id)

upload_file accepts a path, raw bytes, or a seekable binary file object (size is inferred for all three; a file object is read from its current position to the end). The SDK picks single vs. multipart from the size, auto-scales the part size, and bounds part concurrency to a peak-memory budget (override with part_size / max_concurrency / part_retry). A file larger than 8 MiB uploads via a streaming session — the SDK mints each part's URL just before its PUT, so a presigned URL can't expire mid-transfer on a slow upload. Storage PUTs go through a dedicated, header-isolated connection pool (auth and workspace headers never reach object storage, which would otherwise reject the upload) with a 30s connect timeout so a dead endpoint fails fast. Finalize is sent with retries disabled so the exactly-once call is never accidentally replayed.

Every failure is a subclass of hotdata.uploads.UploadError (also importable as from hotdata import UploadError), so a single except UploadError catches the whole flow: SessionCreateError (opening the session — check .status for a 501 PRESIGN_UNSUPPORTED), StorageError (storage returned a non-2xx; .exhausted is True if it outlived every retry round), StorageTransportError (the PUT failed before any response), MissingETagError, MintPartError (minting a part URL), FinalizeError, MalformedSessionError, SizeLimitError, and UploadCancelledError. The phase errors that wrap a control-plane call chain the underlying hotdata.exceptions.ApiException as __cause__ (and expose it as .api_exception / .status). A local file read error surfaces as OSError.

The progress callback receives a cumulative (bytes_done, total) — for a tqdm bar (whose update(n) wants a delta, and which isn't thread-safe under multipart) use the ready-made adapter:

from hotdata import tqdm_progress
from tqdm import tqdm

with tqdm(total=size, unit="B", unit_scale=True) as bar:
    uploads.upload_file("data.parquet", progress=tqdm_progress(bar))

Pass a threading.Event as cancel_event to abort an in-flight upload; tune the control-plane calls with request_timeout (storage-PUT timeouts are automatic and size-scaled).

For that fallback (or to upload from a non-seekable stream), use upload_stream, which sends the bytes to the legacy POST /v1/files endpoint in one request, streaming a file object without buffering it in memory:

with open("data.parquet", "rb") as f:
    resp = uploads.upload_stream(f, content_type="application/parquet")
print(resp.id)

Note upload_file shadows the generated raw-body upload_file(body=...); that raw operation is still reachable at hotdata.api.uploads_api.UploadsApi.upload_file.

API reference

Generated Markdown for every operation and model is in docs/:

  • Resource APIs: docs/*Api.md (for example QueryApi.md)
  • Request and response models: docs/<ModelName>.md

Support

Questions and issues: github.com/hotdata-dev/sdk-python.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hotdata-0.7.0.tar.gz (215.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hotdata-0.7.0-py3-none-any.whl (314.1 kB view details)

Uploaded Python 3

File details

Details for the file hotdata-0.7.0.tar.gz.

File metadata

  • Download URL: hotdata-0.7.0.tar.gz
  • Upload date:
  • Size: 215.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for hotdata-0.7.0.tar.gz
Algorithm Hash digest
SHA256 ef4484a74c0f33ee543b0a0dbfb5a6c04e7b5812b277f42a6448000d41c2405d
MD5 3cc3667f10db1c688999125eb083e7c6
BLAKE2b-256 613008681132e019f6c9ddb566dc38be6a302763c4b4848b2d97e3d7ace8d257

See more details on using hashes here.

Provenance

The following attestation bundles were made for hotdata-0.7.0.tar.gz:

Publisher: publish.yml on hotdata-dev/sdk-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hotdata-0.7.0-py3-none-any.whl.

File metadata

  • Download URL: hotdata-0.7.0-py3-none-any.whl
  • Upload date:
  • Size: 314.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for hotdata-0.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 d9e3008e3084d22bcc27b7bf1d08d8c2c255fe3726c252c777bb614a3cae0576
MD5 a1008deda078dbe4fc5ec51d6cc28b5a
BLAKE2b-256 dcef6c4236640629688074892759b1a7593c08f30e9d916b6951c2bbfec19d7f

See more details on using hashes here.

Provenance

The following attestation bundles were made for hotdata-0.7.0-py3-none-any.whl:

Publisher: publish.yml on hotdata-dev/sdk-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.10.0

2 files

0.9.1

2 files

0.9.0

2 files

0.8.0

2 files

This release

0.7.0 This release

2 files

0.6.0

2 files

0.5.0

2 files

0.4.1

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.6

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

0.0.1

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page