Skip to main content

hotdata-ibis

Use Ibis to create on-demand databases, upload data, and query with Python expressions — get pandas or Arrow results back without writing SQL.

Requirements: Python 3.10+, ibis-framework ≥12,<13, hotdata ≥0.7,<0.9.

Install

pip install hotdata-ibis
# or: uv pip install hotdata-ibis

Quickstart: create a database and query it

import time
import pandas as pd
import ibis

con = ibis.hotdata.connect(
    api_url="https://api.hotdata.dev",
    token="YOUR_API_KEY",
    workspace_id="ws_...",
)

# 1. Create a database and declare the tables you'll load.
#    Hotdata database names are not unique — create_database returns the id
#    you'll use for every subsequent operation on this database.
database_id = con.create_database("sales", tables=["orders"])

# 2. Upload a pandas DataFrame (or PyArrow table)
df = pd.DataFrame({
    "order_id": [1, 2, 3],
    "amount": [9.99, 49.99, 5.00],
    "region": ["west", "east", "west"],
})
con.create_table("orders", df, database=(database_id, "main"), overwrite=True)

# 3. Uploads are async — wait briefly before querying
time.sleep(2)

# 4. Query with Ibis expressions
#    Managed tables are always accessed with catalog "default"
t = con.table("orders", database=("default", "main"))
result = (
    t.group_by("region")
    .agg(total=t.amount.sum())
    .order_by(ibis.desc("total"))
    .execute()  # returns a pandas DataFrame
)

# 5. Clean up
con.drop_table("orders", database=(database_id, "main"))
con.drop_database(database_id)

Connect

con = ibis.hotdata.connect(
    api_url="https://api.hotdata.dev",
    token="YOUR_API_KEY",
    workspace_id="ws_...",
    # optional
    timeout=120.0,             # per-request HTTP timeout in seconds
    verify_ssl=True,           # False to skip TLS verification, or path to CA bundle
    default_connection=None,   # default catalog (connection id); auto-detected if only one exists
    default_schema=None,       # default schema; auto-detected if only one exists
    database_id=None,          # bind an existing managed database id at connect time
    poll_interval_s=0.25,      # polling interval for async queries
    poll_timeout_s=600.0,      # max time to wait for a query result
)

URL-style also works, with the same parameters as query string keys:

con = ibis.connect(
    "hotdata://api.hotdata.dev/"
    "?token=...&workspace_id=ws_..."
    "&default_connection=my_conn&default_schema=public"
)

Managed databases

Managed databases are the primary way to bring data into Hotdata with Ibis. Declare a database and its tables, upload data, and query immediately.

Create and load

# Declare the database and all table names up front.
# Hotdata database names are not unique — create_database returns the id
# you'll use for every subsequent operation on this database.
database_id = con.create_database("analytics", tables=["events", "users"])

# Upload from a pandas DataFrame
con.create_table("events", events_df, database=(database_id, "main"), overwrite=True)

# PyArrow tables also work
import pyarrow as pa
table = pa.table({"id": [1, 2], "name": ["alice", "bob"]})
con.create_table("users", table, database=(database_id, "main"), overwrite=True)

# Schema-only (no data): creates an empty table with the declared schema
import ibis.expr.schema as sch
con.create_table(
    "staging",
    schema=sch.Schema({"id": "int64", "ts": "timestamp"}),
    database=(database_id, "main"),
)

Table names must be declared when the database is created — you cannot upload to a table name that was not listed in tables=.

Query

When querying, use "default" as the catalog:

t = con.table("events", database=("default", "main"))

result = (
    t.filter(t.event_type == "click")
    .group_by("user_id")
    .agg(n=t.count())
    .execute()
)

Or with raw SQL:

result = con.sql(
    'SELECT user_id, COUNT(*) AS n '
    'FROM "default"."main"."events" '
    'WHERE event_type = \'click\' '
    'GROUP BY user_id'
).execute()

Delete

Pass force=True to silently skip errors when the database or table does not exist:

con.drop_table("events", database=(database_id, "main"))
con.drop_table("events", database=(database_id, "main"), force=True)  # no-op if missing

con.drop_database(database_id)
con.drop_database(database_id, force=True)  # no-op if missing

Addressing summary

Operation database= argument
create_table / drop_table (database_id, schema) — the id returned by create_database, not its display name
drop_database the id returned by create_database, not its display name
con.table(...) when querying ("default", schema)

Querying

Ibis expressions

t = con.table("orders", database=("default", "main"))

summary = (
    t.filter(t.amount > 10)
    .group_by("region")
    .agg(total=t.amount.sum(), n=t.count())
    .order_by(ibis.desc("total"))
    .execute()
)

.execute() returns a pandas DataFrame. .to_pyarrow() returns an Arrow table. .to_pyarrow_batches() returns a RecordBatchReader — note that Hotdata returns a single Arrow IPC payload per query, so this method downloads the full result first and then splits it into local batches.

Raw SQL

base = con.sql(
    'SELECT * FROM "default"."main"."orders"',
    dialect="postgres",
)
result = base.filter(base.amount > 10).execute()

You can chain Ibis expressions on the result of con.sql(...).

Vector search

ibis_hotdata.vector provides helpers for querying HNSW-indexed vector (embedding) columns:

from ibis_hotdata.vector import semantic_search, l2_distance

t = con.table("docs", database=("default", "main"))

result = semantic_search(t, "embedding", query_vector, k=10).execute()

# or with a different metric
result = semantic_search(t, "embedding", query_vector, k=10, distance_fn=l2_distance).execute()

semantic_search compiles to ORDER BY <distance>(col, ARRAY[...]) ASC LIMIT k with the vector column excluded from the output — the SQL shape Hotdata's query engine requires to route the query through its HNSW index instead of a brute-force scan.

Creating the HNSW index itself isn't wrapped by this package yet — use the hotdata SDK directly:

from hotdata import ApiClient, Configuration
from hotdata.api.indexes_api import IndexesApi
from hotdata.models.create_index_request import CreateIndexRequest

api = IndexesApi(ApiClient(Configuration(...)))
api.create_index(
    connection_id=connection_id,
    var_schema="main",
    table="docs",
    create_index_request=CreateIndexRequest(
        index_name="docs_embedding_idx",
        index_type="vector",
        columns=["embedding"],
        metric="cosine",
    ),
)

Connecting to existing sources

If you have existing databases or warehouses connected to your Hotdata workspace (Postgres, Snowflake, BigQuery, etc.), you can query them through the same Ibis connection:

con = ibis.hotdata.connect(
    api_url="https://api.hotdata.dev",
    token="YOUR_API_KEY",
    workspace_id="ws_...",
    default_connection="my_postgres",
    default_schema="public",
)

t = con.table("orders")  # resolves to my_postgres.public.orders

Discover what's available:

con.list_catalogs()                                    # connection IDs
con.list_databases(catalog="my_postgres")              # schemas
con.list_tables(database=("my_postgres", "public"))    # tables

What's supported

Feature Status
create_database / drop_database (managed) ✅
create_table from pandas / PyArrow / schema-only ✅
drop_table ✅
con.table(...) with full schema metadata ✅
Ibis expressions: filter, select, join, group_by, agg, order_by, limit ✅
con.sql(...) raw SQL ✅
.execute() → pandas, .to_pyarrow(), .to_pyarrow_batches() ✅
list_catalogs, list_databases, list_tables ✅
Arrow / Parquet column types (timestamp, decimal, list, duration, …) ✅
Vector search (ibis_hotdata.vector.semantic_search) ✅
Temporary tables ❌
In-memory tables (ibis.memtable(...)) ❌
Python UDFs ❌
INSERT / UPDATE / DELETE on external connections ❌

SQL compilation uses Ibis's Postgres dialect. Column types returned by Hotdata's information schema are resolved via PyArrow's type system, so Parquet-loaded tables with Arrow-native types (timestamps with time zones, decimals, lists, durations) are mapped correctly to Ibis types.

Development

uv sync   # installs dev group (pytest, ruff, httpx)
uv run pytest
uv run ruff check src tests examples

CI: uv sync --all-groups && uv run pytest -v.

Examples

Set your credentials, then run any example script:

export HOTDATA_API_KEY=...
export HOTDATA_WORKSPACE=...
uv run python examples/01_catalog_introspection.py
uv run python examples/02_execute_sql.py 'SELECT COUNT(*) AS n FROM tpch.tpch_sf1.customer'
uv run python examples/03_connect_via_url.py
uv run python examples/04_ibis_table_workflows.py
uv run python examples/05_roundtrip_demo.py
uv run python examples/06_semantic_search.py

References

Metadata

Release files for hotdata-ibis 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hotdata-ibis 0.5.0
File Size Uploaded
hotdata_ibis-0.5.0.tar.gz 28.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hotdata-ibis 0.5.0
File Interpreter ABI Platform
hotdata_ibis-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 48.1 kB

Release files / hotdata_ibis-0.5.0.tar.gz

Download URL hotdata_ibis-0.5.0.tar.gz
Size 28.4 kB
Tags Source
SHA-256 checksum
How to use checksums
d5e26381779628422ca9bf8cbaccf1dcbfdc589407593b71afcf580afc411c8f
BLAKE2b-256 checksum
How to use checksums
5813447c2c76dad2d4cfa0c1b1216c57525d5ea1546e23cfe4d92cb90d618366
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.

Transparency log

Release files / hotdata_ibis-0.5.0-py3-none-any.whl

Download URL hotdata_ibis-0.5.0-py3-none-any.whl
Size 19.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e41c6ea4fc944486e443e0f4d7864cb5a870fecbbbf8b9954c6860fae843067c
BLAKE2b-256 checksum
How to use checksums
04674a0e129396b3719a264d0d4888aeb2f4c9741ba53995bb6e837b9ee5295e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page