hotdata-langchain
Give your LangChain agents access to Hotdata — run SQL against your workspace connections, full-text search indexed columns, and work with managed databases.
Install
pip install hotdata-langchain
Authentication
Set HOTDATA_API_KEY in your environment. Optionally set HOTDATA_WORKSPACE to pin a specific workspace (the first available workspace is used if unset).
Quickstart
The package itself depends only on langchain-core, and works with any tool-calling model.
Running an agent additionally needs the langchain package and the integration for whichever
model provider you use.
from langchain.agents import create_agent
import hotdata_langchain as hl
client = hl.from_env()
tools = hl.make_hotdata_tools(client, database_id="dbid...")
agent = create_agent(model=your_model, tools=tools)
result = agent.invoke(
{"messages": [{"role": "user", "content": "Which product categories have the most orders?"}]}
)
print(result["messages"][-1].content)
Queries run against a database scope, so pass database_id= (a managed database id).
hl.from_env().list_managed_databases() shows what is available in the workspace, with the
id of each.
Tools
make_hotdata_tools(client) returns a list of LangChain StructuredTool objects ready to pass to any agent:
| Tool | What it does |
|---|---|
hotdata_execute_sql |
Run a SQL query and return rows as JSON |
hotdata_list_managed_databases |
List available managed databases, with the id of each |
hotdata_create_managed_database |
Create a new managed database and return its id |
hotdata_load_managed_table |
Load a parquet file into a managed table, addressed by database id |
hotdata_describe_tables |
List tables, or one table's columns and types |
hotdata_search_text |
Full-text search an indexed column, ranked by relevance (opt-in — see below) |
The descriptions carry the engine's contract — dialect, what SQL can and cannot do, and where to look things up — so an agent does not need a system prompt explaining the query engine.
Letting the agent discover the schema
hotdata_describe_tables is registered by default. Called with no arguments it lists every
table with its column count; called with a table name it returns that table's columns and
types. Without it an agent has to guess column names, and a guess that misses fails the query.
tools = hl.make_hotdata_tools(client, database_id="dbid...") # included
tools = hl.make_hotdata_tools(client, database_id="dbid...", describe_tables=False) # omitted
It reads information_schema in whichever database the tools are scoped to, so it needs no
extra permissions. With it turned off, the SQL tool's description tells the agent to query
information_schema directly instead.
Calling tools directly
You can also invoke tools outside of an agent loop:
import json
tools = {t.name: t for t in hl.make_hotdata_tools(client, database_id="dbid...")}
result = tools["hotdata_execute_sql"].invoke({"sql": "SELECT * FROM orders LIMIT 10"})
print(result) # JSON rows
created = tools["hotdata_create_managed_database"].invoke({
"name": "sales", # a display label, not an identifier
"schema_name": "public",
"tables": "orders,customers",
})
tools["hotdata_load_managed_table"].invoke({
"database_id": json.loads(created)["id"],
"table": "orders",
"file": "/path/to/orders.parquet",
})
Full-text search
Point the agent at a text column carrying a BM25 index and it gets a search tool alongside SQL:
tools = hl.make_hotdata_tools(
client,
database_id="dbid...",
search_table="default.public.listings", # catalog.schema.table
search_column="description", # must have a BM25 index
search_columns=["id", "name", "price", "description"], # what each hit returns
search_k=5,
)
hits = {t.name: t for t in tools}["hotdata_search_text"].invoke(
{"query": "cozy apartment with a view"}
)
Rows come back ranked, each with a score. The agent supplies only query and an optional
k; the table and column are fixed when you build the tool. That is deliberate — nothing in
the tool surface lets an agent discover which columns are indexed, and the engine errors
outright rather than falling back to a scan when a column has no BM25 index.
Inside a managed database the built-in catalog is always default, so a managed table reads
as default.<schema>.<table> when database_id= scopes the query to it.
For more than one searchable corpus, build the tools yourself and give each a distinct name and description — the agent then routes on the descriptions:
tools = [
# Configure the first corpus here, so the SQL tool's description still names a search
# tool to defer text matching to. Passing no search_table/search_column drops that,
# and the agent goes back to trying to match text in SQL.
*hl.make_hotdata_tools(
client,
database_id="dbid...",
search_table="default.public.listings",
search_column="description",
search_tool_name="search_listings",
),
hl.make_hotdata_search_tool(
client, table="default.public.reviews", column="comments",
name="search_reviews", database_id="dbid...",
),
]
Provisioning the index itself is not yet part of this package; create it through the Hotdata
API or CLI. demo/ has a script that does the whole flow — managed database, data load, BM25
index, then an agent that picks between search and SQL.
Scoping queries to a managed database
database_id= scopes all SQL the agent runs to one managed database. The API requires a
database scope, so queries fail with a database is required without it:
tools = hl.make_hotdata_tools(client, database_id="dbid...")
Databases are addressed by id, never by name. A database name is a display label and is
not unique, so a name lookup can silently resolve to the wrong database — and the agent's
hotdata_load_managed_table overwrites the table it loads into. Passing a name raises
KeyError. Ids come from client.list_managed_databases(), the
hotdata_list_managed_databases tool, or the response of a create.
The id is resolved once when the tools are built, so a bad id fails there rather than on the
agent's first query, and no query pays a repeat lookup. If you already hold a
ManagedDatabase — from list_managed_databases() or create_managed_database() — pass it
instead of its id to skip the lookup entirely:
db = client.create_managed_database(description="sales", schema="public", tables=["orders"])
tools = hl.make_hotdata_tools(client, database_id=db)
Controlling result size
Limit how many rows are returned to the LLM. Useful for keeping responses within context limits (default: 100):
tools = hl.make_hotdata_tools(client, max_rows=50)
Run the examples
uv run python examples/langchain_basic.py
uv run python examples/langchain_managed_db.py
For a full end-to-end run against a real workspace — data load, BM25 index build, then an
agent choosing between search and SQL — see demo/.
Development
uv sync --locked
uv run pytest
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hotdata_langchain-0.3.0.tar.gz.
File metadata
- Download URL: hotdata_langchain-0.3.0.tar.gz
- Upload date:
- Size: 216.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
aa217a5c53ff8fa66ac589ef1a39f7c39f7371b09785f022e335efa5a248d74f
|
|
| MD5 |
fefdf1277f95e26e1d0a2ee5f7be7795
|
|
| BLAKE2b-256 |
4b0d36cbf6dc250db9c22e77de9272dd0b6dcf9829efdf54025809c1d8198aeb
|
Provenance
The following attestation bundles were made for hotdata_langchain-0.3.0.tar.gz:
Publisher:
publish.yml on hotdata-dev/hotdata-langchain
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hotdata_langchain-0.3.0.tar.gz -
Subject digest:
aa217a5c53ff8fa66ac589ef1a39f7c39f7371b09785f022e335efa5a248d74f - Sigstore transparency entry: 2357637960
- Sigstore integration time:
-
Permalink:
hotdata-dev/hotdata-langchain@952393666874fb8b3f5c439e7ab23f3802698f47 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/hotdata-dev
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@952393666874fb8b3f5c439e7ab23f3802698f47 -
Trigger Event:
push
-
Statement type:
File details
Details for the file hotdata_langchain-0.3.0-py3-none-any.whl.
File metadata
- Download URL: hotdata_langchain-0.3.0-py3-none-any.whl
- Upload date:
- Size: 15.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d36146521d180a4cf77652b383ba249eb845163baa51668880fd4007ae3c7443
|
|
| MD5 |
db718b4854bfbecc2b17ebf519e7dd18
|
|
| BLAKE2b-256 |
b0d34219f7052fa09b923283d914331a7a50a8d7085b3587e0e2a778f986fe39
|
Provenance
The following attestation bundles were made for hotdata_langchain-0.3.0-py3-none-any.whl:
Publisher:
publish.yml on hotdata-dev/hotdata-langchain
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hotdata_langchain-0.3.0-py3-none-any.whl -
Subject digest:
d36146521d180a4cf77652b383ba249eb845163baa51668880fd4007ae3c7443 - Sigstore transparency entry: 2357638521
- Sigstore integration time:
-
Permalink:
hotdata-dev/hotdata-langchain@952393666874fb8b3f5c439e7ab23f3802698f47 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/hotdata-dev
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@952393666874fb8b3f5c439e7ab23f3802698f47 -
Trigger Event:
push
-
Statement type: