This release is a pre-release and may not be stable for production use.
h2o-connector-service
- Python client: https://pypi.org/project/h2o-connector-service/
- Source: https://github.com/h2oai/connector-service
Python client SDK for the H2O Connector Service. Provides a high-level API to create connectors, open connections, and stream extracted data from supported data sources (PostgreSQL, Snowflake, Hive, Delta Lake, Blob Storage, and more).
pip install h2o-connector-service
Quick Start (H2O Cloud Discovery)
The recommended way to connect when running on H2O AI Cloud:
from h2o_connector_service import ConnectorService
with ConnectorService.from_discovery("https://cloud.h2o.ai", "my-workspace") as svc:
with svc.open_session("postgresql", {
"host": "db.example.com",
"port": "5432",
"database": "mydb",
"username": "user",
"password": "pass",
}, worker_name="pg-worker") as session:
# Stream rows one-by-one (constant memory)
for row in session.stream_records():
print(row)
Quick Start (Manual / Legacy)
For direct connections without H2O Cloud Discovery (deprecated):
from h2o_connector_service import ConnectorService
with ConnectorService("http://localhost:8080", "<your-oidc-token>", "my-workspace") as svc:
with svc.open_session("postgresql", {
"host": "db.example.com",
"port": "5432",
"database": "mydb",
"username": "user",
"password": "pass",
}, worker_name="pg-worker") as session:
for row in session.stream_records():
print(row)
Output Formats
Once you have a session, stream data into various formats:
# CSV file (memory-safe — rows written as they arrive)
session.stream_to_csv("output.csv")
# pandas DataFrame (requires: pip install h2o-connector-service[pandas])
df = session.stream_to_pandas()
# Parquet file (memory-safe, chunked row groups)
# requires: pip install h2o-connector-service[parquet]
session.stream_to_parquet("output.parquet")
# datatable Frame (memory-safe, chunked rbind)
# requires: pip install h2o-connector-service[datatable]
frame = session.stream_to_data_table()
# H2O Frame (requires running H2O cluster + h2o.init())
# requires: pip install h2o-connector-service[h2o]
h2o_frame = session.stream_to_h2o_frame()
# Collect all rows into a list of dicts
records = session.stream_to_records()
Advanced Usage
For full control over the connector lifecycle, use the individual service clients:
from h2o_connector_service import (
Client,
ConnectorServiceClient,
ConnectionServiceClient,
ConnectorSession,
)
with Client.from_discovery("https://cloud.h2o.ai", "my-workspace") as client:
connector_svc = ConnectorServiceClient(client)
conn_svc = ConnectionServiceClient(client)
# 1. Create a connector
connector_svc.create_connector("my-workspace", {
"metadata": {"name": "my-pg"},
"data_source_type": "postgresql",
"data_source_config": {"host": "db.example.com", "port": "5432", "database": "mydb"},
})
# 2. Create a connection (worker must be pre-provisioned by an admin)
connection = conn_svc.create_connection("my-workspace", {
"connector": "workspaces/my-workspace/connectors/my-pg",
"worker": "workspaces/my-workspace/workers/pg-worker",
"extraction": {"query": "SELECT * FROM my_table"},
})
# 3. Wait for the worker pod and stream data
session = ConnectorSession(client, "my-workspace", connection["connection_id"])
session.wait_for_worker_ready(timeout=300)
session.stream_to_csv("output.csv")
Optional Dependencies
Install extras for additional output format support:
pip install h2o-connector-service[pandas] # pandas DataFrames
pip install h2o-connector-service[parquet] # Parquet files (pyarrow)
pip install h2o-connector-service[datatable] # datatable Frames
pip install h2o-connector-service[h2o] # H2O Frames (pandas + pyarrow + h2o)
Supported Data Source Types
data_source_type |
Display Name | Category | Worker Language |
|---|---|---|---|
postgresql |
PostgreSQL | Tabular | Go |
snowflake |
Snowflake | Tabular | Go |
hive |
Apache Hive | Tabular | Java |
delta-lake |
Delta Lake | Tabular | Rust |
s3 |
Amazon S3 | Blob | Go |
gcs |
Google Cloud Storage | Blob | Go |
azure-blob |
Azure Blob Storage | Blob | Go |
minio |
MinIO | Blob | Go |
Release files for h2o-connector-service 0.1.0.dev9001
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| h2o_connector_service-0.1.0.dev9001.tar.gz | 65.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| h2o_connector_service-0.1.0.dev9001-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 158.5 kB
Release files / h2o_connector_service-0.1.0.dev9001.tar.gz
| Download URL | h2o_connector_service-0.1.0.dev9001.tar.gz |
|---|---|
| Size | 65.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8305cf6ea886099fc578d3fea8aad9df7c6c8610ffbe0b64a3363f11f29feb87
|
|
BLAKE2b-256 checksum How to use checksums |
85659cf1923a7e5391619c772c8bad9e5edfd0903757514fd1658556934ff2bb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Release files / h2o_connector_service-0.1.0.dev9001-py3-none-any.whl
| Download URL | h2o_connector_service-0.1.0.dev9001-py3-none-any.whl |
|---|---|
| Size | 93.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
97ec7951c123133033579430a98f2440e92550cb97638f5a5636d52c40e8b8f2
|
|
BLAKE2b-256 checksum How to use checksums |
898d7842f7c9eb90c0dfd491c254933ad5063805e46ec3fec04bc6b35c2ab048
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|