Skip to main content

opteryx-iceberg

Read-only Apache Iceberg Metastore/FileIO backend for opteryx-catalog, letting an Opteryx workspace query tables from an external Iceberg catalog (REST, SQL, Hive, Glue - whatever pyiceberg's own catalog loader supports) side by side with native Firestore/GCS-backed tables.

This is Tier 1 of Opteryx's Iceberg support: reads only. Writing real Iceberg tables from Opteryx (Tier 2) and serving Opteryx's own catalog as an Iceberg REST endpoint (Tier 3) are separate, later work.

Kept as its own package - not merged into opteryx-catalog or opteryx-core - because it depends on pyiceberg, which pulls in pyarrow/pydantic. Both of those repos are deliberately free of that dependency chain; Iceberg support is optional, the same way opteryx-access is.

Usage

Register a workspace against an external Iceberg catalog using Opteryx's existing connector-registration API:

from opteryx.connectors import register_workspace
from opteryx.connectors.opteryx_connector import OpteryxConnector
from opteryx_iceberg import IcebergMetastore

register_workspace(
    "my_iceberg_workspace",
    OpteryxConnector,
    catalog=IcebergMetastore,
    catalog_type="rest",       # or "sql", "hive", "glue" - anything pyiceberg's loader supports
    uri="https://...",
    warehouse="s3://...",
)

Do not pass workspace= yourself — OpteryxConnector injects it automatically (as the registered prefix) when it instantiates IcebergMetastore; passing it explicitly raises a duplicate-keyword-argument error.

Native (Firestore/GCS-backed) workspaces are entirely unaffected - this only applies to workspaces explicitly registered with catalog=IcebergMetastore.

Config passes through to pyiceberg verbatim — nesting included. Every kwarg after catalog= is forwarded untouched to pyiceberg.catalog.load_catalog, so pyiceberg's own config shapes (auth={...}, token=, credential=) are used directly. (Earlier versions required flat auth_type/google_auth_scopes kwargs because opteryx-core's connector cache hashed registration kwargs and a dict value broke it; that cache is now keyed by workspace name in opteryx-core's resolution-first connector layer, the flattening is retired, and passing the old flat kwargs raises a clear ValueError.)

Google auth (BigLake and other Google-fronted REST catalogs)

register_workspace(
    "tarchia",
    OpteryxConnector,
    catalog=IcebergMetastore,
    catalog_type="rest",
    uri="https://biglake.googleapis.com/iceberg/v1/restcatalog",
    warehouse="bl://projects/<project>/catalogs/<catalog>",
    auth={"type": "google", "google": {"scopes": ["https://www.googleapis.com/auth/cloud-platform"]}},
    **{"header.x-goog-user-project": "<project>"},
)

auth={"type": "google", ...} selects pyiceberg's built-in GoogleAuthManager, which authenticates via Application Default Credentials and refreshes the token on every request — safe for a long-lived server (a manually fetched gcloud auth print-access-token bearer token, by contrast, expires within the hour and is only good for one-off scripts/tests). In production this picks up Cloud Run's attached service account automatically, the same way the rest of the deployment already does — no explicit credentials_path needed. Stored-credential catalogs need no code at all: pass pyiceberg's token= or credential= the same way.

This is wired into worker.opteryx as the tarchia workspace, alongside the native mabel_data registration - reads from it go through the exact same query path as any native table (verified with a real SELECT ... FROM tarchia.interop_ns.people).

Published on PyPI: worker.opteryx depends on opteryx-iceberg>=0.1.2 as a real dependency in its pyproject.toml, alongside opteryx-core, opteryx-catalog[kms] and opteryx-access - no sys.path sibling-checkout shim, no vendoring. The tarchia registration therefore works in a production Cloud Run deploy the same way it works locally. (The sys.path convention still applies to this repo's own tests, which resolve sibling opteryx-catalog/opteryx-core checkouts - see Local development.)

SQL catalogs: the pyiceberg catalog name must equal the workspace prefix

IcebergMetastore passes the Opteryx workspace prefix straight through as pyiceberg's catalog name: load_catalog(workspace, ...). For REST/Hive/Glue that name is a local label and nothing on the wire depends on it. For catalog_type="sql" it is part of the data. pyiceberg's SqlCatalog stores its name in the catalog_name column of its metadata tables and filters every lookup on it, so a table written under catalog name warehouse is simply not there when read back under the name my_workspace.

The failure is silent and unhelpful: load_dataset gets NoSuchTableError and raises a bare DatasetNotFound, exactly as if the table had never been created. There is no hint that the metadata row exists under a different catalog name.

So when pointing Opteryx at a local/SQL Iceberg catalog, whoever wrote the tables must have used the same catalog name as the workspace prefix you register:

# writer
SqlCatalog("my_workspace", uri="sqlite:///.../catalog.db", warehouse="file:///...")

# reader - the prefix here becomes the pyiceberg catalog name
register_workspace("my_workspace", OpteryxConnector, catalog=IcebergMetastore, catalog_type="sql", ...)

If you already have a SQL catalog written under a different name, either register the Opteryx workspace under that name or rewrite the catalog_name values in the catalog's metadata table.

What's supported

  • SELECT queries against existing Iceberg tables, including predicate pushdown/pruning via standard Iceberg manifest bounds (min_values/max_values/null_counts).
  • Schema introspection (DESCRIBE, information_schema).

What's not (yet)

  • Any write path: CREATE/DROP/ALTER/INSERT/rename all raise NotImplementedError — that's Tier 2.
  • Iceberg views (Iceberg's view spec has no equivalent here yet).
  • Opteryx's own sketch-based pruning stats (min_k_hashes/histograms) — standard Iceberg manifests don't carry them; queries fall back to standard bounds-based pruning.
  • Nested Iceberg types (struct/map/list) — IcebergDataset.schema() raises rather than silently misrepresenting them.

Local development

Sibling opteryx-catalog/opteryx-core checkouts are referenced via sys.path insertion in test files (see tests/), never pip install -e - see those repos' own conventions.

Testing

python -m pytest tests/ -v

Tests run against pyiceberg's own local SqlCatalog (SQLite metadata + local-disk FileIO) — no server, no Docker required.

Real REST-catalog interop check

Snowflake Open Catalog is closed to new signups as of 2026 (Snowflake now points new customers at Horizon Catalog, which needs a full paid-account trial). Instead, real wire-protocol compatibility is verified against Google Lakehouse for Apache Iceberg (BigLake), reusing the existing mabeldev GCP project:

  • Catalog: projects/mabeldev/catalogs/opteryx-iceberg-tier1-test (type biglake, credential-mode end-user), storing data under gs://tarchia/iceberg-tier1-test.
  • Verified manually (not in CI - needs a live GCP access token): dataset_exists, load_dataset, schema() type mapping, scan() including real Iceberg bounds-byte decoding (min_values/max_values/field_ids), and get_relation for both hit and miss, all through opteryx_iceberg.IcebergMetastore against a table (interop_ns.people) written independently via plain pyiceberg.catalog.rest.RestCatalog.
  • Connecting needs GOOGLE_APPLICATION_CREDENTIALS set in-process (not just gcloud auth activate-service-account) — PyArrowFileIO's GCS backend otherwise hangs trying to reach the GCE metadata server for ADC. Warehouse URI format is bl://projects/<project>/catalogs/<catalog> (not a bare projects/... path).
  • This catalog/table is being kept around (not torn down) for reuse in future Tier 1/Tier 2 verification.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

opteryx_iceberg-0.1.5.tar.gz (31.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

opteryx_iceberg-0.1.5-py3-none-any.whl (20.3 kB view details)

Uploaded Python 3

File details

Details for the file opteryx_iceberg-0.1.5.tar.gz.

File metadata

  • Download URL: opteryx_iceberg-0.1.5.tar.gz
  • Upload date:
  • Size: 31.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for opteryx_iceberg-0.1.5.tar.gz
Algorithm Hash digest
SHA256 6bb8e38e8e0ddba0c89231f2f75c7805821c33229dcb216efc999d4e94612415
MD5 086354762bc6af1dfb8bb6dd39642911
BLAKE2b-256 7b5d4845a233b52322c8ce157ad10ae7ab677dcc154ade9ad09b340d471cf93d

See more details on using hashes here.

Provenance

The following attestation bundles were made for opteryx_iceberg-0.1.5.tar.gz:

Publisher: release.yaml on mabel-dev/opteryx-iceberg

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file opteryx_iceberg-0.1.5-py3-none-any.whl.

File metadata

  • Download URL: opteryx_iceberg-0.1.5-py3-none-any.whl
  • Upload date:
  • Size: 20.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for opteryx_iceberg-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 83fe6b347d51cb6559e78b7906c2e27e888fa626280cd31c0db69cfa1e41f06f
MD5 c75c37b5034d7cae4ca27cb51a532977
BLAKE2b-256 3da6fddd808e216052eb6d16835d561a9086b486ab4830a861966380cc11242f

See more details on using hashes here.

Provenance

The following attestation bundles were made for opteryx_iceberg-0.1.5-py3-none-any.whl:

Publisher: release.yaml on mabel-dev/opteryx-iceberg

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.1.9

2 files

0.1.8

2 files

0.1.7

2 files

This release

0.1.5 This release

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page