Skip to main content

opteryx-iceberg

Read-only Apache Iceberg Metastore/FileIO backend for opteryx-catalog, letting an Opteryx workspace query tables from an external Iceberg catalog (REST, SQL, Hive, Glue - whatever pyiceberg's own catalog loader supports) side by side with native Firestore/GCS-backed tables.

This is Tier 1 of Opteryx's Iceberg support: reads only. Writing real Iceberg tables from Opteryx (Tier 2) and serving Opteryx's own catalog as an Iceberg REST endpoint (Tier 3) are separate, later work.

Kept as its own package - not merged into opteryx-catalog or opteryx-core - because it depends on pyiceberg, which pulls in pyarrow/pydantic. Both of those repos are deliberately free of that dependency chain; Iceberg support is optional, the same way opteryx-access is.

Usage

Register a workspace against an external Iceberg catalog using Opteryx's existing connector-registration API:

from opteryx.connectors import register_workspace
from opteryx.connectors.opteryx_connector import OpteryxConnector
from opteryx_iceberg import IcebergMetastore

register_workspace(
    "my_iceberg_workspace",
    OpteryxConnector,
    catalog=IcebergMetastore,
    catalog_type="rest",       # or "sql", "hive", "glue" - anything pyiceberg's loader supports
    uri="https://...",
    warehouse="s3://...",
)

Do not pass workspace= yourself — OpteryxConnector injects it automatically (as the registered prefix) when it instantiates IcebergMetastore; passing it explicitly raises a duplicate-keyword-argument error.

Native (Firestore/GCS-backed) workspaces are entirely unaffected - this only applies to workspaces explicitly registered with catalog=IcebergMetastore.

Config passes through to pyiceberg verbatim — nesting included. Every kwarg after catalog= is forwarded untouched to pyiceberg.catalog.load_catalog, so pyiceberg's own config shapes (auth={...}, token=, credential=) are used directly. (Earlier versions required flat auth_type/google_auth_scopes kwargs because opteryx-core's connector cache hashed registration kwargs and a dict value broke it; that cache is now keyed by workspace name in opteryx-core's resolution-first connector layer, the flattening is retired, and passing the old flat kwargs raises a clear ValueError.)

Google auth (BigLake and other Google-fronted REST catalogs)

register_workspace(
    "tarchia",
    OpteryxConnector,
    catalog=IcebergMetastore,
    catalog_type="rest",
    uri="https://biglake.googleapis.com/iceberg/v1/restcatalog",
    warehouse="bl://projects/<project>/catalogs/<catalog>",
    auth={"type": "google", "google": {"scopes": ["https://www.googleapis.com/auth/cloud-platform"]}},
    **{"header.x-goog-user-project": "<project>"},
)

auth={"type": "google", ...} selects pyiceberg's built-in GoogleAuthManager, which authenticates via Application Default Credentials and refreshes the token on every request — safe for a long-lived server (a manually fetched gcloud auth print-access-token bearer token, by contrast, expires within the hour and is only good for one-off scripts/tests). In production this picks up Cloud Run's attached service account automatically, the same way the rest of the deployment already does — no explicit credentials_path needed. Stored-credential catalogs need no code at all: pass pyiceberg's token= or credential= the same way.

This is wired into worker.opteryx as the tarchia workspace, alongside the native mabel_data registration - reads from it go through the exact same query path as any native table (verified with a real SELECT ... FROM tarchia.interop_ns.people).

Published on PyPI: worker.opteryx depends on opteryx-iceberg>=0.1.2 as a real dependency in its pyproject.toml, alongside opteryx-core, opteryx-catalog[kms] and opteryx-access - no sys.path sibling-checkout shim, no vendoring. The tarchia registration therefore works in a production Cloud Run deploy the same way it works locally. (The sys.path convention still applies to this repo's own tests, which resolve sibling opteryx-catalog/opteryx-core checkouts - see Local development.)

SQL catalogs: the pyiceberg catalog name must equal the workspace prefix

IcebergMetastore passes the Opteryx workspace prefix straight through as pyiceberg's catalog name: load_catalog(workspace, ...). For REST/Hive/Glue that name is a local label and nothing on the wire depends on it. For catalog_type="sql" it is part of the data. pyiceberg's SqlCatalog stores its name in the catalog_name column of its metadata tables and filters every lookup on it, so a table written under catalog name warehouse is simply not there when read back under the name my_workspace.

The failure is silent and unhelpful: load_dataset gets NoSuchTableError and raises a bare DatasetNotFound, exactly as if the table had never been created. There is no hint that the metadata row exists under a different catalog name.

So when pointing Opteryx at a local/SQL Iceberg catalog, whoever wrote the tables must have used the same catalog name as the workspace prefix you register:

# writer
SqlCatalog("my_workspace", uri="sqlite:///.../catalog.db", warehouse="file:///...")

# reader - the prefix here becomes the pyiceberg catalog name
register_workspace("my_workspace", OpteryxConnector, catalog=IcebergMetastore, catalog_type="sql", ...)

If you already have a SQL catalog written under a different name, either register the Opteryx workspace under that name or rewrite the catalog_name values in the catalog's metadata table.

What's supported

  • SELECT queries against existing Iceberg tables, including predicate pushdown/pruning via standard Iceberg manifest bounds (min_values/max_values/null_counts).
  • Schema introspection (DESCRIBE, information_schema).
  • Time travel: VERSION AS OF <snapshot-id>, VERSION AS OF PREVIOUS (walks Iceberg's parent_snapshot_id), and TIMESTAMP AS OF '<ts>' (point-in-time, resolved against the commit history; a timestamp before the first commit is an error, not an empty result).
  • SHOW MANIFEST FOR <table> — one row per live data file, with the real decoded Iceberg bounds.
  • SHOW SNAPSHOTS FOR <table> — the commit history, newest first.
  • information_schema.tables / .columns, and SHOW COLUMNS / DESCRIBE.
  • Partitioned tables, including predicates over the partition column.
  • Schema evolution: a time-travel read resolves the historical schema the snapshot was written under, so a snapshot taken before an ADD COLUMN does not report the column that did not exist yet.

What SHOW SNAPSHOTS can and cannot tell you about an Iceberg table

opteryx-core's snapshot output was defined against opteryx-catalog's own commit records, which are richer than Iceberg's. IcebergSnapshot (in dataset.py) adapts a pyiceberg snapshot into that shape; three columns are always NULL for an Iceberg-backed table, because the Iceberg spec has nowhere to record them:

Column Why it is null
author Iceberg records no committer identity
commit_message Iceberg records no commit message
user_created opteryx's user-vs-system commit distinction has no Iceberg equivalent

operation_type is the Iceberg spec's own lowercase name (append, overwrite, delete, replace). Note that a row-level delete is reported as overwrite — that is what pyiceberg commits, not a mistranslation here.

The counter columns come from the snapshot summary, which Iceberg holds as strings and populates per operation: an append records no deleted-* counters at all. Those arrive as NULL — meaning "this commit does not report it", not zero, which would claim the commit deleted nothing.

What's not (yet)

  • Merge-on-read delete files. Iceberg v2 can express a delete either by rewriting the data file (copy-on-write) or by committing a delete file beside it (merge-on-read) that the reader must subtract at scan time. This package does not do that subtraction, so a table carrying delete files is refused with a NotImplementedError naming the cause, rather than read — reading it without applying them would return deleted rows as live data, silently.

    Copy-on-write tables are unaffected, and that is everything pyiceberg itself writes (it has no merge-on-read write path). Spark, Flink and Trino do write merge-on-read deletes, so this is the limitation most likely to be met on a remote catalog written by something other than pyiceberg. Lifting it means implementing positional and equality delete application.

  • Any write path: CREATE/DROP/ALTER/INSERT/rename all raise NotImplementedError — that's Tier 2.

  • Iceberg views (Iceberg's view spec has no equivalent here yet).

  • Opteryx's own sketch-based pruning stats (min_k_hashes/histograms) — standard Iceberg manifests don't carry them; queries fall back to standard bounds-based pruning.

  • Nested Iceberg types (struct/map/list) — IcebergDataset.schema() raises rather than silently misrepresenting them. Note this makes a table with any nested column unreadable, including for a query that touches only its scalar columns.

  • Iceberg's own triggers/views listings (list_views/list_triggers return empty, which is the truthful answer for a Tier 1 reader rather than a stub).

Local development

Sibling opteryx-catalog/opteryx-core checkouts are referenced via sys.path insertion in test files (see tests/), never pip install -e - see those repos' own conventions.

Testing

python -m pytest tests/ -v

Tests run against pyiceberg's own local SqlCatalog (SQLite metadata + local-disk FileIO) — no server, no Docker required.

Real REST-catalog interop check

Snowflake Open Catalog is closed to new signups as of 2026 (Snowflake now points new customers at Horizon Catalog, which needs a full paid-account trial). Instead, real wire-protocol compatibility is verified against Google Lakehouse for Apache Iceberg (BigLake), reusing the existing mabeldev GCP project:

  • Catalog: projects/mabeldev/catalogs/opteryx-iceberg-tier1-test (type biglake, credential-mode end-user), storing data under gs://tarchia/iceberg-tier1-test.
  • Verified manually (not in CI - needs a live GCP access token): dataset_exists, load_dataset, schema() type mapping, scan() including real Iceberg bounds-byte decoding (min_values/max_values/field_ids), and get_relation for both hit and miss, all through opteryx_iceberg.IcebergMetastore against a table (interop_ns.people) written independently via plain pyiceberg.catalog.rest.RestCatalog.
  • Connecting needs GOOGLE_APPLICATION_CREDENTIALS set in-process (not just gcloud auth activate-service-account) — PyArrowFileIO's GCS backend otherwise hangs trying to reach the GCE metadata server for ADC. Warehouse URI format is bl://projects/<project>/catalogs/<catalog> (not a bare projects/... path).
  • This catalog/table is being kept around (not torn down) for reuse in future Tier 1/Tier 2 verification.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

opteryx_iceberg-0.1.8.tar.gz (46.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

opteryx_iceberg-0.1.8-py3-none-any.whl (26.7 kB view details)

Uploaded Python 3

File details

Details for the file opteryx_iceberg-0.1.8.tar.gz.

File metadata

  • Download URL: opteryx_iceberg-0.1.8.tar.gz
  • Upload date:
  • Size: 46.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for opteryx_iceberg-0.1.8.tar.gz
Algorithm Hash digest
SHA256 9e2de1bd1865b52ee6cfb476845499848803ac9cdba248ddf044ee4dfdd70b06
MD5 7513fa4b276f336ac79b4a0328876d79
BLAKE2b-256 09771ebef23691d6d8efc1819713995cb9d51ec31a5e4a61afa95b7461e3b732

See more details on using hashes here.

Provenance

The following attestation bundles were made for opteryx_iceberg-0.1.8.tar.gz:

Publisher: release.yaml on mabel-dev/opteryx-iceberg

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file opteryx_iceberg-0.1.8-py3-none-any.whl.

File metadata

  • Download URL: opteryx_iceberg-0.1.8-py3-none-any.whl
  • Upload date:
  • Size: 26.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for opteryx_iceberg-0.1.8-py3-none-any.whl
Algorithm Hash digest
SHA256 67a1b0694faa02be2dcc3254eeece238b4f2e1795a18409375826898b93f1791
MD5 d3848bd96654334512de2a3134a985ab
BLAKE2b-256 b250ee67ffdf35ec2819e86bc94e43ce6048e88e6ec9d737a832f8e82f628839

See more details on using hashes here.

Provenance

The following attestation bundles were made for opteryx_iceberg-0.1.8-py3-none-any.whl:

Publisher: release.yaml on mabel-dev/opteryx-iceberg

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.1.9

2 files

This release

0.1.8 This release

2 files

0.1.7

2 files

0.1.5

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page