Skip to main content

Opteryx Core

Opteryx Core is the SQL execution engine behind opteryx.app. It is a fork of Opteryx with a smaller, more opinionated API and configuration surface, shaped around the workloads used by the hosted service.

This library is designed for fast, read-heavy analytical queries over columnar data. It handles SQL parsing, planning, predicate pushdown, projection pruning, and execution so you can query datasets from Python without standing up a separate warehouse.

Query planning is Python; query execution is native. Once the planner has produced a physical plan, the engine runs it in compiled code end to end — scan, operators, scheduling, and dispatch — and neither PyArrow nor NumPy is present anywhere in the engine.

This project is opinionated toward the needs of opteryx.app. It is still useful as a standalone library if you want to query local Parquet, NDJSON, CSV, and .skene datasets, embed SQL into a Python service or notebook, or experiment with engine internals directly.

Requirements

  • Python 3.11 or later
  • A C/C++ toolchain for local source builds
  • Rust/Cargo for the Rust extension in src/

Install

pip install opteryx-core

Import it as:

import opteryx

Quick Start: Query Local Files

If your current working directory contains local Parquet data, the simplest way to use Opteryx Core is to register a local workspace and query it with dot-separated names.

import opteryx
from opteryx.connectors import DiskConnector

opteryx.register_workspace("data", DiskConnector)

session = opteryx.session()

for morsel in session.execute_to_morsels("SELECT id, name FROM data.planets WHERE id < 5"):
    print(morsel)

Results arrive as Draken morsels — batches of columns — streamed as the engine produces them, so a large result never has to fit in memory at once. A morsel prints as a table, and carries num_rows, column_names and column(name).to_pylist() for getting at the values. Once the stream has been read to the end, session.rowcount is the number of rows it delivered.

In this model, dataset names are resolved relative to the current working directory. For example, data.planets resolves to ./data/planets, and Opteryx Core reads the files it finds there, detecting the format from their extension. See File Formats for what it can read.

There is also a command line, for querying without writing Python:

python -m opteryx "SELECT id, name FROM data.planets WHERE id < 5"

What It Is For

  • Powering the execution layer used by opteryx.app
  • Running analytical SQL against local Parquet, CSV, JSONL, and .skene datasets
  • Embedding a query engine inside Python applications, scripts, notebooks, and services
  • Working on engine internals such as planning, native execution, and file-format performance
  • Using the file engine or the .skene format on their own, via the rugo and libskene wheels

Local Development

The supported local build path is the repository Makefile:

make dev-install
make compile
make q

Useful targets:

Target Purpose
make compile Clean in-place build of Cython, C++, and Rust extensions
make c Incremental extension build
make q Fast SQL shape smoke test
make test Full pytest suite after compiling
make dt Draken native unit tests
make check Ruff and import-order checks without modifying files

Do not use pip install . as the primary development build path; make compile matches the layout expected by this repository.

Repository Layout

Path Purpose
opteryx/ SQL engine — parser bindings, binder, optimizer, physical planner, connectors, and the native execution engine
draken/ Native columnar vector substrate (DrakenVector) and morsels; zero external dependencies
rugo/ File engine — Parquet, CSV, and JSONL read and write. Also published standalone; the source is opteryx-free
skene/ The .skene columnar file format — C++ reader, writer, and normative specification. Also published standalone
src/ Opteryx compute extension sources: Rust (opteryx_dialect.rs) and C++ (src/cpp/)
reference/ Generated catalog snapshots (functions, operators, types, joins, clauses). Source of truth for code generation — regenerated by make reference, never hand-edited
tests/ Unit, integration, fuzzing, sqllogictest, and benchmark harnesses
testdata/ Local datasets and benchmark fixtures
docs/ Design documents and engineering notes (user documentation lives at docs.opteryx.app)
dev/ Development, release, vendoring, and analysis scripts; never imported by production code
scripts/ CI helper scripts
scratch/ Experimental prototypes and one-off investigations; not packaged
third_party/ Vendored native dependencies
build_common.py Shared build machinery and the single-source extension definitions for draken, rugo, and skene

Distributions

This is a single repository that produces three wheels from one source tree. They are packagings of the same sources, not separate forks, so they cannot drift: the extension definitions are single-sourced in build_common.py.

Wheel Import as Contains For
opteryx-core opteryx The full SQL engine, bundling draken, rugo, and skene Querying data with SQL — the primary distribution
rugo rugo The file engine (Parquet, CSV, JSONL) plus draken Reading and writing files without the SQL engine
libskene skene The .skene format reader and writer plus draken Lossless draken-vector serialization on its own

draken is not published separately; it ships inside each of the three. rugo and skene are parallel — neither depends on the other, and the rugo wheel does not carry skene. Opteryx never depends on the published rugo or libskene wheels; those components are intrinsic to it, and the standalone wheels are separate packagings of the same code.

Wheels are built in CI, never locally. For local development use make compile, as above.

File Formats

Datasets are read by extension, and a dataset is one format throughout — a directory mixing formats is an error rather than a best-effort read.

  • Parquet — the default for stored data and for interchange, read through rugo
  • CSV and JSONL/NDJSON — read through rugo
  • .skene — the draken-native format. It stores one or more row groups of draken vectors losslessly, including the things Parquet drops: an IPv4 column round-trips as a UINT32 refined by an IPV4 logical descriptor rather than losing the refinement, and dictionary encoding and layout hints are restored rather than re-derived. It is deliberately not portable and no foreign reader is promised, so Parquet remains the right choice for interchange; .skene is for cases where the draken-native round trip is what matters. See skene/FORMAT.md for the specification.

Parquet, CSV, and JSONL files can also be named directly with the read_parquet(), read_csv(), and read_jsonl() table functions. There is no read_skene() — skene datasets are read through a registered workspace like any other dataset.

Best With Opteryx Catalog

Opteryx Core works best when paired with the opteryx_catalog library. That is the intended model for named datasets, catalog-backed tables, and the general experience used in opteryx.app.

Typical setup:

import os

import opteryx

from opteryx import set_default_connector
from opteryx.connectors import OpteryxConnector
from opteryx_catalog import OpteryxCatalog

set_default_connector(
    OpteryxConnector,
    catalog=OpteryxCatalog,
    firestore_project=os.environ["GCP_PROJECT_ID"],
    firestore_database=os.environ["FIRESTORE_DATABASE"],
    gcs_bucket=os.environ["GCS_BUCKET"],
)

Once configured, you can query catalog-backed datasets using dot-separated names such as public.space.planets or opteryx.ops.billing.

For local data, Opteryx Core is typically used through registered workspaces such as testdata, scratch, or data. Queries refer to datasets by dot-separated names relative to the workspace root, for example testdata.planets, testdata.satellites, or scratch.signals.

Where It Fits

Opteryx Core is best thought of as an embedded analytical engine rather than a full end-user platform. If you want a hosted experience, multi-tenant service features, and the broader product workflow, use opteryx.app. If you want the core engine in your own environment, this package gives you that engine directly. If you want the intended table-resolution model, pair it with opteryx_catalog.

Contributing

If you use Opteryx-Core yourself, we want to hear from you.

  • Use it on your own datasets
  • Raise bugs when queries, schemas, or performance do not behave as expected
  • Open pull requests for fixes, tests, docs, or performance improvements
  • Share repro cases, failing queries, and edge-case Parquet files

This project is being actively built, and outside usage helps make it better.

Docs: https://docs.opteryx.app/ Source: https://github.com/mabel-dev/opteryx-core License: Apache-2.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

opteryx_core-0.9.101.tar.gz (13.3 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

opteryx_core-0.9.101-cp314-cp314-manylinux_2_34_x86_64.whl (40.6 MB view details)

Uploaded CPython 3.14manylinux: glibc 2.34+ x86-64

opteryx_core-0.9.101-cp314-cp314-manylinux_2_34_aarch64.whl (37.7 MB view details)

Uploaded CPython 3.14manylinux: glibc 2.34+ ARM64

opteryx_core-0.9.101-cp314-cp314-macosx_11_0_arm64.whl (26.3 MB view details)

Uploaded CPython 3.14macOS 11.0+ ARM64

opteryx_core-0.9.101-cp313-cp313-manylinux_2_34_x86_64.whl (40.6 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.34+ x86-64

opteryx_core-0.9.101-cp312-cp312-manylinux_2_34_x86_64.whl (40.7 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.34+ x86-64

opteryx_core-0.9.101-cp311-cp311-manylinux_2_34_x86_64.whl (41.2 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.34+ x86-64

File details

Details for the file opteryx_core-0.9.101.tar.gz.

File metadata

  • Download URL: opteryx_core-0.9.101.tar.gz
  • Upload date:
  • Size: 13.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for opteryx_core-0.9.101.tar.gz
Algorithm Hash digest
SHA256 dadcfa7b4663696672b9a5511198dbccb7e92b9d02f8807c5294ee06717597da
MD5 cb3397b6284b9a7f4c650771316b94cb
BLAKE2b-256 ffb7ef408127cef6cbe06a4d16e1c3addfcc840cbe60c5ae61b1e2a2c9f46398

See more details on using hashes here.

File details

Details for the file opteryx_core-0.9.101-cp314-cp314-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for opteryx_core-0.9.101-cp314-cp314-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 c65360e29cbe2168c8bf9b7e6ec9cbc4f9f61ba54d2dfce4909fe609537a5a5c
MD5 6e2d783f82811d53bece1035d72902d5
BLAKE2b-256 cc399deef506b840a6b5d705f44009e4131f0ff62e737b42dfd93619554b6819

See more details on using hashes here.

File details

Details for the file opteryx_core-0.9.101-cp314-cp314-manylinux_2_34_aarch64.whl.

File metadata

File hashes

Hashes for opteryx_core-0.9.101-cp314-cp314-manylinux_2_34_aarch64.whl
Algorithm Hash digest
SHA256 a24f789774671971165f1f6c4ed00c12d99dece832cf06ea4b266a450e2cafb0
MD5 48407c426e5633e50d735a2a950803f0
BLAKE2b-256 bbe61d65212fa2036c77e2fcbd203e903ff855b1dc7afab6c26d5ee7ddba40cf

See more details on using hashes here.

File details

Details for the file opteryx_core-0.9.101-cp314-cp314-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for opteryx_core-0.9.101-cp314-cp314-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 4bb3237acad284f113b529b428bbc9737f7ce5a6393a698c7779c4c7f07165ba
MD5 d6f625fa615ae5340ef39349955b5ecf
BLAKE2b-256 eb834c2f1804934680875f8e826cc26f1ff68f54067dc5dfd6c76f0488697600

See more details on using hashes here.

File details

Details for the file opteryx_core-0.9.101-cp313-cp313-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for opteryx_core-0.9.101-cp313-cp313-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 2df275e994fb70b0213b08ab566b79b1be2a138596ddc6e2ef624b484aa61588
MD5 05343f2a91d2b0ba557566135646dfd0
BLAKE2b-256 36717c45b96da284759a370029ec0f91ee3adfe8fbc0702795c8ea1395a97f74

See more details on using hashes here.

File details

Details for the file opteryx_core-0.9.101-cp312-cp312-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for opteryx_core-0.9.101-cp312-cp312-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 ddd2543b98cb44275cf03943cf206b3395aeb29df35a0d513a9ad3ee5883ea62
MD5 9e6a8742dcc36d5244aa0f70d15aa7a0
BLAKE2b-256 90a692a5d4b7b738fd1930c29cc106428f7d407633c2cbc0ca2826d07b7340e9

See more details on using hashes here.

File details

Details for the file opteryx_core-0.9.101-cp311-cp311-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for opteryx_core-0.9.101-cp311-cp311-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 d7464d0e478dbbb21bfc0bce8457da0bad11a95e16cdb59331c1826c575eae70
MD5 a99cbf28b0a18387d9e5bc9d77c48462
BLAKE2b-256 ce3fc987dcfb7b3698e6c068e8064e691cb851a5c3b33301cca0adfb8e15edf9

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.9.101 This release

7 files

0.9.100

7 files

0.9.99

7 files

0.9.97

7 files

0.9.96

7 files

0.9.95

7 files

0.9.94

7 files

0.9.93

7 files

0.9.92

7 files

0.9.91

7 files

0.9.90

7 files

0.9.89

7 files

0.9.88

7 files

0.9.87

7 files

0.9.86

7 files

0.9.85

7 files

0.9.84

7 files

0.9.83

7 files

0.9.82

7 files

0.9.81

7 files

0.9.80

7 files

0.9.79

7 files

0.9.78

7 files

0.9.77

7 files

0.9.76

7 files

0.9.75

7 files

0.9.74

7 files

0.9.73

7 files

0.9.71

10 files

0.9.70

10 files

0.9.69

10 files

0.9.68

6 files

0.9.67

6 files

0.9.66

6 files

0.9.65

6 files

0.9.64

6 files

0.9.63

6 files

0.9.62

6 files

0.9.60

6 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page