Skip to main content

Open Data Stack

CI Status PyPI Version Python Versions License

Architecture Diagram

A curated collection of open-source technologies and an accompanying CLI (odctl) for experimenting with modern data architecture and MLOps locally.

Provisioning a local data environment with distributed systems can be highly complex. The Open Data Stack streamlines this process by resolving dependency conflicts, network routing configurations, and integration challenges across tools like Kafka, Spark, Flink, Iceberg, and Airflow. It provides a cohesive, Docker-based blueprint that operates seamlessly out of the box.

Bundled Technologies

The stack is organized into distinct profiles that can be launched independently or together:

  • Messaging: Real-time event streaming, schema validation, and robust data ingestion.
    • Stack: Kafka (KRaft), Schema Registry (Karapace), Kafka Connect, Kafka UI (kafbat)
  • Stream Processing: Stateful stream processing and continuous real-time data transformations.
    • Stack: Apache Flink
  • Data Processing: Distributed batch processing and large-scale ETL pipelines.
    • Stack: Apache Spark
  • Analytics: Real-time OLAP querying, federated SQL execution, and interactive BI dashboards.
    • Stack: ClickHouse, Trino, Metabase
  • Orchestration: Workflow scheduling, DAG execution, and complex pipeline automation, in one container that reads DAGs from object storage.
    • Stack: Apache Airflow (standalone, with the MLflow and Feast clients and the XGBoost, LightGBM and PyTorch runtimes)
  • MLOps: Machine learning experiment tracking, model registry, HTTP model serving, and a feature store.
    • Stack: MLflow, Feast
  • Metadata: Centralized data catalog, data discovery, and enterprise governance.
    • Stack: OpenMetadata
  • Observability: Metrics, traces and logs over OpenTelemetry, with dashboards, in one container.
    • Stack: grafana/otel-lgtm (OpenTelemetry Collector, Prometheus, Tempo, Loki, Grafana)
  • Lineage: Data provenance, pipeline dependency tracking, and troubleshooting.
    • Stack: OpenLineage, Marquez
  • Foundational Storage, Data Store, & Catalog: Persistent state, S3-compatible object storage, unified table metadata, high-performance caching/vector search, and unified stream storage.
    • Stack: PostgreSQL (pgvector), SeaweedFS (S3), Iceberg REST Catalog, Valkey Bundle, Apache Fluss

Prerequisites & Installation

Requirements

  • Docker: Docker Engine or Docker Desktop must be running. We highly recommend allocating at least 8GB to 16GB of RAM to Docker, as data processing engines are resource-heavy.
  • Python: Version 3.10 or higher.

Installation

Since odctl is a CLI tool, it is highly recommended to install it in an isolated environment using uv tool or pipx.

Using uv (Recommended):

uv tool install odctl

Using pipx:

pipx install odctl

Using pip:

pip install odctl

Quick Start

Get your local cluster up and running in three simple steps.

1. Initialize your workspace This command copies the default Docker Compose files and configurations into a hidden .odctl folder in your current directory.

odctl init

2. Explore available profiles See a full list of technologies you can launch.

odctl list

3. Launch the streaming and batch processing engines Bring up a robust data engineering environment.

odctl up kafka-lite flink-lite spark-lite

Note: You do not need to memorize dependencies. The CLI will automatically detect that these profiles require foundational infrastructure and will launch PostgreSQL, SeaweedFS (S3), and the Iceberg REST Catalog for you before starting the target compute engines.

CLI Command Reference

The odctl CLI orchestrates the Open Data Stack and is logically grouped by functionality. You can append --help to any command for deeper parameter details.

Global Options

  • --verbose: Enable debug-level logging across all commands.
  • -w, --workspace PATH: Path to the ODCTL workspace directory (default: ./.odctl).

Inspection & Info

  • odctl list: List all available profiles and their capabilities.
  • odctl explain <profile>: Explain the details, services, images, and dependencies of a profile.
  • odctl ps: List Docker containers managed by the Open Data Stack.
  • odctl info: View package and system-wide Docker daemon health status.

Workspace

  • odctl init: Initialize a local .odctl workspace for custom configurations.

Cluster Lifecycle

  • odctl pull: Pre-fetch Docker images without starting the containers.
  • odctl up: Launch Open Data profiles (automatically resolves upstream dependencies).
  • odctl down: Stop and remove profile containers and networks.

Management

  • odctl logs: Fetch the logs of containers managed by specific profiles.
  • odctl restart: Restart one or more specific profiles, keeping their containers.
  • odctl recreate: Replace one or more profiles' containers, which is what applies an edited compose file. Add --pull to refresh the images too.

restart and recreate differ in one way that matters: a restart bounces the process inside the existing container, so a new memory limit, port, image tag or environment variable never reaches it, while a recreate replaces the container from its current definition. Recreating therefore discards whatever that container held. Both act only on the profiles you name, unlike up, which stops the ones you leave out.

Examples

# View all profiles and exposed ports
$ odctl list -d

# See exactly what the kafka profile provisions
$ odctl explain kafka-lite

# Launch specific profiles
$ odctl up flink-lite kafka-lite spark-lite

# Complete teardown and wipe all data
$ odctl down --all --volumes

Workspace Customization (.odctl)

The Open Data Stack is designed to be fully hackable. When you run odctl init, a local ./.odctl/ workspace is generated in your current working directory.

This folder contains all the underlying configurations that power the stack:

  • compose-*.yml: The actual Docker Compose definitions. You can edit these to change exposed ports, adjust memory limits, or inject new environment variables.
  • registry.yml: The internal dependency graph.
  • .env: The environment variables used across the stack (e.g., default credentials or timezones).
  • trino/rules.json: Trino's file-based access control. The shipped policy is deliberately generic: it grants every identity full table privileges and denies analyst schema ownership, and it names no catalog, schema or table, because those belong to your project rather than to this tool. Add your own table rules here for row filtering and column masking, and restart Trino afterwards, since the file is read at startup.

The CLI will always prioritize the files in your local ./.odctl/ directory. If you make a mistake, you can always revert to the pristine default state by running odctl init --force.

Blog posts that use this CLI:

Local Development & Contributing

If you want to contribute to the CLI itself, we welcome pull requests!

  1. Clone the repository.
  2. Install uv for dependency management.
  3. Sync the dependencies and install the project in development mode:
    uv sync
    
  4. Install the pre-commit hooks to ensure formatting checks pass:
    uv run pre-commit install
    
  5. Run the test suite:
    uv run pytest tests/
    

License

This project is licensed under the Apache License 2.0. See the LICENSE file for details.

Release files for odctl 0.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for odctl 0.8.0
File Size Uploaded
odctl-0.8.0.tar.gz 312.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for odctl 0.8.0
File Interpreter ABI Platform
odctl-0.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 402.1 kB

Release files / odctl-0.8.0.tar.gz

Download URL odctl-0.8.0.tar.gz
Size 312.4 kB
Tags Source
SHA-256 checksum
How to use checksums
6b52c9b48748063434ba8c761859853a10bd035efe34076e1c6b44d0e06eef25
BLAKE2b-256 checksum
How to use checksums
fe5d9d30cd621c47bea47e66a811eb9ecd476a59130457d67583f83a98c5152a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.

Transparency log

Release files / odctl-0.8.0-py3-none-any.whl

Download URL odctl-0.8.0-py3-none-any.whl
Size 89.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dda1129bf05ef414a6b3c055d849eb58cf19a6459aeb83f7ee975bcaca465ed9
BLAKE2b-256 checksum
How to use checksums
ef576ff4b0f03637d27d9ffb8ff78542ef9e1b40dbd7069393b67f0885c3605f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.8.0 This release

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page