Skip to main content

Open Data Stack

CI Status PyPI Version Python Versions License

Architecture Diagram

A curated collection of open-source technologies and an accompanying CLI (odctl) for experimenting with modern data architecture and MLOps locally.

Provisioning a local data environment with distributed systems can be highly complex. The Open Data Stack streamlines this process by resolving dependency conflicts, network routing configurations, and integration challenges across tools like Kafka, Spark, Flink, Iceberg, and Airflow. It provides a cohesive, Docker-based blueprint that operates seamlessly out of the box.

Bundled Technologies

The stack is organized into distinct profiles that can be launched independently or together:

  • Messaging: Real-time event streaming, schema validation, and robust data ingestion.
    • Stack: Kafka (KRaft), Schema Registry (Karapace), Kafka Connect
  • Stream Processing: Stateful stream processing and continuous real-time data transformations.
    • Stack: Apache Flink
  • Data Processing: Distributed batch processing and large-scale ETL pipelines.
    • Stack: Apache Spark
  • Analytics: Real-time OLAP querying, federated SQL execution, and interactive BI dashboards.
    • Stack: ClickHouse, Trino, Metabase
  • Orchestration: Workflow scheduling, DAG execution, and complex pipeline automation.
    • Stack: Apache Airflow
  • MLOps: Machine learning experiment tracking, model registry, and scalable inference serving.
    • Stack: MLflow, Ray Serve
  • Metadata: Centralized data catalog, data discovery, and enterprise governance.
    • Stack: OpenMetadata
  • Observability: Metrics collection, intelligent alerting, and system telemetry visualization.
    • Stack: Prometheus, Alertmanager, Grafana
  • Lineage: Data provenance, pipeline dependency tracking, and troubleshooting.
    • Stack: OpenLineage, Marquez
  • Foundational Storage, Data Store, & Catalog: Persistent state, S3-compatible object storage, unified table metadata, high-performance caching/vector search, and unified stream storage.
    • Stack: PostgreSQL (pgvector), SeaweedFS (S3), Iceberg REST Catalog, Valkey Bundle, Apache Fluss

Prerequisites & Installation

Requirements

  • Docker: Docker Engine or Docker Desktop must be running. We highly recommend allocating at least 8GB to 16GB of RAM to Docker, as data processing engines are resource-heavy.
  • Python: Version 3.10 or higher.

Installation

Since odctl is a CLI tool, it is highly recommended to install it in an isolated environment using uv tool or pipx.

Using uv (Recommended):

uv tool install odctl

Using pipx:

pipx install odctl

Using pip:

pip install odctl

Quick Start

Get your local cluster up and running in three simple steps.

1. Initialize your workspace This command copies the default Docker Compose files and configurations into a hidden .odctl folder in your current directory.

odctl init

2. Explore available profiles See a full list of technologies you can launch.

odctl list

3. Launch the streaming and batch processing engines Bring up a robust data engineering environment.

odctl up kafka-lite flink-lite spark-lite

Note: You do not need to memorize dependencies. The CLI will automatically detect that these profiles require foundational infrastructure and will launch PostgreSQL, SeaweedFS (S3), and the Iceberg REST Catalog for you before starting the target compute engines.

CLI Command Reference

The odctl CLI orchestrates the Open Data Stack and is logically grouped by functionality. You can append --help to any command for deeper parameter details.

Global Options

  • --verbose: Enable debug-level logging across all commands.
  • -w, --workspace PATH: Path to the ODCTL workspace directory (default: ./.odctl).

Inspection & Info

  • odctl list: List all available profiles and their capabilities.
  • odctl explain <profile>: Explain the details, services, images, and dependencies of a profile.
  • odctl ps: List Docker containers managed by the Open Data Stack.
  • odctl info: View package and system-wide Docker daemon health status.

Workspace

  • odctl init: Initialize a local .odctl workspace for custom configurations.

Cluster Lifecycle

  • odctl pull: Pre-fetch Docker images without starting the containers.
  • odctl up: Launch Open Data profiles (automatically resolves upstream dependencies).
  • odctl down: Stop and remove profile containers and networks.

Management

  • odctl logs: Fetch the logs of containers managed by specific profiles.
  • odctl restart: Restart one or more specific profiles.

Examples

# View all profiles and exposed ports
$ odctl list -d

# See exactly what the kafka profile provisions
$ odctl explain kafka-lite

# Launch specific profiles
$ odctl up flink-lite kafka-lite spark-lite

# Complete teardown and wipe all data
$ odctl down --all --volumes

Workspace Customization (.odctl)

The Open Data Stack is designed to be fully hackable. When you run odctl init, a local ./.odctl/ workspace is generated in your current working directory.

This folder contains all the underlying configurations that power the stack:

  • compose-*.yml: The actual Docker Compose definitions. You can edit these to change exposed ports, adjust memory limits, or inject new environment variables.
  • registry.yml: The internal dependency graph.
  • .env: The environment variables used across the stack (e.g., default credentials or timezones).

The CLI will always prioritize the files in your local ./.odctl/ directory. If you make a mistake, you can always revert to the pristine default state by running odctl init --force.

Local Development & Contributing

If you want to contribute to the CLI itself, we welcome pull requests!

  1. Clone the repository.
  2. Install uv for dependency management.
  3. Sync the dependencies and install the project in development mode:
    uv sync
    
  4. Install the pre-commit hooks to ensure formatting checks pass:
    uv run pre-commit install
    
  5. Run the test suite:
    uv run pytest tests/
    

License

This project is licensed under the Apache License 2.0. See the LICENSE file for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

odctl-0.4.5.tar.gz (400.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

odctl-0.4.5-py3-none-any.whl (74.6 kB view details)

Uploaded Python 3

File details

Details for the file odctl-0.4.5.tar.gz.

File metadata

  • Download URL: odctl-0.4.5.tar.gz
  • Upload date:
  • Size: 400.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for odctl-0.4.5.tar.gz
Algorithm Hash digest
SHA256 8bb196bc5f93a925919891d67e4fc1332ab5fb1903597322cf86231719ce5748
MD5 6b4a97d2ab31f17ce6fe7f2582d55e6d
BLAKE2b-256 48b081216d028094d5d02ad8e2be892a1d21c0fe99f66646f2278ca56cbf0135

See more details on using hashes here.

Provenance

The following attestation bundles were made for odctl-0.4.5.tar.gz:

Publisher: publish.yml on jaehyeon-kim/odctl

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file odctl-0.4.5-py3-none-any.whl.

File metadata

  • Download URL: odctl-0.4.5-py3-none-any.whl
  • Upload date:
  • Size: 74.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for odctl-0.4.5-py3-none-any.whl
Algorithm Hash digest
SHA256 6ac3251a1c2326e3203cf1b6a3e8db4f270660936cccc326313418093a9bafe1
MD5 2ed7fd48396fe0971d0ce153a70c3565
BLAKE2b-256 c1d581760b5159e1733c410102350d0f3b683854b9b6138ec329afb8db2e2d02

See more details on using hashes here.

Provenance

The following attestation bundles were made for odctl-0.4.5-py3-none-any.whl:

Publisher: publish.yml on jaehyeon-kim/odctl

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.7

2 files

0.4.6

2 files

This release

0.4.5 This release

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page