Skip to main content

Ubunye Engine

Ubunye (oo-BOON-yeh) — isiZulu for "unity"

One framework. Every pipeline. Any environment.

DocsQuickstartWhy UbunyeCommunity


Hey there 👋

A data pipeline is a program that moves data from one place to another — a database to a file, a REST API to a data warehouse — and usually reshapes the data along the way. Building one from scratch is mostly plumbing: wire up the connection, juggle credentials, learn a framework's quirks, write the same "read → transform → write" scaffold for the tenth time this year. It's a lot of glue code standing between you and the three lines that actually matter.

Ubunye Engine writes that plumbing for you. You describe the pipeline in a short YAML file and put your transformation in a normal Python class. Ubunye takes care of connections, the compute engine (Apache Spark), and the read/write loop.

Same pipeline runs on your laptop today and on a production cluster tomorrow, with no code changes. That sentence is tested, not hoped: a build job runs one pipeline on six environments and fails if the outputs differ by a single byte.


What the engine gives you today

  • Proven portability. The same task runs on a laptop, in Docker, on Kubernetes, against object storage, through the cloud submit path, and on Databricks. One output hash across all of them, checked on every change.
  • Models saved anywhere. The model registry writes to a local folder, a Databricks volume, S3 or GCS, chosen purely by the path. New storage kinds are one class and one entry point.
  • A truly open plugin system. Connectors, storage backends and registries are all added from the outside, with no engine edits and no inheritance required. A test proves it with a connector the engine has never seen.
  • Types that ship. The package carries its type information, and a type checker guards every merge.
  • Errors that help. Failures say what went wrong, show the context, and suggest the fix.
  • Nothing here is decorative. Eleven worked examples have been run for real. Claims that could not be executed were removed from these docs rather than left to mislead.

Quickstart

Install it:

pip install ubunye-engine

Scaffold a new pipeline folder:

ubunye init -d ./pipelines -u demo -p starter -t filter_adults

You get:

pipelines/demo/starter/filter_adults/
  config.yaml              ← describes the pipeline (inputs, outputs, settings)
  transformations.py       ← your code goes here
  notebooks/               ← an interactive dev notebook for exploring

ubunye init gives you a working starting point you can customise. For a minimal run-it-on-your-laptop example, edit config.yaml to read a local CSV and write Parquet:

CONFIG:
  inputs:
    people:
      format: s3              # generic file reader; "file://" paths work too
      file_format: csv
      path: "file:///tmp/people.csv"
      options:
        header: "true"
        inferSchema: "true"

  outputs:
    adults:
      format: s3
      file_format: parquet
      path: "file:///tmp/adults/"
      mode: overwrite

Then open transformations.py and write your logic:

from typing import Any, Dict
from ubunye.core.interfaces import Task


class FilterAdults(Task):
    """Keep only rows where age is 18 or older."""

    def transform(self, sources: Dict[str, Any]) -> Dict[str, Any]:
        people = sources["people"]
        return {"adults": people.filter("age >= 18")}

Two things to notice:

  • sources["people"] matches the inputs.people name from the YAML.
  • The return key "adults" matches the outputs.adults name.

Run it:

ubunye run -d ./pipelines -u demo -p starter -t filter_adults

That's the whole loop. Ubunye reads /tmp/people.csv, hands you a Spark DataFrame, and writes whatever you return to /tmp/adults/.

Running on Databricks? Call it from a notebook instead:

import ubunye
outputs = ubunye.run_task(task_dir="./pipelines/demo/starter/filter_adults")

Ubunye detects Databricks' active Spark session and reuses it — same pipeline, no code change.

Want realistic end-to-end examples? They live in their own repository: ubunye-examples. Every one of them has been run for real, and most run on several environments.


Why Ubunye

We've all been there. You join a new team, open the repo, and find five Spark projects — each structured differently, each with its own way of handling configs, credentials, and deployment. One uses a JSON file, another has everything hardcoded, a third has a 300-line bash script that "Dave wrote and it just works."

Ubunye says: let's agree on how pipelines look. One folder structure. One config format. One CLI. Whether you're building an ETL job, a feature pipeline, or an ML training run.

Without Ubunye With Ubunye
Every project looks different One standard: use_case / pipeline / task
Spark setup scattered everywhere Engine handles it from YAML config
Credentials hardcoded or inconsistent {{ env.DB_PASSWORD }} everywhere
"Works on my machine" Same config runs local, YARN, K8s, Databricks
New teammate needs a week to onboard ubunye init and they're running in minutes

How It Works

Three simple ideas:

Config over code. Your pipeline is a YAML file. Inputs, outputs, Spark settings, scheduling — all declared, not coded.

Plugins for everything. The format field in your config picks which connector to use. A connector is a small Python class that knows how to read from or write to one specific place (a database, a REST API, a cloud bucket). Built-ins include hive, jdbc, delta, s3, unity, and rest_api. Need a new data source? Write one and register it — Ubunye discovers plugins automatically.

Folders as architecture. Pipelines are organized as project / use_case / pipeline / task. The CLI uses this structure for scaffolding, execution, and discovery:

pipelines/
  fraud_detection/
    ingestion/
      claim_etl/
      policy_etl/
    feature_engineering/
      claim_features/
    risk_scoring/
      train_model/
      score_claims/

What Can You Build With It

ETL pipelines — move data between Hive, JDBC databases, Delta Lake, S3, REST APIs. Config-driven, scheduled, reproducible.

ML training and inference — define your model behind a simple contract, let the engine handle versioning, storage, and deployment.

RAG document pipelines — ingest documents, extract text, chunk, compute embeddings, load into a vector store. All from YAML.

Feature engineering — compute features once, write to a shared table, reuse across use cases.

Data drift detection — monitor feature distributions between runs, flag when things shift.

Check out the Patterns section in our docs for full examples.


Examples

All worked examples live in ubunye-examples. They install the engine from PyPI, so what you run there is what you get from pip install.

They are grouped by where they run. The same task folder is used in every environment. Only environment variables change.

Databricks (free workspace is enough)

Example What it shows
Tables and SQL pushdown read a table, push a join down as SQL, write with merge
REST API ingestion the rest_api connector against a public weather API
Unstructured files binary file reading and text chunking
The ML lifecycle features, train, quality gate, registry, promote, score
RAG embeddings and a chat model, retrieval, grounded answers
Fine-tune an open LLM a large model labels data, DistilBERT learns from it
Data quality a contract with severities, bad rows quarantined
Model monitoring drift, decay, and a rollback that is a decision, not a reflex

Local machine, Docker, and Kubernetes (no cloud account needed)

Example What it shows
Run anywhere one task, identical output hash on local Spark, Docker, Kubernetes, object storage and Databricks
JDBC a partitioned parallel read from a real public PostgreSQL database
RAG and fine-tuning on open models the same pipelines with local open source models instead of hosted endpoints
The three-framework race the same model in scikit-learn, PyTorch and TensorFlow behind identical configs

AWS and GCP

Submit scripts and CI jobs exist for EMR Serverless and Dataproc Serverless. They are written and documented but need a cloud account to run, and the CI jobs say plainly when they were skipped rather than run.

Connectors

Format Read Write Description
hive Apache Hive tables
jdbc PostgreSQL, MySQL, Teradata, and more
delta Delta Lake (standalone or Unity Catalog)
s3 S3, HDFS, or local filesystem
unity Databricks Unity Catalog
binary Binary files (images, PDFs)
rest_api REST APIs with pagination and auth

Want to add one? See the plugin guide.


Run Anywhere

The same task folder runs on every environment below. This is tested, not claimed: a CI job runs one task on each and fails the build if the output hashes differ.

Environment How you launch it
Your laptop ubunye run ... with local Spark
Docker one container image, same entry point
Kubernetes the same image as a Job
Databricks ubunye.run_task() from a notebook or a bundle job
AWS EMR Serverless, GCP Dataproc spark-submit with python -m ubunye

One rule makes this work: the task never chooses its own cluster. Which Spark master to use belongs to whoever launches the job, so leave spark.master out of your config. On a laptop the runner sets local[*]. On a cloud the platform sets it, and the engine refuses a config that tries to override it, because a silent single-node run on paid compute is worse than an error.

---------------------------------------|-----------------------------------------------| | Your laptop | spark.master: "local[*]" | | Hadoop / YARN cluster | spark.master: "yarn" | | Kubernetes | spark.master: "k8s://..." | | Databricks notebooks or jobs | Call ubunye.run_task() from Python — Ubunye picks up the active session | | AWS EMR | Runs as an EMR Step |

Don't recognise some of these? That's fine — you only need one. If you're starting out, local[*] runs Spark on your own machine with no setup.


Jinja Templating

Anywhere a string appears in your YAML, you can plug in a variable using {{ … }} syntax (this is called Jinja templating). That's how you keep secrets out of your config, change paths per environment, and inject the run date from the CLI:

# Environment variables
password: "{{ env.DB_PASSWORD }}"

# CLI variables (--var ds=2025-01-01)
path: "s3a://bucket/{{ ds }}/"

# Defaults
path: "s3a://bucket/{{ ds | default('2025-01-01') }}/"

CLI

ubunye init     -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # scaffold
ubunye validate -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # check config
ubunye plan     -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # preview plan
ubunye run      -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # execute
ubunye test run -d ./pipelines -u <use_case> -p <pipeline> -t <task>   # test mode
ubunye lineage list -d ./pipelines -u <use_case> -p <pipeline> -t <task>  # run history
ubunye models list -u <use_case> -m <model> -s <store>                 # model versions

Python API

import ubunye

# Run from Databricks or any Python environment
outputs = ubunye.run_task(task_dir="./pipelines/...", mode="DEV", dt="2024-06-01")

# Multiple tasks
results = ubunye.run_pipeline(
    usecase_dir="./pipelines", usecase="fraud", package="etl",
    tasks=["claim_etl", "features"], mode="DEV",
)

What Ubunye Is Not

It's not an agent framework — use LangChain or CrewAI for that. It's not an orchestrator — use Airflow, Prefect, or Dagster. It's not a compute engine — it runs on Spark.

Ubunye is the standardization layer between your data sources and your applications. It makes the plumbing boring so you can focus on what matters.


Roadmap

  • Config-driven ETL pipelines
  • Multi-environment profiles
  • Jinja templating
  • Plugin-based connectors
  • CLI scaffolding and execution
  • Pydantic config validation
  • ML model contract
  • Model registry with versioning
  • Lineage tracking
  • Python API for Databricks
  • Databricks Asset Bundles deployment
  • Dev notebook scaffolding
  • Data drift detection
  • ubunye deploy CLI command

Get Involved

We'd love your help. Whether it's a new connector, a bug fix, a typo, or just telling us what you're building — all contributions matter.


License

MIT License


Built with 🇿🇦 by Ubunye AI Ecosystems

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ubunye_engine-0.5.0.tar.gz (143.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ubunye_engine-0.5.0-py3-none-any.whl (176.7 kB view details)

Uploaded Python 3

File details

Details for the file ubunye_engine-0.5.0.tar.gz.

File metadata

  • Download URL: ubunye_engine-0.5.0.tar.gz
  • Upload date:
  • Size: 143.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for ubunye_engine-0.5.0.tar.gz
Algorithm Hash digest
SHA256 59543934ada646a9d7b266fa56a94496cef38b1140d783573f394a41d156e524
MD5 a9df914f77eb88ff643eea4c3c865337
BLAKE2b-256 8d22c722bcc380628892d74604a178479cf44230d6b0ba2b182768c4f7699ba4

See more details on using hashes here.

Provenance

The following attestation bundles were made for ubunye_engine-0.5.0.tar.gz:

Publisher: publish_pypip.yml on ubunye-ai-ecosystems/ubunye_engine

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ubunye_engine-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: ubunye_engine-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 176.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for ubunye_engine-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 575c0087e15b51078f58347a3e60404b87d9101fe303b95713745d224ac4f3c6
MD5 63cf9582e6a09016d1abb4f1992533f9
BLAKE2b-256 b79c78d0a29b4e8f56f8ae1d3b5d3ce23223aec79ef8143d0c34af261fcc0870

See more details on using hashes here.

Provenance

The following attestation bundles were made for ubunye_engine-0.5.0-py3-none-any.whl:

Publisher: publish_pypip.yml on ubunye-ai-ecosystems/ubunye_engine

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page