Skip to main content

Hydra ETL

Your data pipelines are described, not programmed.

CI PyPI Python License: AGPL v3 Status: beta

Website · Docs & playground · MCP server · VS Code extension · Changelog

Hydra is an open-source declarative ETL engine. You write what a pipeline is in YAML; Hydra validates it before touching any data, then runs it. The same manifests run from the terminal, a REST API, or a visual editor in your browser, with one pip install and nothing else to deploy.

Hydra Studio: visual editor for jobs and workflows

Quick start

pip install "hydra-etl[server]"
hdrctl serve --open            # Studio + API on http://localhost:5678

Or stay in the terminal:

hdrctl init my_job             # scaffold one of six templates
hdrctl validate my_job         # strict validation, no data touched
hdrctl run my_job              # execute

No Node, no build step, no database, no message broker. Python 3.9+ on Linux, macOS or Windows.


Why Hydra

  • Validate before you run. hdrctl validate checks sources, steps, types and destinations deterministically, before any data is read or written.
  • Declarative, visual, zero deployment. YAML manifests, a browser editor that outputs the same YAML, and a single Python process. No cluster, no JVM, no container required.
  • Built for CI/CD. Manifests are versioned and reviewed like code; hdrctl scaffolds, validates, tests and runs from a terminal or a CI job.
  • Built-in scheduler. DAG workflows with dependencies, parallel branches, delayed retries, cron triggers, conditional guards and eleven actions (webhook, email, Bash, PowerShell, SSH, Python…).
  • One job, many environments. {{ param: }}, {{ env: }} and ${SECRET:} keep the manifest identical across dev, staging and production.
  • Two engines per operation. pandas or DuckDB, chosen step by step, with optional Rust acceleration (see Native acceleration).
  • AI-ready. An MCP server lets Claude, Cursor or VS Code write pipelines that Hydra validates.

Validation before execution Declarative, visual, zero deployment

How it compares

Hydra Airflow dbt Airbyte NiFi / Apache Hop
Pipelines defined in YAML Python SQL + YAML UI / config UI flows
Scope Extract, transform, load Orchestration Transform in the warehouse Extract & load Extract, transform, load
Visual editor Included Monitoring UI Not in dbt Core Included Included
To get started pip install Scheduler, webserver, metadata DB A data warehouse Docker / Kubernetes JVM

Hydra is not a replacement for all of these. It targets file-and-database pipelines that should stay readable and run without extra infrastructure.


What you get

Component Description
Engine Declarative jobs: one source, N transformations, one destination
CLI hdrctl: scaffold, validate, run, inspect. English and Spanish
API FastAPI, with interactive docs at /docs
Studio Visual editor for jobs and workflows, served by the same process
Workflows Multi-job DAG with dependencies, actions, retries and runtime parameters
MCP Natural-language pipeline authoring, validated by Hydra

Connectors (sources and destinations):


CSV

JSON

Parquet

MySQL / MariaDB

PostgreSQL

MongoDB

Web API

Transformation engines: pandas and DuckDB.


Who it is for

  • Data engineers and analysts who need repeatable file-and-database pipelines without standing up infrastructure for them.
  • Teams where pipelines must stay readable by people who do not write Python: a YAML manifest reviewed in a pull request, not a script.
  • Developers embedding ETL in a product, who want a CLI and a REST API over the same engine.

A job in four files

A job is a folder. Four manifests describe it, and each one answers a single question.

sources.yaml: where the data comes from

version: "1.0"
sources:
  src_input:
    type: csv
    extract:
      table: ./input.csv

transformations.yaml: how it is reshaped

version: "1.0"
steps:
  - cast:
      mapping:
        amount: float
  - filter:
      expr: "amount > 0"
  - aggregate:
      by: [name]
      agg:
        total: { func: sum, col: amount }

A CSV carries no types, so cast comes before any numeric comparison.

destinations.yaml: where it goes

version: "1.0"
destinations:
  dest_output:
    type: csv
    load:
      table: ./output.csv
      mode: replace          # append | replace | upsert

pipeline.yaml: which source feeds which destination

version: "1.0"
pipeline:
  from: src_input
  to: dest_output
hdrctl validate my_job && hdrctl run my_job

Workflows: order several jobs

version: "1.0"
workflow:
  name: daily_etl
  trigger:
    type: schedule
    cron: "0 8 * * *"
  steps:
    - name: extract
      type: job
      job: ./jobs/extract
      depends_on: []

    - name: transform
      type: job
      job: ./jobs/transform
      depends_on: ["extract"]     # always a list, supports fan-in

    - name: notify
      type: action
      action: webhook
      params: { url: "{{ env:WEBHOOK_URL }}" }
      depends_on: ["transform"]
      on_failure: skip
hdrctl workflow validate ./workflow.yaml
hdrctl workflow run      ./workflow.yaml

An edge is a dependency, not a pipe: it decides when a job runs, never what data reaches it. Steps that share no dependency run in parallel.

DAG workflows with a built-in scheduler

Parameters

Values can be declared once and reused, or created while the workflow runs.

- filter:
    expr: "region == '{{ param:region }}'"

{{ param:NAME }} reads a parameter, {{ env:NAME }} an environment variable, ${SECRET:NAME} a secret. The set_param and assign_param actions create and change parameters mid-run, so two jobs can share a placeholder and produce different results.


Install what you need

The base install is the engine and the CLI. Everything else is opt-in.

pip install hydra-etl                  # engine + CLI
pip install "hydra-etl[server]"        # + API + Studio
pip install "hydra-etl[postgres]"      # + PostgreSQL driver
pip install "hydra-etl[all]"           # everything

Available extras: server, native, duckdb, parquet, mysql, postgres, mongodb, http, mcp, all.


Serving

hdrctl serve                 # Studio and API on port 5678
hdrctl serve --open          # and open the browser
hdrctl serve --no-studio     # API only, for a headless server
hdrctl serve --port 8080

The server writes projects into the directory you launch it from. Interactive API docs are at /docs.


Your AI assistant, connected (MCP)

Hydra ships an MCP server. Point Claude Desktop, Cursor, VS Code or any MCP client at it, and ask for a pipeline in plain language.

pip install "hydra-etl[mcp]"
hydra-mcp

Your assistant writes the manifests; Hydra validates them before anything is written, and nothing runs until you ask. Requires Python 3.10+. See docs/MCP.md for client configuration.


Native acceleration (optional)

Parts of the engine have a Rust implementation. It is optional and off by default: without it, Hydra behaves exactly the same.

pip install "hydra-etl[native]"
HYDRA_BACKEND=rust hdrctl run ./jobs/sales

Or per operation, with a hydra.backends.yaml file next to the job:

default: python
overrides:
  csv.read: rust      # currently the only accelerated operation

On one million rows, CSV reading is about four times faster, with byte-for-byte identical results (parity checked on 26,000 CSV files and one million floats against CPython repr()). Hydra falls back to Python automatically whenever that guarantee cannot be kept. Benchmarks: bench/.


VS Code extension

Completion and validation for Hydra manifests inside the editor, without installing Hydra. Download the .vsix from the latest release, or build it yourself from vscode-extension/ (python build_vsix.py, no Node required).

Hydra VS Code extension

Project status

Beta. The documented connectors and steps are implemented and covered by tests. The shape of the YAML DSL is settled; any breaking change will be announced in the release notes before 1.0. Use it on real work, and pin your version.


Contributing

Issues, ideas and pull requests are welcome. See CONTRIBUTING.md. If Hydra is useful to you, a ⭐ on GitHub helps others find it.


License

Hydra ETL is released under the GNU Affero General Public License v3 or later. See LICENSE.

You may use, modify and redistribute it freely, including commercially. If you modify Hydra and let others use it, even only over a network, you must make your modified source available under the same terms.

Commercial license: for use without that obligation, contact the author via hydraetl.com or by opening an issue.

Release files for hydra-etl 0.10.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hydra-etl 0.10.1
File Size Uploaded
hydra_etl-0.10.1.tar.gz 3.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for hydra-etl 0.10.1
File Interpreter ABI Platform
hydra_etl-0.10.1-py3-none-any.whl Python 3 none any Details

Total release size: 6.3 MB

Release files / hydra_etl-0.10.1.tar.gz

Download URL hydra_etl-0.10.1.tar.gz
Size 3.1 MB
Tags Source
SHA-256 checksum
How to use checksums
438f9febe820147d87c8680890e278330c841232f1262eea333e80df62225186
BLAKE2b-256 checksum
How to use checksums
2762e48cd77a9fea6b44f514a4eb23a6db2ebe27241bdf2f44bdd8b084e82b05
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / hydra_etl-0.10.1-py3-none-any.whl

Download URL hydra_etl-0.10.1-py3-none-any.whl
Size 3.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
ebfaba6b15bbc8e2d423e864e3ed89fedd58aae28a7685ebf899ed0047efacdf
BLAKE2b-256 checksum
How to use checksums
aa28211b649389c373e5e2532733d9ba4a8afc33a7133a633973198229050351
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.10.1 This release

2 release files

0.10.0

2 release files

0.9.6

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page