Skip to main content

A configuration-driven and programmatic ETL helper for DuckDB.

Project description

Quackpipe

The missing link between your Python scripts and your data infrastructure.

Quackpipe is a powerful ETL helper library that uses DuckDB to create a unified, high-performance data plane for Python applications. It bridges the gap between writing raw, complex connection code and adopting a full-scale data transformation framework.

With a simple YAML configuration, you can instantly connect to multiple data sources like PostgreSQL, S3, Azure Blob Storage, and SQLite, and even orchestrate complex DuckLake setups, all from a single, clean Python interface.

codecov

What Gap Does Quackpipe Fill?

In the modern data stack, you often face a choice:

  • Low-Level: Write boilerplate code with multiple database drivers (psycopg2, boto3, etc.) to connect and move data manually. This is flexible but repetitive and error-prone.
  • High-Level: Adopt a full DataOps framework like SQLMesh or dbt. These are powerful for building production-grade data warehouses but can be overkill for ad-hoc analysis, rapid prototyping, or simple scripting.

Quackpipe provides the perfect middle ground. It gives you the power of a unified query engine and the simplicity of a Python library, allowing you to:

  • Prototype Rapidly: Spin up a multi-source data environment in seconds.
  • Simplify ETL Scripts: Replace complex driver code with a single, clean session or a one-line move_data command.
  • Explore Data Interactively: Use the built-in CLI to launch a web UI with all your sources pre-connected for instant ad-hoc querying.
  • Bridge to Production: Automatically generate configuration for frameworks like SQLMesh when you're ready to graduate from a script to a versioned data model.

Core Capabilities

  • Unified Data Access: Query across PostgreSQL, S3, Azure, and SQLite as if they were all schemas in a single database.
  • Declarative Configuration: Define all your data sources in one human-readable config.yml file.
  • Powerful ETL Utilities: Move data between any two configured sources with the move_data() function.
  • Programmatic API: Use the QuackpipeBuilder for dynamic, on-the-fly connection setups in your code.
  • Secure Secret Management: Load credentials safely from .env files, keeping them out of your code and configuration.
  • Interactive UI: Launch an interactive DuckDB web UI with all your sources pre-connected using a single CLI command.
  • Framework Integration: Automatically generate a sqlmesh_config.yml file to seamlessly transition your project to a full DataOps framework.

Installation

pip install quackpipe

Install support for the sources you need:

# Example: Install support for Postgres, S3, Azure, and the UI
pip install "quackpipe[postgres,s3,azure,ui]"

Configuration

quackpipe uses a simple config.yml file to define your sources and an .env file to manage your secrets.

Configuration Priority

Quackpipe loads its configuration in the following order of priority:

  1. Directly in code: You can pass a config_path or a list of configs directly to functions like quackpipe.session() or move_data(). This always takes the highest priority.
  2. Environment Variable: If no configuration is provided in code, Quackpipe will check for the QUACKPIPE_CONFIG_PATH environment variable. You can set this to the path of your config.yml file.
  3. Local config.yml: When using the CLI, if a config.yml file exists in the current directory, it will be used automatically.

This layered approach provides flexibility for both interactive use and production deployments.

config.yml Example

# config.yml
sources:
  # A writeable PostgreSQL database.
  pg_warehouse:
    type: postgres
    secret_name: "pg_prod" # See Secret Management section below
    read_only: false       # Allows writing data back to this source

  # An S3 data lake for Parquet files.
  s3_datalake:
    type: s3
    secret_name: "aws_prod"
    region: "us-east-1"

  # An Azure Blob Storage container.
  azure_datalake:
    type: azure
    provider: connection_string
    secret_name: "azure_prod"

  # A composite DuckLake source.
  my_lake:
    type: ducklake
    catalog:
      type: sqlite
      path: "/path/to/lake_catalog.db"
    storage:
      type: local
      path: "/path/to/lake_storage/"

Secret Management with .env

Quackpipe uses a secret_name in the config to refer to a bundle of credentials. These are loaded from an .env file using a simple prefix convention: SECRET_NAME_KEY.

Create an .env file in your project root:

# .env

# Secrets for secret_name: "pg_prod"
PG_PROD_HOST=db.example.com
PG_PROD_USER=myuser
PG_PROD_PASSWORD=mypassword
PG_PROD_DATABASE=production

# Secrets for secret_name: "aws_prod"
AWS_PROD_ACCESS_KEY_ID=YOUR_AWS_ACCESS_KEY
AWS_PROD_SECRET_ACCESS_KEY=YOUR_AWS_SECRET_KEY

# Secrets for secret_name: "azure_prod"
AZURE_PROD_CONNECTION_STRING="DefaultEndpointsProtocol=https..."

Usage Highlights

1. Interactive Querying with session

Need to join a CSV in S3 with a table in Postgres? quackpipe makes it trivial.

import quackpipe

# quackpipe automatically loads your .env file
with quackpipe.session(config_path="config.yml", env_file=".env") as con:
    df = con.execute("""
        SELECT u.name, o.order_total
        FROM pg_warehouse.users u
        JOIN read_parquet('s3://my-bucket/orders/*.parquet') o ON u.id = o.user_id
        WHERE u.signup_date > '2024-01-01';
    """).fetchdf()

    print(df.head())

2. One-Line Data Movement with move_data

Archive old records from your production database to your data lake with a single command.

from quackpipe.etl_utils import move_data

move_data(
    config_path="config.yml",
    env_file=".env",
    source_query="SELECT * FROM pg_warehouse.logs WHERE timestamp < '2024-01-01'",
    destination_name="s3_datalake",
    table_name="logs_archive_2023"
)

3. Instant Data Exploration with the CLI

Launch a web browser UI with all your sources attached and ready for ad-hoc queries.

# This command reads your config.yml and .env file
quackpipe ui

# Or connect to specific sources
quackpipe ui pg_warehouse s3_datalake

4. Validate Your Configuration

Before running your scripts, you can validate your config.yml file against the built-in schema to catch errors early.

# Validate the default config.yml
quackpipe validate

# Or validate a specific file
quackpipe validate --config /path/to/your/config.yml

If the configuration is valid, you'll see a success message:

✅ Configuration file at 'config.yml' is valid.

If it's invalid, quackpipe will tell you why:

❌ Configuration file at 'config.yml' is invalid.
   Reason: 'port' in source 'pg_main' should be an integer.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

quackpipe-0.6.8.tar.gz (49.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

quackpipe-0.6.8-py3-none-any.whl (42.7 kB view details)

Uploaded Python 3

File details

Details for the file quackpipe-0.6.8.tar.gz.

File metadata

  • Download URL: quackpipe-0.6.8.tar.gz
  • Upload date:
  • Size: 49.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.11

File hashes

Hashes for quackpipe-0.6.8.tar.gz
Algorithm Hash digest
SHA256 11e024eb20509e6b856b8ce5e2089360f5f616b88301dcfc2bc5f163dfd47e0f
MD5 0c241c095ab9a1e813989c1b4c5da91c
BLAKE2b-256 37d3eec9ee886c9ddc068e30879dad3e65c7690e3d4937c799bd634377a3b265

See more details on using hashes here.

File details

Details for the file quackpipe-0.6.8-py3-none-any.whl.

File metadata

  • Download URL: quackpipe-0.6.8-py3-none-any.whl
  • Upload date:
  • Size: 42.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.11

File hashes

Hashes for quackpipe-0.6.8-py3-none-any.whl
Algorithm Hash digest
SHA256 c8ed81b1faa521e3c7c4bf74cd32eddc08bac5344d640dae0dcb15cee8f50065
MD5 b5b253feb8cf2fbc1b91693c4b64a273
BLAKE2b-256 166d56482e1190380196c65940a0d974c29e8d1954ba998e1c640d5a0177b4ec

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page