Skip to main content

A configuration-driven and programmatic ETL helper for DuckDB.

Project description

Quackpipe

The missing link between your Python scripts and your data infrastructure.

Quackpipe is a powerful ETL helper library that uses DuckDB to create a unified, high-performance data plane for Python applications. It bridges the gap between writing raw, complex connection code and adopting a full-scale data transformation framework.

With a simple YAML configuration, you can instantly connect to multiple data sources like PostgreSQL, S3, Azure Blob Storage, and SQLite, and even orchestrate complex DuckLake setups, all from a single, clean Python interface.

codecov

What Gap Does Quackpipe Fill?

In the modern data stack, you often face a choice:

  • Low-Level: Write boilerplate code with multiple database drivers (psycopg2, boto3, etc.) to connect and move data manually. This is flexible but repetitive and error-prone.
  • High-Level: Adopt a full DataOps framework like SQLMesh or dbt. These are powerful for building production-grade data warehouses but can be overkill for ad-hoc analysis, rapid prototyping, or simple scripting.

Quackpipe provides the perfect middle ground. It gives you the power of a unified query engine and the simplicity of a Python library, allowing you to:

  • Prototype Rapidly: Spin up a multi-source data environment in seconds.
  • Simplify ETL Scripts: Replace complex driver code with a single, clean session or a one-line move_data command.
  • Explore Data Interactively: Use the built-in CLI to launch a web UI with all your sources pre-connected for instant ad-hoc querying.
  • Bridge to Production: Automatically generate configuration for frameworks like SQLMesh when you're ready to graduate from a script to a versioned data model.

Core Capabilities

  • Unified Data Access: Query across PostgreSQL, S3, Azure, and SQLite as if they were all schemas in a single database.
  • Declarative Configuration: Define all your data sources in one human-readable config.yml file.
  • Powerful ETL Utilities: Move data between any two configured sources with the move_data() function.
  • Programmatic API: Use the QuackpipeBuilder for dynamic, on-the-fly connection setups in your code.
  • Secure Secret Management: Load credentials safely from .env files, keeping them out of your code and configuration.
  • Interactive UI: Launch an interactive DuckDB web UI with all your sources pre-connected using a single CLI command.
  • Framework Integration: Automatically generate a sqlmesh_config.yml file to seamlessly transition your project to a full DataOps framework.

Installation

pip install quackpipe

Install support for the sources you need:

# Example: Install support for Postgres, S3, Azure, and the UI
pip install "quackpipe[postgres,s3,azure,ui]"

Configuration

quackpipe uses a simple config.yml file to define your sources and an .env file to manage your secrets.

Configuration Priority

Quackpipe loads its configuration in the following order of priority:

  1. Directly in code: You can pass a config_path or a list of configs directly to functions like quackpipe.session() or move_data(). This always takes the highest priority.
  2. Environment Variable: If no configuration is provided in code, Quackpipe will check for the QUACKPIPE_CONFIG_PATH environment variable. You can set this to the path of your config.yml file.
  3. Local config.yml: When using the CLI, if a config.yml file exists in the current directory, it will be used automatically.

This layered approach provides flexibility for both interactive use and production deployments.

config.yml Example

# config.yml
sources:
  # A writeable PostgreSQL database.
  pg_warehouse:
    type: postgres
    secret_name: "pg_prod" # See Secret Management section below
    read_only: false       # Allows writing data back to this source

  # An S3 data lake for Parquet files.
  s3_datalake:
    type: s3
    secret_name: "aws_prod"
    region: "us-east-1"

  # An Azure Blob Storage container.
  azure_datalake:
    type: azure
    provider: connection_string
    secret_name: "azure_prod"

  # A composite DuckLake source.
  my_lake:
    type: ducklake
    catalog:
      type: sqlite
      path: "/path/to/lake_catalog.db"
    storage:
      type: local
      path: "/path/to/lake_storage/"

Secret Management with .env

Quackpipe uses a secret_name in the config to refer to a bundle of credentials. These are loaded from an .env file using a simple prefix convention: SECRET_NAME_KEY.

Create an .env file in your project root:

# .env

# Secrets for secret_name: "pg_prod"
PG_PROD_HOST=db.example.com
PG_PROD_USER=myuser
PG_PROD_PASSWORD=mypassword
PG_PROD_DATABASE=production

# Secrets for secret_name: "aws_prod"
AWS_PROD_ACCESS_KEY_ID=YOUR_AWS_ACCESS_KEY
AWS_PROD_SECRET_ACCESS_KEY=YOUR_AWS_SECRET_KEY

# Secrets for secret_name: "azure_prod"
AZURE_PROD_CONNECTION_STRING="DefaultEndpointsProtocol=https..."

Usage Highlights

1. Interactive Querying with session

Need to join a CSV in S3 with a table in Postgres? quackpipe makes it trivial.

import quackpipe

# quackpipe automatically loads your .env file
with quackpipe.session(config_path="config.yml", env_file=".env") as con:
    df = con.execute("""
        SELECT u.name, o.order_total
        FROM pg_warehouse.users u
        JOIN read_parquet('s3://my-bucket/orders/*.parquet') o ON u.id = o.user_id
        WHERE u.signup_date > '2024-01-01';
    """).fetchdf()

    print(df.head())

2. One-Line Data Movement with move_data

Archive old records from your production database to your data lake with a single command.

from quackpipe.etl_utils import move_data

move_data(
    config_path="config.yml",
    env_file=".env",
    source_query="SELECT * FROM pg_warehouse.logs WHERE timestamp < '2024-01-01'",
    destination_name="s3_datalake",
    table_name="logs_archive_2023"
)

3. Instant Data Exploration with the CLI

Launch a web browser UI with all your sources attached and ready for ad-hoc queries.

# This command reads your config.yml and .env file
quackpipe ui

# Or connect to specific sources
quackpipe ui pg_warehouse s3_datalake

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

quackpipe-0.6.6.tar.gz (43.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

quackpipe-0.6.6-py3-none-any.whl (38.6 kB view details)

Uploaded Python 3

File details

Details for the file quackpipe-0.6.6.tar.gz.

File metadata

  • Download URL: quackpipe-0.6.6.tar.gz
  • Upload date:
  • Size: 43.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.11

File hashes

Hashes for quackpipe-0.6.6.tar.gz
Algorithm Hash digest
SHA256 3a46a16feb6025981c82234a1153abb5d282c246efca8437f57cfe093c78c751
MD5 d15511e6ac9fd2516a3442cb17ffc5b6
BLAKE2b-256 d1dc81c31e26d61eda253347c84515798350e0cfa100afc63689b0a7a5f9ff23

See more details on using hashes here.

File details

Details for the file quackpipe-0.6.6-py3-none-any.whl.

File metadata

  • Download URL: quackpipe-0.6.6-py3-none-any.whl
  • Upload date:
  • Size: 38.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.11

File hashes

Hashes for quackpipe-0.6.6-py3-none-any.whl
Algorithm Hash digest
SHA256 ba8676c18f470add956e0472f781346804525317f42fe50f78d02d46866c7c05
MD5 0ac07190408b9789374e0d9254a1613f
BLAKE2b-256 bb670638feff3d09818b8a374d3b092f6975184ce4093ff734a1002005e5d68d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page