Skip to main content

No project description provided

Project description

duckberg


DuckBerg
query your iceberg data easily and efficiently

Hatch project PyPI - Version PyPI - Python Version linting - Ruff code style - Black License: Apache 2.0


Table of Contents

About

Duckberg is a Python package that synthesizes the power of PyIceberg and DuckDb. PyIceberg enables efficient interaction with Apache Iceberg, a format for handling large datasets, while DuckDb offers swift in-memory data analysis. When combined, these tools create Duckberg, which simplifies the querying process for large Iceberg datasets stored on blob storage with a user-friendly Pythonic approach.

The underlying principle of the Duckberg Python package is to execute your SQL queries only on those data lake files that contain the necessary data for your results. To fully utilize the benefits of this package, it's assumed that your data is partitioned in a manner that suits your query and use case.

Iceberg catalog types

Duckberg supports the same Iceberg catalogs as PyIceberg, including REST, SQL, Hive, Glue, and DynamoDB. These catalogs are sources of information about Iceberg datasets, tables, partitions, etc. Before using Duckberg, ensure that you have access to an Iceberg catalog that can be utilized.

Installation

pip install duckberg

Features

Easy initialisation

Following initialisation is using the REST Iceberg catalog with Amazon S3 as a iceberg data storage.

from duckberg import DuckBerg

catalog_config: dict[str, str] = {
  "type": "rest", # Iceberg catalog type 
  "uri": "http://iceberg-rest:8181/", # url for Iceberg catalog
  "credentials": "user:password", # credentials for Iceberg catalog
  "s3.endpoint": S3_ENDPOINT, # s3 
  "s3.access-key-id": S3_ACCESS_KEY_ID,
  "s3.secret-access-key": S3_SECET_KEY
}

db = DuckBerg(
     catalog_name="warehouse",
     catalog_config=catalog_config)

Listing tables

db.list_tables()

Listing partitions for particular table

db.list_partitions(table="nyc.taxis")

Querying data to Pandas dataframe

query = "SELECT * FROM nyc.taxis WHERE trip_distance > 40 ORDER BY tolls_amount DESC"
df = db.select(table="nyc.taxis", partition_filter="payment_type = 1", sql=query).read_pandas()

Playground

You can run the playground environment running docker compose in the playground

cd playground
docker-compose up -d

The initial run could take additional time for jupyter docker image build. Then you can access

Iceberg data init

Once all the containers have been initiated run the Spark Iceberg Jupyter notebook that will init the Iceberg data and catalog.

Duckberg playground

Navigate to localhost:8888. Then select example Jupyter notebook you want to run and enjoy Duckberg!

Development

For the development, there is recommendation to use Python 3.10. If you manage your Python versions by Pyenv use

pyenv install 3.10.13
pyenv global 3.10.13

then create and activate virtual environment

python -m venv venv
source venv/bin/activate 

upgrade pip and install dependencies

pip install --upgrade pip
pip install .

then run dockers that contains Iceberg catalog and file storage containing iceberg files

cd playground
docker-compose up -d

init data by running Init Jupyter notebook and run/test Duckberg in the file tests/duckberg-sample.py

Style & Formatting

Use

hatch run lint:fmt
hatch run lint:style

Building package

The Duckberg project is managed by Hatch. Follow [Hatch docs] for an installation or just install by command

brew install hatch

or

pip install hatch

Increase package by

hatch version "x.x.x"

Build

hatch build

and publish

hatch publish

License

duckberg is distributed under the terms of the Apache 2.0 license.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

duckberg-0.0.3.tar.gz (534.3 kB view details)

Uploaded Source

Built Distribution

duckberg-0.0.3-py3-none-any.whl (8.7 kB view details)

Uploaded Python 3

File details

Details for the file duckberg-0.0.3.tar.gz.

File metadata

  • Download URL: duckberg-0.0.3.tar.gz
  • Upload date:
  • Size: 534.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: python-httpx/0.25.2

File hashes

Hashes for duckberg-0.0.3.tar.gz
Algorithm Hash digest
SHA256 2a3cc150a690ccd534c9e1e8ae7c0d82133f2c1f2cab2f62d5c9d4ee2be4b7c5
MD5 29a413c022c76cf018e1a8d30defbcd4
BLAKE2b-256 b0f67b210514f465fd8ca4c5d8746c357a333db9cd960700b6ec2668748eb5b0

See more details on using hashes here.

File details

Details for the file duckberg-0.0.3-py3-none-any.whl.

File metadata

  • Download URL: duckberg-0.0.3-py3-none-any.whl
  • Upload date:
  • Size: 8.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: python-httpx/0.25.2

File hashes

Hashes for duckberg-0.0.3-py3-none-any.whl
Algorithm Hash digest
SHA256 4e7ae8ed13e10b19d22f6d6ee94ddf25aaa7b4119b0cac3e167bf3639a44dace
MD5 20c11c7e07f8ab82980c9e69f98bfe0d
BLAKE2b-256 c5ed5c703db856ac4e2edd4ae926863b73adb57e75cb6dd63e29e248b1f6562f

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page