dpetl — Data package ETL
The dpetl is a command-line interface (CLI) tool designed to run the three ETL phases (Extract, Transform, Load).
It is designed to work alongside the Data Package standard specification.
Installation
It requires Python 3.10 or more. Install:
# using pip
pip install dpetl
# using poetry
poetry add dpetl
Usage
Activate your virtual environment!
Use the --help flag to inspect the CLI documentation:
dpetl --help
How It Works
The CLI loads Data Package descriptor(s) (via the frictionless-py Python package) and iterates over its resources.
By default, the CLI looks for:
-
datapackage.yamlwhen runningextractortransform -
datapackage.jsonwhen runningload
If you have multiple data packages, place them in a datapackages/ folder (each in its own subdirectory) and dpetl will process all of them.
You can also specify one or more descriptors manually using the -d or --descriptor flag (it can be passed multiple times). This is useful when you want to test a specific configuration or process a file that is not in the default location.
# Single descriptor
dpetl transform -d configs/datapackage_payroll.yaml
# Multiple descriptors
dpetl extract -d datapackages/sales/datapackage.yaml -d datapackages/hr/datapackage.yaml
Environment variables (EMAIL_USER, EMAIL_PWD, EMAIL_IMAP, GH_TOKEN, proxy variables) can also be defined in a .env file in the current working directory — it is loaded automatically.
# .env
EMAIL_USER=user@example.com
EMAIL_PWD=your-email-password-or-app-token
EMAIL_IMAP=imap.gmail.com
GH_TOKEN=ghp_xxxxxxxxxxxxxxxxxxxx
HTTPS_PROXY=http://user:password@proxy-host:8080
extract
Runs the ETL extraction phase. Downloads data from external sources (API, email) and saves them locally as configured in the datapackage.
# Run extract using the default datapackage.yaml descriptor
dpetl extract
# Extract emails received today only
dpetl extract --today-email
# Include package name in email subject search pattern
dpetl extract --add-package-name
Each resource in the descriptor must declare an extraction mode (email or api) inside its dpetl_extract property (see Example Data Package Configuration). If a resource is missing this property, the whole package extraction stops immediately.
Email Extractor
Connects to your IMAP server using environment variables (EMAIL_USER, EMAIL_PWD, EMAIL_IMAP), applying proxy settings (HTTP_PROXY/HTTPS_PROXY) if set.
It then searches for the most recent e-mail matching criteria:
-
subject: if you don't set it in the datapackage (criteria.subject), it defaults to the resource name (or{package_name}_{resource_name}with--add-package-name). -
date_gte: if you don't set it in the datapackage (criteria.date_gte), it defaults to the most recent e-mail (no date filter). Passing--today-emailsets it to today's date. -
any other filter supported by imap-tools (sender, folder, etc.) can also be set.
Once a matching e-mail is found, its first attachment is saved to resource.path. If there is more than one attachment, the extra ones are saved next to it, named {name}_1{ext}, {name}_2{ext}, and so on.
If the resource declares extrapaths (multiple files for the same resource), the same search-and-save logic runs once per path.
API Extractor
Reads the request settings from the resource's sources property:
-
method: required (e.g.get,post). -
path: required. Treated as a base URL, not the full file URL.dpetl appends the file name from
resource.pathto it (e.g.https://api.example.com+data/invoices.csv→https://api.example.com/invoices.csv). -
timeout: optional, in seconds (defaults to 30). -
params,headers,stream: optional, passed directly to the request.
Once the URL is built, dpetl makes the request and saves the response to resource.path.
If the resource declares extrapaths (multiple files for the same resource), the same steps run once per path, reusing the same base URL.
transform
Runs the ETL transformation phase. Applies column renaming, format conversion, and other transformations defined in the datapackage.
# Run transform using the default datapackage.yaml descriptor
dpetl transform
Reads the transformation settings from the resource's dpetl_transform property:
-
path: optional (defaults todata). Folder where the transformed file is saved. -
format: optional (defaults tocsv.gz, a gzip-compressed CSV).Supported formats:
csv,txt,xlsx— optionally followed by a compression, likecsv.gz. -
encoding: optional (defaults toutf-8). Used forcsv/txtfiles.
Any field in the resource schema may define a target property. If defined, the field is renamed to the specified target value.
Once all resources are transformed, dpetl updates the descriptor to match the generated files:
-
Each resource is converted to a single file, with updated
path,extrapaths(removed),scheme,formatandcompression(if any) values, and inferredstats. -
Field names are updated based on their target values, and the
targetproperty is removed. -
An
updated_attimestamp is added to the package, and the updated descriptor is saved as a JSON file.
load
Runs the ETL load phase. Uploads transformed data and the updated datapackage.json to a GitHub repository, creating a single commit with all files.
# Run load using the default datapackage.json descriptor
dpetl load
Requires the GH_TOKEN environment variable and reads its settings from the package's dpetl_load property:
-
repo: optional. Name of the target repository (defaults to the current repository). -
owner: required ifrepois set. GitHub user or organization name. -
level: optional (defaults touser). Useorgsto target a GitHub organization instead of a user account. -
visibility: optional (defaults toprivate). Usepublicfor a public repository.
If repo is set and the repository doesn't exist, dpetl creates it automatically.
Before publishing, the dpetl_extract, dpetl_transform and dpetl_load properties are removed from the descriptor.
The transformed data folder and the package descriptor, exported as datapackage.json, are published in a single commit.
Example Data Package Configuration
The following example shows a complete datapackage.yaml configuration.
Note that dpetl_extract and dpetl_transform are defined per resource, while dpetl_load is defined at the package level.
resources:
# Example 1: API extraction
- name: invoices
path: data/invoices.csv
sources: # Request settings
- method: get # required (e.g. get, post)
path: https://api.example.com # required
timeout: 30 # optional (Defaults to 30 seconds)
params: {} # optional
headers: {} # optional
stream: false # optional (Defaults to false)
dpetl_extract:
mode: api
# Example 2: Email extraction
- name: payroll_from_email
path: data/payroll.xlsx
dpetl_extract:
mode: email
mailbox: INBOX # optional (Defaults to INBOX)
criteria: # optional
subject: "Payroll Report" # optional (Defaults to resource name. See also the flag --add-package-name)
from_: "finance@example.com" # optional
date_gte: 2024-01-01 # optional (See also the flag --today-email)
# Example 3: Static resource with column renaming
- name: employees
path: data/employees.csv
schema:
fields:
- name: Name
type: string
target: employee_name
- name: Department
type: string
target: employee_department
dpetl_transform:
format: csv.gz # optional (Defaults to csv.gz)
path: data/processed # optional (Defaults to data)
encoding: utf-8 # optional (Defaults to utf-8)
# Load configuration (defined once per package)
dpetl_load:
owner: github-username
repo: my-data-repo
level: user # optional (Defaults to user)
visibility: private # optional (Defaults to private)
Validation
To validate datapackage resources and schemas, dpetl uses frictionless-py.
Validation can occur in two stages:
-
Before processing: validates the entire datapackage when
--validate-beforeis used. -
After processing: validates each processed resource.
Validation behavior depends on the command:
-
extract: validates each resource after download.--validate-beforeis not supported. -
transform: validates each resource after transformation. Use--validate-beforeto validate the entire datapackage before processing. -
load: validates before publishing only if--validate-beforeis used.
Use --no-validate to skip all validation.
Use --no-stop to continue processing resources even when validation errors are found.
Global Flags
Flags that can be used with any command:
| Flag | Description |
|---|---|
--descriptor, -d |
Path to one or more datapackage descriptors (repeatable) |
--no-validate, -nv |
Skip datapackage validation |
--no-stop, -ns |
Continue even if validation fails (do not exit with error) |
--validate-before, -vb |
Run validation before processing (not supported for extract) |
--version, -v |
Show version number and exit |
--help |
Show help message |
Environment Variables
| Variable | Used By | Description |
|---|---|---|
EMAIL_USER |
extract (email mode) | Username for IMAP email connection |
EMAIL_PWD |
extract (email mode) | Password for IMAP email connection |
EMAIL_IMAP |
extract (email mode) | IMAP server address (e.g., imap.gmail.com) |
HTTP_PROXY |
extract (email mode) | Proxy settings for IMAP connections* |
GH_TOKEN |
load | GitHub Personal Access Token |
All variables above can also be set in a .env file in the current directory instead of the shell environment.
* Proxy notes: If your network requires a proxy, dpetl supports both uppercase and lowercase proxy environment variables. Use the format http://user:pwd@host:port when authentication is required.
Design Philosophy
The dpetl package follows a convention over configuration philosophy, treating the Data Package descriptor as the single source of truth for ETL process.
Each resource declares how it should be processed through structured metadata, enabling reproducible, declarative, and version-controlled data workflows.
The goal is to keep the CLI simple while allowing flexible strategies driven entirely by configuration rather than imperative scripting.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dpetl-0.10.0.tar.gz.
File metadata
- Download URL: dpetl-0.10.0.tar.gz
- Upload date:
- Size: 16.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/2.2.1 CPython/3.10.12 Linux/7.0.11-76070011-generic
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
217841a21a7008d7316fdbd2deb031ec33d9734a57dbe758674843c145a42b98
|
|
| MD5 |
6b8a1c750491f2d0f448714a17bb23f5
|
|
| BLAKE2b-256 |
de305805ad7bd4c919722d0b6cadfd27325fef822e9ce7d7dee2ee6a9347272d
|
File details
Details for the file dpetl-0.10.0-py3-none-any.whl.
File metadata
- Download URL: dpetl-0.10.0-py3-none-any.whl
- Upload date:
- Size: 18.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/2.2.1 CPython/3.10.12 Linux/7.0.11-76070011-generic
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2407f7e72f0467e0e9a545e7e9048ce40604481d3608c8b6036f99781dbdd754
|
|
| MD5 |
4506032b3ecd54bfa920747a2019b23b
|
|
| BLAKE2b-256 |
f58d9090ead13c6939f04ea4b3f37a8b6879b3e52522592605469be642f69617
|