Skip to main content

Core ETL pipeline framework for mkpipe.

Project description

MkPipe

MkPipe is a modular, open-source ETL (Extract, Transform, Load) tool that allows you to integrate various data sources and sinks easily. It is designed to be extensible with a plugin-based architecture that supports extractors, transformers, and loaders.

Features

  • Extract data from multiple sources (e.g., PostgreSQL, MongoDB).
  • Transform data using custom Python logic and Apache Spark.
  • Load data into various sinks (e.g., ClickHouse, PostgreSQL, Parquet).
  • Plugin-based architecture that supports future extensions.
  • Cloud-native architecture, can be deployed on Kubernetes and other environments.

Quick Setup

You can deploy MkPipe using one of the following strategies:

1. Using Docker Compose

This method sets up all required services automatically using Docker Compose.

Steps:

  1. Clone or copy the deploy folder from the repository.

  2. Modify the configuration files:

  3. Run the following command to start the services:

    docker-compose up --build
    

    This will set up the following services:

    • PostgreSQL: Required for data storage.
    • RabbitMQ: Required for the Celery run_coordinator=celery.
    • Celery Worker: Required for running the Celery run_coordinator=celery.
    • Flower UI: Optional, but required for monitoring Celery tasks.

    Note: If you only want to use the run_coordinator=singlewithout Celery, only PostgreSQL is necessary.

2. Running Locally

You can also set up the environment manually and run MkPipe locally.

Steps:

  1. Set up and configure the following services:
    • RabbitMQ: Required for the Celery run_coordinator.
    • PostgreSQL: Required for data storage.
    • Flower UI: Optional, but required for monitoring Celery tasks.
  2. Update the following configuration files in the deploy folder:
    • .env for environment variables.
    • mkpipe_project.yaml for your ETL configurations.
  3. Install the python packages
    pip install mkpipe mkpipe-extractor-postgres mkpipe-loader-postgres
    
  4. Set the project directory environment variable:
    export MKPIPE_PROJECT_DIR={YOUR_PROJECT_PATH}
    
  5. Start MkPipe using the following command:
    mkpipe run
    

Documentation

For more detailed documentation, please visit the GitHub repository.

License

This project is licensed under the Apache 2.0 License - see the LICENSE file for details.

Db Support Plan

For actively supported databases/plugins, please visit the MkPipe-hub repository!

Core Relational Databases

  • PostgreSQL
  • MySQL
  • MariaDB
  • SQL Server
  • Oracle Database
  • SQLite
  • Snowflake
  • Google BigQuery
  • Amazon Redshift
  • ClickHouse
  • Amazon S3

NoSQL Databases

  • MongoDB
  • Cassandra
  • DynamoDB
  • Redis
  • Azure Data Lake Storage (ADLS)
  • Google Cloud Storage
  • Elasticsearch
  • TimescaleDB
  • HDFS
  • InfluxDB

ERP/CRM Systems

  • Salesforce
  • SAP
  • Microsoft Dynamics
  • NetSuite
  • Workday
  • HubSpot
  • Zoho CRM
  • Freshsales
  • Zendesk
  • Oracle NetSuite

Emerging Databases & Analytical Tools

Apache Druid

  • Vertica
  • SingleStore (MemSQL)
  • Exasol
  • SAP HANA
  • IBM Db2
  • Neo4j (Graph Database)
  • Greenplum
  • CockroachDB
  • AWS Athena

Streaming Systems

  • Kafka
  • RabbitMQ
  • Pulsar
  • Apache Flink
  • Amazon Kinesis
  • Google Pub/Sub
  • Azure Event Hubs
  • Apache NiFi
  • ActiveMQ
  • Redpanda

File Formats & Data Lakes

  • Parquet
  • Avro
  • JSON
  • CSV
  • XML
  • ORC
  • Google Drive (for raw files)
  • Dropbox
  • Box
  • FTP/SFTP Servers

Specialized Analytics Tools

  • Metabase (Data Visualization)
  • Tableau Data Extracts
  • Power BI
  • Looker
  • Google Analytics (GA4)
  • Mixpanel
  • Amplitude
  • Adobe Analytics
  • Heap
  • Klipfolio

Industry-Specific Databases

  • Aerospike
  • RocksDB
  • FaunaDB
  • ScyllaDB
  • ArangoDB
  • MarkLogic
  • CrateDB
  • TigerGraph
  • HarperDB
  • SAP ASE (Sybase)

Legacy Databases

  • Teradata
  • Netezza
  • Informix
  • Ingres
  • Firebird
  • Progress OpenEdge
  • ParAccel
  • MaxDB
  • HP Vertica
  • Sybase IQ

Emerging Cloud & Hybrid Databases

  • PlanetScale (MySQL-based)
  • YugabyteDB
  • TiDB
  • OceanBase
  • Citus (PostgreSQL-based)
  • Snowplow Analytics
  • Spanner (Google Cloud)
  • MariaDB ColumnStore
  • CockroachDB Serverless
  • Weaviate (Vector Search)

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mkpipe-0.4.5.tar.gz (26.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mkpipe-0.4.5-py3-none-any.whl (33.6 kB view details)

Uploaded Python 3

File details

Details for the file mkpipe-0.4.5.tar.gz.

File metadata

  • Download URL: mkpipe-0.4.5.tar.gz
  • Upload date:
  • Size: 26.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.13

File hashes

Hashes for mkpipe-0.4.5.tar.gz
Algorithm Hash digest
SHA256 975e07afd422117c2e9b16653c77689dffe6cb8576f9b610a554b42e9a0e4e6e
MD5 1ea882cc77f5717e5f8738c8067ede1f
BLAKE2b-256 4a6dc66876a6e4d9e8a6154dc01652bae9676610f7dd2dbea3c09b3d1eda5d06

See more details on using hashes here.

File details

Details for the file mkpipe-0.4.5-py3-none-any.whl.

File metadata

  • Download URL: mkpipe-0.4.5-py3-none-any.whl
  • Upload date:
  • Size: 33.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.13

File hashes

Hashes for mkpipe-0.4.5-py3-none-any.whl
Algorithm Hash digest
SHA256 33b467aed102dcbc6a02345241b7cdc0e44d32d1ec64fab735f41e189fba7de9
MD5 c2af75bda0eccc21ead464beaa199bfe
BLAKE2b-256 712f56a5a6c17ec249d2367524f9964b4d5182bd8802a67fe643766f635b8bf8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page