Skip to main content

A short description of your project

Project description

Modular ETL Framework

About

This framework is designed for Extract - Transform - Load (ETL) operations, allowing you to process data from multiple source systems and load it into HDFS (Hadoop Distributed File System).

Project Tree

├── src

│ ├── config

│ ├── extract

│ ├── transform

│ ├── load

│ └── db_interaction

├── logs

├── tests

├── data

├── docs

├── output

├── .gitignore

├── config.ini

├── etl_db.db

├── main.py

├── requirements.txt

└── README.md

Descriptions

  • src: Main source code directory.
    • config: Configuration files for the ETL process.
    • extract: Code for extracting data from source systems.
    • transform: Code for transforming the extracted data.
    • load: Code for loading the transformed data into HDFS.
    • db_interaction: Code for interacting with databases if needed.
  • logs: Directory for storing log files generated during ETL process.
  • tests: Directory for unit tests and testing-related files.
  • data: Storage for input or intermediate data.
  • docs: Documentation related to the ETL framework.
  • output: Directory for final output files or data.
  • .gitignore: File specifying excluded files/directories from version control.
  • config.ini: Configuration file for connection details, API keys, etc.
  • etl_db.db: Local database file if used.
  • main.py: Main entry point for the ETL framework.
  • requirements.txt: List of required Python packages.
  • README.md: Main project documentation in Markdown.

Features

This ETL framework includes the following core features:

  • Extract: Fetch data from various source systems.
  • Transform: Clean, process, and transform the extracted data.
  • Load: Load the transformed data into HDFS.

Tech Stack

The framework utilizes the following technologies:

  • Programming Language: Python
  • ETL Libraries/Frameworks: (Specify if applicable)
  • Database System: (Specify if applicable)
  • HDFS Setup: (Specify if applicable)

Installation

Follow these steps to set up the ETL framework:

  1. Clone this repository: git clone <repository_url>
  2. Navigate to the project directory: cd ETL_Framework
  3. Install required packages: pip install -r requirements.txt
  4. Configure settings in config.ini as needed.

How to Use

Perform ETL operations using the following steps:

  1. Navigate to the project directory: cd ETL_Framework
  2. Run the ETL process: python main.py
  3. Monitor the logs in the logs directory for progress and issues.
  4. Find the final output in the output directory.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bigdata_vanhaydoi_framework-0.2.0.tar.gz (2.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

bigdata_vanhaydoi_framework-0.2.0-py3-none-any.whl (3.2 kB view details)

Uploaded Python 3

File details

Details for the file bigdata_vanhaydoi_framework-0.2.0.tar.gz.

File metadata

File hashes

Hashes for bigdata_vanhaydoi_framework-0.2.0.tar.gz
Algorithm Hash digest
SHA256 30d1b5bd341cf022a2b88ee5041e0adeebbe789a42382a4889a85591a2f28e34
MD5 8ec761b8728fca173a4c7b8a21f6e97f
BLAKE2b-256 ff33b43eb4056c2c3d171b0849716836b6dbf4e0ea7b3080cf64f3760075cff9

See more details on using hashes here.

File details

Details for the file bigdata_vanhaydoi_framework-0.2.0-py3-none-any.whl.

File metadata

File hashes

Hashes for bigdata_vanhaydoi_framework-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0f5f5b5c1098d12dbec49c6da076dd5a2900db71de4a19e11f7238f73de00de6
MD5 6e58b525991b3b6ce54206f64161fc96
BLAKE2b-256 cf13d7960bbe8325a9a6478012acad6cd08c905155e61945f9abcadaf6b3fb9f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page