Skip to main content

A short description of your project

Project description

Modular ETL Framework

About

This framework is designed for Extract - Transform - Load (ETL) operations, allowing you to process data from multiple source systems and load it into HDFS (Hadoop Distributed File System).

Project Tree

├── src

│ ├── config

│ ├── extract

│ ├── transform

│ ├── load

│ └── db_interaction

├── logs

├── tests

├── data

├── docs

├── output

├── .gitignore

├── config.ini

├── etl_db.db

├── main.py

├── requirements.txt

└── README.md

Descriptions

  • src: Main source code directory.
    • config: Configuration files for the ETL process.
    • extract: Code for extracting data from source systems.
    • transform: Code for transforming the extracted data.
    • load: Code for loading the transformed data into HDFS.
    • db_interaction: Code for interacting with databases if needed.
  • logs: Directory for storing log files generated during ETL process.
  • tests: Directory for unit tests and testing-related files.
  • data: Storage for input or intermediate data.
  • docs: Documentation related to the ETL framework.
  • output: Directory for final output files or data.
  • .gitignore: File specifying excluded files/directories from version control.
  • config.ini: Configuration file for connection details, API keys, etc.
  • etl_db.db: Local database file if used.
  • main.py: Main entry point for the ETL framework.
  • requirements.txt: List of required Python packages.
  • README.md: Main project documentation in Markdown.

Features

This ETL framework includes the following core features:

  • Extract: Fetch data from various source systems.
  • Transform: Clean, process, and transform the extracted data.
  • Load: Load the transformed data into HDFS.

Tech Stack

The framework utilizes the following technologies:

  • Programming Language: Python
  • ETL Libraries/Frameworks: (Specify if applicable)
  • Database System: (Specify if applicable)
  • HDFS Setup: (Specify if applicable)

Installation

Follow these steps to set up the ETL framework:

  1. Clone this repository: git clone <repository_url>
  2. Navigate to the project directory: cd ETL_Framework
  3. Install required packages: pip install -r requirements.txt
  4. Configure settings in config.ini as needed.

How to Use

Perform ETL operations using the following steps:

  1. Navigate to the project directory: cd ETL_Framework
  2. Run the ETL process: python main.py
  3. Monitor the logs in the logs directory for progress and issues.
  4. Find the final output in the output directory.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bigdata_vanhaydoi_framework-0.1.0.tar.gz (2.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

bigdata_vanhaydoi_framework-0.1.0-py3-none-any.whl (3.1 kB view details)

Uploaded Python 3

File details

Details for the file bigdata_vanhaydoi_framework-0.1.0.tar.gz.

File metadata

File hashes

Hashes for bigdata_vanhaydoi_framework-0.1.0.tar.gz
Algorithm Hash digest
SHA256 432f63df42a5fc26aa896890cbc01eb394203e9c80a9e2591961340c9398724b
MD5 e26ebc0a5cfa57ab11f41f72652208d7
BLAKE2b-256 1249b536b9e2930e7c861dfda836639a3267a477b6b2983a4be56fc56a4d6d40

See more details on using hashes here.

File details

Details for the file bigdata_vanhaydoi_framework-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for bigdata_vanhaydoi_framework-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a96148d258a1e8b2a883c14a4a47e2931c5831fdec84b969fef51662eccd1ab1
MD5 ad83511728b3375b80d938df840d8565
BLAKE2b-256 02f77f57d04781fcebe076912ec321f38165b2ebb5452ada2f42a5ce68ff5319

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page