Skip to main content

Python package for fast and parallel transferring a bulk of files to S3, Azure Blob Storage, and Google Cloud Storage

Project description


Logo

Cloud Bulk Upload (cloudbulkupload)

Python package for fast and parallel transferring a bulk of files to S3, Azure Blob Storage, and Google Cloud Storage!
See on PyPI · View Examples · Report Bug/Request Feature

Python Version License Downloads

Table of Contents
  1. About cloudbulkupload
  2. Getting Started
  3. Usage by Provider
  4. Testing and Performance
  5. Documentation
  6. Contributing
  7. Contributors
  8. License

About cloudbulkupload

Boto3 is the official Python SDK for accessing and managing all AWS resources such as Amazon Simple Storage Service (S3). Generally, it's pretty ok to transfer a small number of files using Boto3. However, transferring a large number of small files impede performance. Although it only takes a few milliseconds per file to transfer, it can take up to hours to transfer hundreds of thousands, or millions, of files if you do it sequentially. Moreover, because Amazon S3 does not have folders/directories, managing the hierarchy of directories and files manually can be a bit tedious especially if there are many files located in different folders.

The cloudbulkupload package solves these issues. It speeds up transferring of many small files to Amazon AWS S3, Azure Blob Storage, and Google Cloud Storage by executing multiple download/upload operations in parallel by leveraging the Python multiprocessing module and async/await patterns. Depending on the number of cores of your machine, Cloud Bulk Upload can make cloud storage transfers even 100X faster than sequential mode using traditional Boto3! Furthermore, Cloud Bulk Upload can keep the original folder structure of files and directories when transferring them.

🚀 Main Functionalities

  • 🔄 Multi-Cloud Support: AWS S3, Azure Blob Storage, and Google Cloud Storage
  • ⚡ High Performance: Multi-thread and async operations for maximum speed
  • 📁 Directory Operations: Upload/download entire directories with structure preservation
  • 🎯 Bulk Operations: Efficient handling of thousands of files
  • 📊 Progress Tracking: Built-in progress bars for long-running operations
  • 🧪 Comprehensive Testing: Full test suite with performance comparisons
  • 🔧 Configurable: Customizable concurrency, timeouts, and error handling
  • 📈 Performance Monitoring: Built-in metrics and comparison tools

🏆 Performance Benefits

  • 100X faster than sequential uploads
  • Async operations for Azure and Google Cloud
  • Multi-threading for AWS S3
  • Configurable concurrency for optimal performance
  • Memory efficient for large file sets

Getting Started

Prerequisites

Note: You can deploy a free S3-compatible server using MinIO on your local machine for testing. See our documentation for setup instructions.

Installation

Use the package manager pip to install cloudbulkupload.

pip install cloudbulkupload

For development and testing:

pip install "cloudbulkupload[test]"

Quick Start

# AWS S3
from cloudbulkupload import BulkBoto3

aws_client = BulkBoto3(
    endpoint_url="your-endpoint",
    aws_access_key_id="your-key",
    aws_secret_access_key="your-secret",
    verbose=True
)

# Upload directory
aws_client.upload_dir_to_storage(
    bucket_name="my-bucket",
    local_dir="path/to/files",
    storage_dir="uploads",
    n_threads=50
)
# Azure Blob Storage
import asyncio
from cloudbulkupload import BulkAzureBlob

async def azure_example():
    azure_client = BulkAzureBlob(
        connection_string="your-connection-string",
        verbose=True
    )
    
    await azure_client.upload_directory(
        container_name="my-container",
        local_dir="path/to/files",
        storage_dir="uploads"
    )

asyncio.run(azure_example())
# Google Cloud Storage
import asyncio
from cloudbulkupload import BulkGoogleStorage

async def google_example():
    google_client = BulkGoogleStorage(
        project_id="your-project-id",
        verbose=True
    )
    
    await google_client.upload_directory(
        bucket_name="my-bucket",
        local_dir="path/to/files",
        storage_dir="uploads"
    )

asyncio.run(google_example())

Usage by Provider

AWS S3

AWS S3 support uses multi-threading for optimal performance on the AWS platform.

Basic Setup

from cloudbulkupload import BulkBoto3

client = BulkBoto3(
    endpoint_url="https://s3.amazonaws.com",  # or your custom endpoint
    aws_access_key_id="your-access-key",
    aws_secret_access_key="your-secret-key",
    max_pool_connections=300,
    verbose=True
)

Directory Operations

# Upload entire directory
client.upload_dir_to_storage(
    bucket_name="my-bucket",
    local_dir="path/to/local/directory",
    storage_dir="uploads/my-files",
    n_threads=50
)

# Download entire directory
client.download_dir_from_storage(
    bucket_name="my-bucket",
    storage_dir="uploads/my-files",
    local_dir="downloads",
    n_threads=50
)

Individual File Operations

from cloudbulkupload import StorageTransferPath

# Upload specific files
upload_paths = [
    StorageTransferPath("file1.txt", "uploads/file1.txt"),
    StorageTransferPath("file2.txt", "uploads/file2.txt")
]

client.upload(bucket_name="my-bucket", upload_paths=upload_paths)

# Download specific files
download_paths = [
    StorageTransferPath("uploads/file1.txt", "local/file1.txt"),
    StorageTransferPath("uploads/file2.txt", "local/file2.txt")
]

client.download(bucket_name="my-bucket", download_paths=download_paths)

Bucket Management

# Create bucket
client.create_new_bucket("new-bucket-name")

# List objects
objects = client.list_objects(bucket_name="my-bucket", storage_dir="uploads")

# Check if object exists
exists = client.check_object_exists(bucket_name="my-bucket", object_path="uploads/file.txt")

# Empty bucket
client.empty_bucket("my-bucket")

Azure Blob Storage

Azure Blob Storage support uses async/await patterns for optimal performance.

Basic Setup

import asyncio
from cloudbulkupload import BulkAzureBlob

async def main():
    client = BulkAzureBlob(
        connection_string="your-azure-connection-string",
        max_concurrent_operations=50,
        verbose=True
    )
    
    # Your operations here
    await client.upload_directory(
        container_name="my-container",
        local_dir="path/to/files",
        storage_dir="uploads"
    )

asyncio.run(main())

Directory Operations

# Upload directory
await client.upload_directory(
    container_name="my-container",
    local_dir="path/to/local/directory",
    storage_dir="uploads/my-files"
)

# Download directory
await client.download_directory(
    container_name="my-container",
    storage_dir="uploads/my-files",
    local_dir="downloads"
)

Individual File Operations

from cloudbulkupload import StorageTransferPath

# Upload specific files
upload_paths = [
    StorageTransferPath("file1.txt", "uploads/file1.txt"),
    StorageTransferPath("file2.txt", "uploads/file2.txt")
]

await client.upload_files("my-container", upload_paths)

# Download specific files
download_paths = [
    StorageTransferPath("uploads/file1.txt", "local/file1.txt"),
    StorageTransferPath("uploads/file2.txt", "local/file2.txt")
]

await client.download_files("my-container", download_paths)

Container Management

# Create container
await client.create_container("new-container")

# List blobs
blobs = await client.list_blobs("my-container", prefix="uploads/")

# Check if blob exists
exists = await client.check_blob_exists("my-container", "uploads/file.txt")

# Empty container
await client.empty_container("my-container")

Convenience Functions

from cloudbulkupload import bulk_upload_blobs, bulk_download_blobs

# Bulk upload
files = ["file1.txt", "file2.txt", "file3.txt"]
await bulk_upload_blobs(
    connection_string="your-connection-string",
    container_name="my-container",
    files_to_upload=files,
    max_concurrent=50,
    verbose=True
)

# Bulk download
await bulk_download_blobs(
    connection_string="your-connection-string",
    container_name="my-container",
    files_to_download=files,
    local_dir="downloads",
    max_concurrent=50,
    verbose=True
)

Google Cloud Storage

Google Cloud Storage support uses async/await patterns and includes a hybrid approach with Google's Transfer Manager for maximum performance.

Basic Setup

import asyncio
from cloudbulkupload import BulkGoogleStorage

async def main():
    client = BulkGoogleStorage(
        project_id="your-project-id",
        credentials_path="/path/to/service-account.json",  # Optional
        max_concurrent_operations=50,
        verbose=True
    )
    
    # Your operations here
    await client.upload_directory(
        bucket_name="my-bucket",
        local_dir="path/to/files",
        storage_dir="uploads"
    )

asyncio.run(main())

Authentication Options

# Method 1: Service Account Key File
client = BulkGoogleStorage(
    project_id="your-project-id",
    credentials_path="/path/to/service-account.json"
)

# Method 2: Service Account JSON String (for cloud/container environments)
client = BulkGoogleStorage(
    project_id="your-project-id",
    credentials_json='{"type": "service_account", ...}'
)

# Method 3: Application Default Credentials
client = BulkGoogleStorage(project_id="your-project-id")

Directory Operations

# Upload directory
await client.upload_directory(
    bucket_name="my-bucket",
    local_dir="path/to/local/directory",
    storage_dir="uploads/my-files"
)

# Download directory
await client.download_directory(
    bucket_name="my-bucket",
    storage_dir="uploads/my-files",
    local_dir="downloads"
)

Individual File Operations

from cloudbulkupload import StorageTransferPath

# Upload specific files
upload_paths = [
    StorageTransferPath("file1.txt", "uploads/file1.txt"),
    StorageTransferPath("file2.txt", "uploads/file2.txt")
]

await client.upload_files("my-bucket", upload_paths)

# Download specific files
download_paths = [
    StorageTransferPath("uploads/file1.txt", "local/file1.txt"),
    StorageTransferPath("uploads/file2.txt", "local/file2.txt")
]

await client.download_files("my-bucket", download_paths)

Hybrid Approach: Standard vs Transfer Manager

# Standard Mode (Consistent API across all providers)
await client.upload_files("my-bucket", upload_paths)

# Transfer Manager Mode (High Performance - Google Cloud only)
await client.upload_files("my-bucket", upload_paths, use_transfer_manager=True)

Bucket Management

# Create bucket
await client.create_bucket("new-bucket-name")

# List blobs
blobs = await client.list_blobs("my-bucket", prefix="uploads/")

# Check if blob exists
exists = await client.check_blob_exists("my-bucket", "uploads/file.txt")

# Empty bucket
await client.empty_bucket("my-bucket")

Convenience Functions

from cloudbulkupload import google_bulk_upload_blobs, google_bulk_download_blobs

# Bulk upload
files = ["file1.txt", "file2.txt", "file3.txt"]
await google_bulk_upload_blobs(
    project_id="your-project-id",
    bucket_name="my-bucket",
    files_to_upload=files,
    max_concurrent=50,
    verbose=True,
    use_transfer_manager=True  # Optional: Use Google's Transfer Manager
)

# Bulk download
await google_bulk_download_blobs(
    project_id="your-project-id",
    bucket_name="my-bucket",
    files_to_download=files,
    local_dir="downloads",
    max_concurrent=50,
    verbose=True
)

Testing and Performance

Running Tests

The package includes a comprehensive test suite for all providers and performance comparisons.

Install Test Dependencies

pip install "cloudbulkupload[test]"

Run Different Test Types

# Unit tests
python run_tests.py --type unit

# Performance tests
python run_tests.py --type performance

# AWS S3 tests
python run_tests.py --type aws

# Azure Blob Storage tests
python run_tests.py --type azure

# Google Cloud Storage tests
python run_tests.py --type google-cloud

# AWS vs Azure comparison
python run_tests.py --type azure-comparison

# Three-way comparison (AWS, Azure, Google)
python run_tests.py --type three-way-comparison

# All tests
python run_tests.py --type all

Individual Test Files

# Run specific test files
python tests/aws_s3_test.py
python tests/azure_blob_test.py
python tests/google_cloud_test.py
python tests/performance_comparison_three_way.py

Performance Comparison

The package includes built-in performance comparison tools to test and compare different cloud providers.

Three-Way Performance Comparison

python tests/performance_comparison_three_way.py

This will:

  • Test AWS S3, Azure Blob Storage, and Google Cloud Storage
  • Compare upload/download speeds
  • Generate performance reports
  • Create CSV files with detailed metrics

Performance Metrics

The tests measure:

  • Upload Speed: MB/s for different file sizes
  • Download Speed: MB/s for different file sizes
  • Concurrency Impact: Performance with different thread counts
  • File Size Impact: Performance with different file sizes
  • Provider Comparison: Direct comparison between AWS, Azure, and Google Cloud

Expected Performance

Based on our testing:

  • AWS S3: 5-8 MB/s with multi-threading
  • Azure Blob Storage: 6-9 MB/s with async operations
  • Google Cloud Storage: 6-9 MB/s with async operations
  • Google Transfer Manager: 8-12 MB/s for large files

Test Results

Test results are automatically generated and saved to:

  • performance_comparison_results.csv - AWS vs Azure comparison
  • performance_comparison_three_way_results.csv - Three-way comparison
  • test_results.csv - General test results
  • google_cloud_test_results.json - Google Cloud specific results

For detailed test documentation, see docs/TESTING.md.

Documentation

Comprehensive documentation is available in the docs/ directory:

📚 Implementation Guides

📋 Implementation Summaries

🧪 Testing Documentation

📦 PyPI Publishing

📖 Original Documentation

Contributing

Any contributions you make are greatly appreciated. If you have a suggestion that would make this better, please fork the repo and create a pull request. You can also simply open an issue with the tag "enhancement". To contribute to cloudbulkupload, follow these steps:

  1. Fork this repository
  2. Create a feature branch (git checkout -b feature/AmazingFeature)
  3. Make your changes and commit them (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a pull request

Alternatively, see the GitHub documentation on creating a pull request.

Development Setup

# Clone the repository
git clone https://github.com/dynamicdeploy/cloudbulkupload.git
cd cloudbulkupload

# Create virtual environment
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

# Install development dependencies
pip install -e ".[test,dev]"

# Run tests
python run_tests.py --type all

Contributors

Thanks to the following people who have contributed to this project:

License

Distributed under the MIT License. See LICENSE for more information.


Credits

This project is based on the original work by Amir Masoud Sefidian who created the bulk upload concept and initial implementation. The original repository can be found at: https://github.com/iamirmasoud/bulkboto3

The project has been significantly expanded to support multiple cloud providers (AWS S3, Azure Blob Storage, and Google Cloud Storage) while maintaining the core performance benefits of the original implementation.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cloudbulkupload-2.0.0.tar.gz (27.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cloudbulkupload-2.0.0-py3-none-any.whl (18.7 kB view details)

Uploaded Python 3

File details

Details for the file cloudbulkupload-2.0.0.tar.gz.

File metadata

  • Download URL: cloudbulkupload-2.0.0.tar.gz
  • Upload date:
  • Size: 27.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.3

File hashes

Hashes for cloudbulkupload-2.0.0.tar.gz
Algorithm Hash digest
SHA256 883d1eb4a5a0ae2cf155b7c26cf13bb69efa5729b057a41118fe180784b13c1b
MD5 2d983d004c38e85914991801eff3dcd3
BLAKE2b-256 74b86636c565eecb2b64ac46671e3478d1f4cdcaec92cbf506e4ea31c13b9f94

See more details on using hashes here.

File details

Details for the file cloudbulkupload-2.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for cloudbulkupload-2.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e2c4ec51767bca0342f11c7d2c47801ce2125e90b573f5b699b7f74fc4eaacea
MD5 a702ae0b6f5a6144f46f9c1641d3f1ee
BLAKE2b-256 a2accc78d9cc60b6bf0516c4e61cc9bfadf82d34c6e04f192f07e2547779e35a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page