Skip to main content

S3 Parquetifier

Build Status PyPI version fury.io MIT license

S3 Parquetifier is an ETL tool that can take a file from an S3 bucket convert it to Parquet format and save it to another bucket.

S3 Parquetifier supports the following file types

  • CSV
  • JSON
  • TSV

Instructions

How to install

To install the package just run the following

sudo apt-get install -y libssl-dev libffi-dev &&
sudo apt-get install -y libxml2-dev libxslt1-dev &&
sudo apt-get install -y libsnappy-dev
pip install s3-parquetifier

How to use it

S3 parquetifier needs an AWS Account that will have at least read rights for the target bucket and read-write rights for the destination bucket.

You can read the following article on how to set up S3 roles and policies here

Running the Script

from s3_parquetifier import S3Parquetifier

# Call the covertor
S3Parquetifier(
    source_bucket="<the bucket's name where the CSVs are>",
    target_bucket="<the bucket's name where you want the parquet file to be saved>",
    verbose=True,  # for verbosity or not
).convert_from_s3(
    source_key="<the key of the S3 object>",
    target_key="<the key of the S3 object>",
    chunk_size=100000  # The number of rows per parquet
)
from s3_parquetifier import S3Parquetifier

# Call the covertor
S3Parquetifier(
    target_bucket="<the bucket's name where you want the parquet file to be saved>",
    verbose=True,  # for verbosity or not
).convert_from_local(
    file_name='<The CSV file that you want to transform>',
    target_key='<The S3 bucket key where the file will be saved>',
    chunk_size=100000,
)

Adding custom pre-processing function

You can add custom pre-processing function on your source file. Because this tool is designed for large files the preprocessing is taking place on every chunk separately. If the full file is needed for the preprocessing then a local preprocessing is needed in the source file.

In the following example, we are going to add custom columns on the chunk with some custom values. We are going to add the columns test1, test2, test3 with the values 1, 2, 3 respectively.

We define our function bellow named pre_process and we also define the arguments for the function kwargs. The chunk DataFrame is not needed in the kwargs, it is taken by default. You have to pass your function as an argument in pre_process_chunk and the arguments in kwargs in the convert_from_s3 method.

from s3_parquetifier import S3Parquetifier


# Add three new columns with custom values
def pre_process(chunk, columns=None, values=None):

    for index, column in enumerate(columns):
        chunk[column] = values[index]

    return chunk

# define the arguments for the pre-processor
kwargs = {
    'columns': ['test1', 'test2', 'test3'],
    'values': [1, 2, 3]
}

# Call the covertor
S3Parquetifier(
    source_bucket="<the bucket's name where the CSVs are>",
    target_bucket="<the bucket's name where you want the parquet file to be saved>",
    verbose=True,  # for verbosity or not
).convert_from_s3(
    source_key='<the key of the S3 object>',
    target_key='<the key of the S3 object>',
    chunk_size=100000  # The number of rows per parquet
    pre_process_chunk=pre_process,  # A preprocessing function that will pre-process the each chunk
    kwargs=kwargs  # potential extra arguments for the pre-preocess function
)

ToDo

  • Add support to handle local files too
  • Add support for JSON
  • Add streaming from url support

Release files for s3-parquetifier 0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for s3-parquetifier 0.2
File Size Uploaded
s3-parquetifier-0.2.tar.gz 6.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for s3-parquetifier 0.2
File Interpreter ABI Platform
s3_parquetifier-0.2-py3-none-any.whl Python 3 none any Details

Total release size: 14.2 kB

Release files / s3-parquetifier-0.2.tar.gz

Download URL s3-parquetifier-0.2.tar.gz
Size 6.4 kB
Tags Source
SHA-256 checksum
How to use checksums
3e9ec61f140be0b30fd5c5a28d38346da5957c62822057ed4b38c3cc7bada72f
BLAKE2b-256 checksum
How to use checksums
ff84ace06ebba86d9fec789ce7e7a608b70f506dc33bad379a3b7bd2b88f697f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/1.13.0 pkginfo/1.5.0.1 requests/2.21.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.31.1 CPython/3.6.9

Release files / s3_parquetifier-0.2-py3-none-any.whl

Download URL s3_parquetifier-0.2-py3-none-any.whl
Size 7.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cae454a15776635ec81749fac1a32f0c40dd395c63eb116f78da83d22e0fb16e
BLAKE2b-256 checksum
How to use checksums
433bfe22662b9775685fb0d3ad1eb6a2d13690025d428cda3cd6ab04ed89c4e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/1.13.0 pkginfo/1.5.0.1 requests/2.21.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.31.1 CPython/3.6.9

Release history Release notifications | RSS feed

This release

0.2 This release

2 release files

0.1.1

1 release file

0.1.0

1 release file

0.0.9

1 release file

0.0.8

1 release file

0.0.7

1 release file

0.0.6

1 release file

0.0.3

1 release file

0.0.2

2 release files

0.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page