fastparquet

Python support for Parquet file format

These details have not been verified by PyPI

Project links

Homepage

Project description

https://travis-ci.org/jcrobak/parquet-python.svg?branch=master

fastparquet is a python implementation of the parquet format, aiming integrate into python-based big data work-flows.

Not all parts of the parquet-format have been implemented yet or tested e.g. see the Todos linked below. With that said, fastparquet is capable of reading all the data files from the parquet-compatability project.

Introduction

This software is alpha, expect frequent API changes and breakages.

A list of expected features and their status in this branch can be found in this issue, and further Please feel free to comment on that list as to missing items and priorities.

In the meantime, the more eyes on this code, the more example files and the more use cases the better.

Requirements

(all development is against recent versions in the default anaconda channels)

Required:

numba
numpy
pandas

Optional (compression algorithms; gzip is always available):

snappy
lzo
brotli

Installation

Install using conda:

conda install -c conda-forge fastparquet

install from pypi:

pip install fastparquet

or install latest version from github:

pip install git+https://github.com/dask/fastparquet

For the pip methods, numba must have been previously installed (using conda).

Usage

Reading

from fastparquet import ParquetFile
pf = ParquetFile('myfile.parq')
df = pf.to_pandas()
df2 = pf.to_pandas(['col1', 'col2'], categories=['col1'])

You may specify which columns to load, which of those to keep as categoricals (if the data uses dictionary encoding). The file-path can be a single file, a metadata file pointing to other data files, or a directory (tree) containing data files. The latter is what is typically output by hive/spark.

Writing

from fastparquet import write
write('outfile.parq', df)
write('outfile2.parq', df, row_group_offsets=[0, 10000, 20000],
      compression='GZIP', file_scheme='hive')

The default is to produce a single output file with a single row-group (i.e., logical segment) and no compression. At the moment, only simple data-types and plain encoding are supported, so expect performance to be similar to numpy.savez.

History

Since early October 2016, this fork of parquet-python has been undergoing considerable redevelopment. The aim is to have a small and simple and performant library for reading and writing the parquet format from python.

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

2026.3.0

Mar 17, 2026

2025.12.0

Dec 18, 2025

2024.11.0

Nov 12, 2024

2024.5.0

May 21, 2024

2024.2.0

Feb 8, 2024

2023.10.1

Oct 26, 2023

2023.10.0

Oct 25, 2023

2023.8.0

Aug 30, 2023

2023.7.0

Jul 1, 2023

2023.4.0

Apr 27, 2023

2023.2.0

Feb 8, 2023

2023.1.0

Jan 19, 2023

2022.12.0

Dec 5, 2022

2022.11.0

Nov 17, 2022

0.8.3

Aug 28, 2022

0.8.2

Aug 19, 2022

0.8.1

Apr 1, 2022

0.8.0

Jan 26, 2022

0.7.2

Nov 22, 2021

0.7.1

Aug 3, 2021

0.7.0

Jul 16, 2021

0.6.3

May 13, 2021

0.6.2

May 12, 2021

0.6.1

May 11, 2021

0.6.0.post1

May 6, 2021

0.6.0

May 6, 2021

0.5.0

Dec 29, 2020

0.4.2 yanked

Dec 14, 2020

Reason this release was yanked:

deps not updated

0.4.1

Jul 16, 2020

0.4.0

May 12, 2020

0.3.3

Feb 5, 2020

0.3.2

Aug 1, 2019

0.3.1

Apr 25, 2019

0.3.0

Mar 30, 2019

0.2.1

Dec 18, 2018

0.2.0

Nov 22, 2018

0.1.6

Aug 19, 2018

0.1.5

Apr 1, 2018

0.1.4

Jan 27, 2018

0.1.3

Oct 8, 2017

0.1.2

Aug 28, 2017

0.1.1

Jul 21, 2017

0.1.0

Jun 13, 2017

0.0.6

May 4, 2017

0.0.5

Feb 16, 2017

0.0.4.post1

Dec 27, 2016

This version

0.0.4

Dec 27, 2016

0.0.3

Dec 1, 2016

0.0.2

Nov 15, 2016

0.0.1.post2

Nov 1, 2016

0.0.1.post1

Nov 1, 2016

0.0.1

Nov 1, 2016

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fastparquet-0.0.4.tar.gz (54.5 kB view details)

Uploaded Dec 27, 2016 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

fastparquet-0.0.4-py2.py3-none-any.whl (53.6 kB view details)

Uploaded Dec 27, 2016 Python 2Python 3

File details

Details for the file fastparquet-0.0.4.tar.gz.

File metadata

Download URL: fastparquet-0.0.4.tar.gz
Upload date: Dec 27, 2016
Size: 54.5 kB
Tags: Source
Uploaded using Trusted Publishing? No

File hashes

Hashes for fastparquet-0.0.4.tar.gz
Algorithm	Hash digest
SHA256	`4293f90e0ab03dcfee384f55332bb8464cf7e043dfddd19902e916a1dfadb1fb`
MD5	`3d035d79874e11cb01921833b8d621b8`
BLAKE2b-256	`e53935be1384886dd9505d6adb782a57e9a08004209bbbe88d91dd843bc14f53`

See more details on using hashes here.

File details

Details for the file fastparquet-0.0.4-py2.py3-none-any.whl.

File metadata

Download URL: fastparquet-0.0.4-py2.py3-none-any.whl
Upload date: Dec 27, 2016
Size: 53.6 kB
Tags: Python 2, Python 3
Uploaded using Trusted Publishing? No

File hashes

Hashes for fastparquet-0.0.4-py2.py3-none-any.whl
Algorithm	Hash digest
SHA256	`1fb04c181b27385c4f4953f90da9faf9cafc945424cb72dc5939380fe8e68ea3`
MD5	`7199547baf412df53037563b1ca9a33c`
BLAKE2b-256	`d6a4643ab14d90e00e826f2e417306cdddf739c4235b4842827cebf8952697eb`

See more details on using hashes here.

fastparquet 0.0.4

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Introduction

Requirements

Installation

Usage

History

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes