Skip to main content

A unified file reader for Parquet and Avro files with automatic format detection

Project description

Snavro

CI PyPI version Python versions

A unified file reader for Parquet and Avro files with automatic format detection.

Features

  • 🚀 Automatic format detection - Just point to a file and Snavro figures out the format
  • 📊 Multiple format support - Handles .parquet, .parquet.snappy, and .avro files
  • 🖥️ CLI and Python API - Use from command line or import in your Python code
  • 📈 Smart data preview - Shows file info, shape, columns, and sample data
  • 🔧 Type hints - Full type annotation support for better IDE experience

Supported Formats

  • Parquet files (.parquet)
  • Snappy-compressed Parquet files (.parquet.snappy)
  • Avro files (.avro)

Installation

From PyPI (when published)

pip install snavro

From source

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e .

Usage

Command Line Interface

Interactive mode

snavro

This will show you all supported files in the current directory and let you choose which one to examine.

Direct file access

# Show first 5 rows (default)
snavro myfile.parquet

# Show first 10 rows
snavro myfile.avro 10

# Works with snappy-compressed files too
snavro data.parquet.snappy 3

Python API

Basic usage

import snavro

# Simple function to read and display file info
snavro.read_and_display_file("myfile.parquet", num_rows=5)

Using the FileReader class

from snavro import FileReader

reader = FileReader()

# Check if a file is supported
if reader.is_supported_file("myfile.parquet"):
    # Read the file (returns pandas DataFrame for Parquet, list of dicts for Avro)
    data = reader.read_file("myfile.parquet")
    
    # Or display file info with preview
    reader.display_file_info("myfile.parquet", num_rows=10)

Working with the data

from snavro import FileReader

reader = FileReader()

# For Parquet files - returns pandas DataFrame
df = reader.read_parquet("data.parquet")
print(f"Shape: {df.shape}")
print(df.head())

# For Avro files - returns list of dictionaries
records = reader.read_avro("data.avro")
print(f"Number of records: {len(records)}")
print(records[0])  # First record

Utility functions

from snavro import get_supported_files

# Get all supported files in current directory
files = get_supported_files()
print(files)

# Get supported files in specific directory
files = get_supported_files("/path/to/data")

Example Output

Parquet File

Reading Parquet file: sales_data.parquet
--------------------------------------------------
Shape: (1000, 5)
Columns: ['date', 'product', 'sales', 'region', 'msg']

First few rows:
        date product  sales region                    msg
0 2023-01-01   Widget    100   North  {"status": "active"}
1 2023-01-02   Gadget    150   South  {"status": "pending"}
2 2023-01-03   Widget     75    East  {"status": "active"}

First 'msg' value:
{"status": "active"}

Avro File

Reading Avro file: events.avro
--------------------------------------------------
Total records: 500

First 3 records:
Record 1:
{'timestamp': '2023-01-01T10:00:00Z', 'event': 'login', 'user_id': 12345}

Record 2:
{'timestamp': '2023-01-01T10:05:00Z', 'event': 'page_view', 'user_id': 12345}

Record 3:
{'timestamp': '2023-01-01T10:10:00Z', 'event': 'logout', 'user_id': 12345}

Development

Setup development environment

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e ".[dev]"

Run tests

pytest

Code formatting

black snavro/

Type checking

mypy snavro/

Dependencies

  • pandas >= 2.0.0 - For Parquet file handling
  • pyarrow >= 10.0.0 - Parquet engine with compression support
  • fastavro >= 1.8.0 - Fast Avro file reading

License

MIT License - see LICENSE file for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Development Setup

git clone https://github.com/SkinnyPigeon/snavro.git
cd snavro
pip install -e ".[dev]"

Running Tests

pytest tests/ -v

Code Quality

black snavro/ tests/
flake8 snavro/ tests/
mypy snavro/

Releases

This project uses automated releases via GitHub Actions. See RELEASE.md for details.

Changelog

0.1.0

  • Initial release
  • Support for Parquet and Avro files
  • CLI and Python API
  • Automatic format detection

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

snavro-0.1.3.tar.gz (16.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

snavro-0.1.3-py3-none-any.whl (8.7 kB view details)

Uploaded Python 3

File details

Details for the file snavro-0.1.3.tar.gz.

File metadata

  • Download URL: snavro-0.1.3.tar.gz
  • Upload date:
  • Size: 16.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for snavro-0.1.3.tar.gz
Algorithm Hash digest
SHA256 8c6c02d7ea25ee1b44a8355eaf00f230792c26d3d3b58b6bbf187d0b2ff1cb25
MD5 72c514c9ca7df07e8ce1b3ea860d7af2
BLAKE2b-256 460cf6b98580d17437d90ecb1d61b0ab7e22a22f95ca1f49c12eb76ec44096ee

See more details on using hashes here.

File details

Details for the file snavro-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: snavro-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 8.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for snavro-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 143a3b1b5ba34f2664cbb445aabf30073f69bf269306b4a6476a38a2deb7603c
MD5 97cd8723e137c6c540b04a9fdd082fda
BLAKE2b-256 76fe5216308fd4080d4120feae99757a32c25d9c4ceb664e7db93ddadee979fc

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page