Skip to main content

A unified file reader for Parquet and Avro files with automatic format detection

Project description

Snavro

CI PyPI version Python versions

A unified file reader for Parquet and Avro files with automatic format detection.

Features

  • 🚀 Automatic format detection - Just point to a file and Snavro figures out the format
  • 📊 Multiple format support - Handles .parquet, .parquet.snappy, and .avro files
  • 🖥️ CLI and Python API - Use from command line or import in your Python code
  • 📈 Smart data preview - Shows file info, shape, columns, and sample data
  • 🔧 Type hints - Full type annotation support for better IDE experience

Supported Formats

  • Parquet files (.parquet)
  • Snappy-compressed Parquet files (.parquet.snappy)
  • Avro files (.avro)

Installation

From PyPI (when published)

pip install snavro

From source

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e .

Usage

Command Line Interface

Interactive mode

snavro

This will show you all supported files in the current directory and let you choose which one to examine.

Direct file access

# Show first 5 rows (default)
snavro myfile.parquet

# Show first 10 rows
snavro myfile.avro 10

# Works with snappy-compressed files too
snavro data.parquet.snappy 3

Python API

Basic usage

import snavro

# Simple function to read and display file info
snavro.read_and_display_file("myfile.parquet", num_rows=5)

Using the FileReader class

from snavro import FileReader

reader = FileReader()

# Check if a file is supported
if reader.is_supported_file("myfile.parquet"):
    # Read the file (returns pandas DataFrame for Parquet, list of dicts for Avro)
    data = reader.read_file("myfile.parquet")
    
    # Or display file info with preview
    reader.display_file_info("myfile.parquet", num_rows=10)

Working with the data

from snavro import FileReader

reader = FileReader()

# For Parquet files - returns pandas DataFrame
df = reader.read_parquet("data.parquet")
print(f"Shape: {df.shape}")
print(df.head())

# For Avro files - returns list of dictionaries
records = reader.read_avro("data.avro")
print(f"Number of records: {len(records)}")
print(records[0])  # First record

Utility functions

from snavro import get_supported_files

# Get all supported files in current directory
files = get_supported_files()
print(files)

# Get supported files in specific directory
files = get_supported_files("/path/to/data")

Example Output

Parquet File

Reading Parquet file: sales_data.parquet
--------------------------------------------------
Shape: (1000, 5)
Columns: ['date', 'product', 'sales', 'region', 'msg']

First few rows:
        date product  sales region                    msg
0 2023-01-01   Widget    100   North  {"status": "active"}
1 2023-01-02   Gadget    150   South  {"status": "pending"}
2 2023-01-03   Widget     75    East  {"status": "active"}

First 'msg' value:
{"status": "active"}

Avro File

Reading Avro file: events.avro
--------------------------------------------------
Total records: 500

First 3 records:
Record 1:
{'timestamp': '2023-01-01T10:00:00Z', 'event': 'login', 'user_id': 12345}

Record 2:
{'timestamp': '2023-01-01T10:05:00Z', 'event': 'page_view', 'user_id': 12345}

Record 3:
{'timestamp': '2023-01-01T10:10:00Z', 'event': 'logout', 'user_id': 12345}

Development

Setup development environment

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e ".[dev]"

Run tests

pytest

Code formatting

black snavro/

Type checking

mypy snavro/

Dependencies

  • pandas >= 2.0.0 - For Parquet file handling
  • pyarrow >= 10.0.0 - Parquet engine with compression support
  • fastavro >= 1.8.0 - Fast Avro file reading

License

MIT License - see LICENSE file for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Development Setup

git clone https://github.com/SkinnyPigeon/snavro.git
cd snavro
pip install -e ".[dev]"

Running Tests

pytest tests/ -v

Code Quality

black snavro/ tests/
flake8 snavro/ tests/
mypy snavro/

Releases

This project uses automated releases via GitHub Actions. See RELEASE.md for details.

Changelog

0.1.0

  • Initial release
  • Support for Parquet and Avro files
  • CLI and Python API
  • Automatic format detection

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

snavro-0.1.5.tar.gz (16.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

snavro-0.1.5-py3-none-any.whl (8.7 kB view details)

Uploaded Python 3

File details

Details for the file snavro-0.1.5.tar.gz.

File metadata

  • Download URL: snavro-0.1.5.tar.gz
  • Upload date:
  • Size: 16.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for snavro-0.1.5.tar.gz
Algorithm Hash digest
SHA256 8a26e5b52c9da0361bc320a07c921d0a709e584962b430fe4b81d1344db4c95f
MD5 a94f18230fe0bc5c0e5c07e38162de51
BLAKE2b-256 cf966d619b86621f3d87056183a728b1f7a13d4faffeecc78c27873ef0894547

See more details on using hashes here.

File details

Details for the file snavro-0.1.5-py3-none-any.whl.

File metadata

  • Download URL: snavro-0.1.5-py3-none-any.whl
  • Upload date:
  • Size: 8.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for snavro-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 f0f54d1a0d49185b56931b414c18f0c64cd874fac2b2e1fdebf2b5266a5e6918
MD5 077adb2a265ed21e902db1c0ab5032a3
BLAKE2b-256 dc081ca055340216ef9fe39ce9590ac2b69330f970bd8f984854ecbc7981f2e2

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page