Skip to main content

A unified file reader for Parquet and Avro files with automatic format detection

Project description

Snavro

A unified file reader for Parquet and Avro files with automatic format detection.

Features

  • 🚀 Automatic format detection - Just point to a file and Snavro figures out the format
  • 📊 Multiple format support - Handles .parquet, .parquet.snappy, and .avro files
  • 🖥️ CLI and Python API - Use from command line or import in your Python code
  • 📈 Smart data preview - Shows file info, shape, columns, and sample data
  • 🔧 Type hints - Full type annotation support for better IDE experience

Supported Formats

  • Parquet files (.parquet)
  • Snappy-compressed Parquet files (.parquet.snappy)
  • Avro files (.avro)

Installation

From PyPI (when published)

pip install snavro

From source

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e .

Usage

Command Line Interface

Interactive mode

snavro

This will show you all supported files in the current directory and let you choose which one to examine.

Direct file access

# Show first 5 rows (default)
snavro myfile.parquet

# Show first 10 rows
snavro myfile.avro 10

# Works with snappy-compressed files too
snavro data.parquet.snappy 3

Python API

Basic usage

import snavro

# Simple function to read and display file info
snavro.read_and_display_file("myfile.parquet", num_rows=5)

Using the FileReader class

from snavro import FileReader

reader = FileReader()

# Check if a file is supported
if reader.is_supported_file("myfile.parquet"):
    # Read the file (returns pandas DataFrame for Parquet, list of dicts for Avro)
    data = reader.read_file("myfile.parquet")
    
    # Or display file info with preview
    reader.display_file_info("myfile.parquet", num_rows=10)

Working with the data

from snavro import FileReader

reader = FileReader()

# For Parquet files - returns pandas DataFrame
df = reader.read_parquet("data.parquet")
print(f"Shape: {df.shape}")
print(df.head())

# For Avro files - returns list of dictionaries
records = reader.read_avro("data.avro")
print(f"Number of records: {len(records)}")
print(records[0])  # First record

Utility functions

from snavro import get_supported_files

# Get all supported files in current directory
files = get_supported_files()
print(files)

# Get supported files in specific directory
files = get_supported_files("/path/to/data")

Example Output

Parquet File

Reading Parquet file: sales_data.parquet
--------------------------------------------------
Shape: (1000, 5)
Columns: ['date', 'product', 'sales', 'region', 'msg']

First few rows:
        date product  sales region                    msg
0 2023-01-01   Widget    100   North  {"status": "active"}
1 2023-01-02   Gadget    150   South  {"status": "pending"}
2 2023-01-03   Widget     75    East  {"status": "active"}

First 'msg' value:
{"status": "active"}

Avro File

Reading Avro file: events.avro
--------------------------------------------------
Total records: 500

First 3 records:
Record 1:
{'timestamp': '2023-01-01T10:00:00Z', 'event': 'login', 'user_id': 12345}

Record 2:
{'timestamp': '2023-01-01T10:05:00Z', 'event': 'page_view', 'user_id': 12345}

Record 3:
{'timestamp': '2023-01-01T10:10:00Z', 'event': 'logout', 'user_id': 12345}

Development

Setup development environment

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e ".[dev]"

Run tests

pytest

Code formatting

black snavro/

Type checking

mypy snavro/

Dependencies

  • pandas >= 2.0.0 - For Parquet file handling
  • pyarrow >= 10.0.0 - Parquet engine with compression support
  • fastavro >= 1.8.0 - Fast Avro file reading

License

MIT License - see LICENSE file for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Changelog

0.1.0

  • Initial release
  • Support for Parquet and Avro files
  • CLI and Python API
  • Automatic format detection

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

snavro-0.1.1.tar.gz (9.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

snavro-0.1.1-py3-none-any.whl (7.8 kB view details)

Uploaded Python 3

File details

Details for the file snavro-0.1.1.tar.gz.

File metadata

  • Download URL: snavro-0.1.1.tar.gz
  • Upload date:
  • Size: 9.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for snavro-0.1.1.tar.gz
Algorithm Hash digest
SHA256 862d42f1cea1619e84d392644e745c63efbb171628d8bb42498c911fcfbef90b
MD5 fe0243b4b252cdb8a81d0f90f1479999
BLAKE2b-256 3442234d889c48751c73b35c3c7f843a37fa93041618f04a85a0f011530bfcb8

See more details on using hashes here.

File details

Details for the file snavro-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: snavro-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 7.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for snavro-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 b28ef55e54854d0fe9706ce2e52cae3e2e9bcdd16ce25c067b2d8014b0bdb3ed
MD5 d8c40820ef409de3357b61af7658115e
BLAKE2b-256 32bfb8a4cab1cb2d83247b28b2c9e52c3e70169ff88a4d6918b36367e1a3f06d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page