Skip to main content

A unified file reader for Parquet and Avro files with automatic format detection

Project description

Snavro

A unified file reader for Parquet and Avro files with automatic format detection.

Features

  • 🚀 Automatic format detection - Just point to a file and Snavro figures out the format
  • 📊 Multiple format support - Handles .parquet, .parquet.snappy, and .avro files
  • 🖥️ CLI and Python API - Use from command line or import in your Python code
  • 📈 Smart data preview - Shows file info, shape, columns, and sample data
  • 🔧 Type hints - Full type annotation support for better IDE experience

Supported Formats

  • Parquet files (.parquet)
  • Snappy-compressed Parquet files (.parquet.snappy)
  • Avro files (.avro)

Installation

From PyPI (when published)

pip install snavro

From source

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e .

Usage

Command Line Interface

Interactive mode

snavro

This will show you all supported files in the current directory and let you choose which one to examine.

Direct file access

# Show first 5 rows (default)
snavro myfile.parquet

# Show first 10 rows
snavro myfile.avro 10

# Works with snappy-compressed files too
snavro data.parquet.snappy 3

Python API

Basic usage

import snavro

# Simple function to read and display file info
snavro.read_and_display_file("myfile.parquet", num_rows=5)

Using the FileReader class

from snavro import FileReader

reader = FileReader()

# Check if a file is supported
if reader.is_supported_file("myfile.parquet"):
    # Read the file (returns pandas DataFrame for Parquet, list of dicts for Avro)
    data = reader.read_file("myfile.parquet")
    
    # Or display file info with preview
    reader.display_file_info("myfile.parquet", num_rows=10)

Working with the data

from snavro import FileReader

reader = FileReader()

# For Parquet files - returns pandas DataFrame
df = reader.read_parquet("data.parquet")
print(f"Shape: {df.shape}")
print(df.head())

# For Avro files - returns list of dictionaries
records = reader.read_avro("data.avro")
print(f"Number of records: {len(records)}")
print(records[0])  # First record

Utility functions

from snavro import get_supported_files

# Get all supported files in current directory
files = get_supported_files()
print(files)

# Get supported files in specific directory
files = get_supported_files("/path/to/data")

Example Output

Parquet File

Reading Parquet file: sales_data.parquet
--------------------------------------------------
Shape: (1000, 5)
Columns: ['date', 'product', 'sales', 'region', 'msg']

First few rows:
        date product  sales region                    msg
0 2023-01-01   Widget    100   North  {"status": "active"}
1 2023-01-02   Gadget    150   South  {"status": "pending"}
2 2023-01-03   Widget     75    East  {"status": "active"}

First 'msg' value:
{"status": "active"}

Avro File

Reading Avro file: events.avro
--------------------------------------------------
Total records: 500

First 3 records:
Record 1:
{'timestamp': '2023-01-01T10:00:00Z', 'event': 'login', 'user_id': 12345}

Record 2:
{'timestamp': '2023-01-01T10:05:00Z', 'event': 'page_view', 'user_id': 12345}

Record 3:
{'timestamp': '2023-01-01T10:10:00Z', 'event': 'logout', 'user_id': 12345}

Development

Setup development environment

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e ".[dev]"

Run tests

pytest

Code formatting

black snavro/

Type checking

mypy snavro/

Dependencies

  • pandas >= 2.0.0 - For Parquet file handling
  • pyarrow >= 10.0.0 - Parquet engine with compression support
  • fastavro >= 1.8.0 - Fast Avro file reading

License

MIT License - see LICENSE file for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Changelog

0.1.0

  • Initial release
  • Support for Parquet and Avro files
  • CLI and Python API
  • Automatic format detection

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

snavro-0.1.0.tar.gz (7.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

snavro-0.1.0-py3-none-any.whl (7.4 kB view details)

Uploaded Python 3

File details

Details for the file snavro-0.1.0.tar.gz.

File metadata

  • Download URL: snavro-0.1.0.tar.gz
  • Upload date:
  • Size: 7.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for snavro-0.1.0.tar.gz
Algorithm Hash digest
SHA256 62a078d3b88c43f2cd0d013989188bb1a420bc3d3f5c1660ec6f2826ab695e26
MD5 243538a70beddc869a188245fc8d8403
BLAKE2b-256 0aca5132792b8dfda6ac3dfddb6faeaaa035262e782e5df55994c03728a9c749

See more details on using hashes here.

File details

Details for the file snavro-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: snavro-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 7.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for snavro-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 cb09fcc51100aaa27f669674b7a92a7ea36965e8de8ae9d6024d1250f95fc5d4
MD5 968fb0078d076129468ea30fb321dd24
BLAKE2b-256 df221f0a911c375188c4325bc33057316a614e29a24ab59a4065f7a0b0c034ad

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page