Skip to main content

A unified file reader for Parquet and Avro files with automatic format detection

Project description

Snavro

CI PyPI version Python versions

A unified file reader for Parquet and Avro files with automatic format detection.

Features

  • 🚀 Automatic format detection - Just point to a file and Snavro figures out the format
  • 📊 Multiple format support - Handles .parquet, .parquet.snappy, and .avro files
  • 🖥️ CLI and Python API - Use from command line or import in your Python code
  • 📈 Smart data preview - Shows file info, shape, columns, and sample data
  • 🔧 Type hints - Full type annotation support for better IDE experience

Supported Formats

  • Parquet files (.parquet)
  • Snappy-compressed Parquet files (.parquet.snappy)
  • Avro files (.avro)

Installation

From PyPI (when published)

pip install snavro

From source

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e .

Usage

Command Line Interface

Interactive mode

snavro

This will show you all supported files in the current directory and let you choose which one to examine.

Direct file access

# Show first 5 rows (default)
snavro myfile.parquet

# Show first 10 rows
snavro myfile.avro 10

# Works with snappy-compressed files too
snavro data.parquet.snappy 3

Python API

Basic usage

import snavro

# Simple function to read and display file info
snavro.read_and_display_file("myfile.parquet", num_rows=5)

Using the FileReader class

from snavro import FileReader

reader = FileReader()

# Check if a file is supported
if reader.is_supported_file("myfile.parquet"):
    # Read the file (returns pandas DataFrame for Parquet, list of dicts for Avro)
    data = reader.read_file("myfile.parquet")
    
    # Or display file info with preview
    reader.display_file_info("myfile.parquet", num_rows=10)

Working with the data

from snavro import FileReader

reader = FileReader()

# For Parquet files - returns pandas DataFrame
df = reader.read_parquet("data.parquet")
print(f"Shape: {df.shape}")
print(df.head())

# For Avro files - returns list of dictionaries
records = reader.read_avro("data.avro")
print(f"Number of records: {len(records)}")
print(records[0])  # First record

Utility functions

from snavro import get_supported_files

# Get all supported files in current directory
files = get_supported_files()
print(files)

# Get supported files in specific directory
files = get_supported_files("/path/to/data")

Example Output

Parquet File

Reading Parquet file: sales_data.parquet
--------------------------------------------------
Shape: (1000, 5)
Columns: ['date', 'product', 'sales', 'region', 'msg']

First few rows:
        date product  sales region                    msg
0 2023-01-01   Widget    100   North  {"status": "active"}
1 2023-01-02   Gadget    150   South  {"status": "pending"}
2 2023-01-03   Widget     75    East  {"status": "active"}

First 'msg' value:
{"status": "active"}

Avro File

Reading Avro file: events.avro
--------------------------------------------------
Total records: 500

First 3 records:
Record 1:
{'timestamp': '2023-01-01T10:00:00Z', 'event': 'login', 'user_id': 12345}

Record 2:
{'timestamp': '2023-01-01T10:05:00Z', 'event': 'page_view', 'user_id': 12345}

Record 3:
{'timestamp': '2023-01-01T10:10:00Z', 'event': 'logout', 'user_id': 12345}

Development

Setup development environment

git clone https://github.com/yourusername/snavro.git
cd snavro
pip install -e ".[dev]"

Run tests

pytest

Code formatting

black snavro/

Type checking

mypy snavro/

Dependencies

  • pandas >= 2.0.0 - For Parquet file handling
  • pyarrow >= 10.0.0 - Parquet engine with compression support
  • fastavro >= 1.8.0 - Fast Avro file reading

License

MIT License - see LICENSE file for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Development Setup

git clone https://github.com/SkinnyPigeon/snavro.git
cd snavro
pip install -e ".[dev]"

Running Tests

pytest tests/ -v

Code Quality

black snavro/ tests/
flake8 snavro/ tests/
mypy snavro/

Releases

This project uses automated releases via GitHub Actions. See RELEASE.md for details.

Changelog

0.1.0

  • Initial release
  • Support for Parquet and Avro files
  • CLI and Python API
  • Automatic format detection

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

snavro-0.1.9.tar.gz (17.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

snavro-0.1.9-py3-none-any.whl (9.0 kB view details)

Uploaded Python 3

File details

Details for the file snavro-0.1.9.tar.gz.

File metadata

  • Download URL: snavro-0.1.9.tar.gz
  • Upload date:
  • Size: 17.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for snavro-0.1.9.tar.gz
Algorithm Hash digest
SHA256 5b1a60df473e09674102b9567c00bfa3bf7344586ce7a5fb311f0279d18f443e
MD5 8d49ab941ef11cbbbbadeb35a6d5baff
BLAKE2b-256 a7c231a20eae747f6ba5946a6934127629c67696d9d7da3cdb8b600f4b74f4e6

See more details on using hashes here.

File details

Details for the file snavro-0.1.9-py3-none-any.whl.

File metadata

  • Download URL: snavro-0.1.9-py3-none-any.whl
  • Upload date:
  • Size: 9.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for snavro-0.1.9-py3-none-any.whl
Algorithm Hash digest
SHA256 41a4b4045ef6d6e6b705e994a5b129def415ae65a85579e9ee7a09592e7ed985
MD5 db7b6602f66152c79be58adcc7cc43f1
BLAKE2b-256 44f8fc015e38fba3aa01d12eb01fa3decd9cb5ca256b3d5159ac8bdbab93b17e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page