High-performance compressed binary format for DNA/RNA sequences with 2-bit encoding and Zstd compression
Project description
ZNA: Compressed Nucleic Acid Format
ZNA (Compressed Z-Nucleic N-Acid A) is a high-performance binary format for storing DNA/RNA sequences with exceptional compression and I/O speed.
Performance
- 135 MB/s roundtrip throughput (9.5x faster than Python baseline)
- 2.8+ GB/s encoding/decoding for long reads
- 3.7-4.0x compression ratio with Zstd
- C++ acceleration with pure Python fallback
Features
- High Compression: 2-bit encoding (4 bases per byte) + optional Zstd compression
- Ultra-Fast I/O: C++ accelerated encode/decode with block-based architecture
- Minimal Dependencies:
zstandardonly (C++ extension auto-compiled) - Flexible: Single-end, paired-end, and interleaved reads
- Strand-Specific Support: dUTP, TruSeq, and custom strand protocols
- Metadata Rich: Read groups, descriptions, and custom flags
- Unix-Friendly: Pipe-compatible CLI for seamless workflow integration
- Streaming: Memory-efficient block-based processing
Installation
# From source (recommended - includes C++ acceleration)
git clone https://github.com/mkiyer/zna.git
cd zna
pip install -e .
# Check if C++ acceleration is available
python -c "from zna.core import is_accelerated; print(f'Accelerated: {is_accelerated()}')"
Requirements:
- Python ≥3.10
- C++ compiler (for optimal performance)
- CMake ≥3.15 (auto-installed via pip)
Quick Start
# Encode FASTQ to compressed ZNA
zna encode sample.fastq.gz -o sample.zzna
# Decode back to FASTA
zna decode sample.zzna -o sample.fasta
# Inspect file statistics
zna inspect sample.zzna
# Pipe-friendly workflows
cat reads.fastq | zna encode -o reads.zzna
zna decode reads.zzna | head -n 1000
Performance Benchmarks
Throughput by Read Length
| Read Type | Encode (MB/s) | Decode (MB/s) | Compression |
|---|---|---|---|
| Short (Illumina, 100-150bp) | 189.5 | 668.8 | 3.68x |
| Medium (300-500bp) | 540.5 | 1,280.9 | 3.87x |
| Long (PacBio, 1-5kb) | 1,921.5 | 2,864.6 | 3.98x |
| Very Long (Nanopore, 5-15kb) | 2,824.7 | 3,392.7 | 3.99x |
Key Insights:
- Performance scales dramatically with read length
- Compression ratio remains consistent across workloads
- C++ acceleration provides 9.5x speedup over pure Python
See PERFORMANCE.md for detailed benchmarking.
File Format Specification
Overview
ZNA files use a binary format optimized for nucleic acid sequences:
- Magic Number:
ZNA\x1A(4 bytes) - Version: 1 (1 byte)
- 2-bit Encoding: A=00, C=01, G=10, T=11
- Block Structure: Data organized in compressed/uncompressed blocks
- Metadata: Read groups, descriptions, and custom information
File Structure
┌─────────────────────────────────────┐
│ File Header │
│ - Magic (4 bytes) │
│ - Version (1 byte) │
│ - Sequence length encoding (1 byte)│
│ - Flags (1 byte) │
│ - Compression method (1 byte) │
│ - Compression level (1 byte) │
│ - Metadata lengths (6 bytes) │
│ - Variable metadata strings │
├─────────────────────────────────────┤
│ Block 1 │
│ - Block Header (12 bytes) │
│ * Compressed size (4 bytes) │
│ * Uncompressed size (4 bytes) │
│ * Record count (4 bytes) │
│ - Compressed/Raw Payload │
│ * Record 1: flags, length, seq │
│ * Record 2: flags, length, seq │
│ * ... │
├─────────────────────────────────────┤
│ Block 2 │
│ ... │
└─────────────────────────────────────┘
Record Format
Each record in a block contains:
- Flags (1 byte): IS_READ1, IS_READ2, IS_PAIRED
- Length (1-4 bytes): Sequence length (configurable)
- Sequence (variable): 2-bit encoded bases
Compression
- Method 0: Uncompressed (
.zna) - Method 1: Zstd compression (
.zzna, levels 1-22) - Block Size: Default 128KB (configurable)
Usage Guide
Encoding
Single-End Reads
# From FASTQ file
zna encode sample.fastq -o sample.zna
# From FASTA file
zna encode sample.fasta -o sample.zna
# From gzipped input
zna encode sample.fastq.gz -o sample.zzna
# With compression
zna encode sample.fastq --zstd --level 5 -o sample.zzna
# From stdin
cat sample.fastq | zna encode -o sample.zna
# Force format (when extension detection fails)
cat data.txt | zna encode --fastq -o sample.zna
Paired-End Reads
# Separate R1/R2 files
zna encode R1.fastq.gz R2.fastq.gz -o paired.zzna
# Interleaved file (strict alternating R1/R2 pairs)
zna encode interleaved.fastq --interleaved -o paired.zzna
# Interleaved from stdin
cat interleaved.fastq | zna encode --interleaved -o paired.zzna
Mixed Paired-End and Single-End Reads (Interleaved)
The --interleaved mode intelligently detects both paired-end and single-end reads in the same file by analyzing read names. This is useful for output from tools like fastp that produce mixed merged (single) and unmerged (paired) reads.
How it works:
- Reads with matching base names (e.g.,
read1/1andread1/2) are paired - Reads without matching pairs are treated as single-end
- Read names are used to determine pairing (not just alternating order)
# Mixed interleaved input (fastp output with merged + unmerged reads)
zna encode fastp_output.fastq --interleaved -o mixed.zzna
# Example input structure:
# @read1/1 → paired with next read
# @read1/2
# @merged1 → single-end (no pair)
# @read2/1 → paired with next read
# @read2/2
# @merged2 → single-end (no pair)
Read name formats supported:
/1and/2suffixes:read1/1,read1/2- No suffix: treated as single-end unless next read has matching base name
- Comments ignored:
read1/1 merged_length:150extractsread1/1
Advanced Options
# Custom metadata
zna encode sample.fastq \
--read-group "Sample_01" \
--description "Experiment XYZ" \
-o sample.zzna
# Strand-specific library (default: R1 antisense, R2 sense)
zna encode R1.fastq.gz R2.fastq.gz \
--strand-specific \
-o stranded.zzna
# Custom strand orientation (e.g., fr-secondstrand protocol)
zna encode R1.fastq.gz R2.fastq.gz \
--strand-specific --read1-sense --read2-antisense \
-o stranded.zzna
# Handle sequences with N nucleotides
zna encode sample.fastq --npolicy drop -o clean.zzna # Skip sequences with N
zna encode sample.fastq --npolicy random -o clean.zzna # Replace N with random base
zna encode sample.fastq --npolicy A -o clean.zzna # Replace N with A
# Control compression
zna encode sample.fastq \
--zstd --level 9 \
--block-size 262144 \
-o sample.zzna
# Sequence length encoding (max sequence length)
zna encode sample.fastq \
--seq-len-bytes 1 \ # Max 255 bp
-o sample.zna
zna encode sample.fastq \
--seq-len-bytes 2 \ # Max 65,535 bp (default)
-o sample.zna
zna encode sample.fastq \
--seq-len-bytes 4 \ # Max 4.2 billion bp
-o sample.zna
Decoding
Basic Decoding
# To FASTA file
zna decode sample.zzna -o output.fasta
# To gzipped FASTA
zna decode sample.zzna -o output.fasta.gz
# To stdout (pipe-friendly)
zna decode sample.zzna | head -n 1000
# From stdin
cat sample.zzna | zna decode -o output.fasta
Paired-End Decoding
# Interleaved output (default)
zna decode paired.zzna -o interleaved.fasta
# Split to R1/R2 files (use # placeholder)
zna decode paired.zzna -o reads#.fasta
# Creates: reads_1.fasta and reads_2.fasta
# Split with gzip
zna decode paired.zzna -o reads#.fasta.gz
# Creates: reads_1.fasta.gz and reads_2.fasta.gz
# Restore original strand for strand-specific libraries
zna decode stranded.zzna --restore-strand -o reads.fasta
Piping Examples
# Extract first 1M reads
zna decode large.zzna | head -n 2000000 > subset.fasta
# Count sequences
zna decode sample.zzna | grep -c "^>"
# Convert to gzipped output via pipe
zna decode sample.zzna --gzip > output.fasta.gz
# Chain operations
zna decode sample.zzna | seqtk seq -r - | gzip > reversed.fasta.gz
Inspecting Files
# Show file statistics
zna inspect sample.zzna
Example Output:
File: sample.zzna
Total Size: 45.32 MB
--- Header Metadata ---
Read Group: Sample_01
Description: Experiment XYZ
Seq Length: 2 bytes (Max: 65535 bp)
Strand Specific: True
R1 Antisense: True
R2 Antisense: False
Compression: ZSTD (Level 3)
--- Content Statistics ---
Total Blocks: 356
Total Records: 1000000
Compressed Payload: 42.15 MB
Uncompressed Data: 125.50 MB
Compression Ratio: 2.98x
Command Reference
zna encode
Convert FASTQ/FASTA to ZNA format.
Usage:
zna encode [FILE1] [FILE2] [OPTIONS]
Positional Arguments:
FILE1 [FILE2] Input files (0=stdin, 1=single/interleaved, 2=paired R1 R2)
Options:
--interleaved Treat input as interleaved (auto-detects mixed paired/single reads)
--fasta Force FASTA format (overrides extension detection)
--fastq Force FASTQ format (overrides extension detection)
Metadata:
--read-group TEXT Read group ID (default: "Unknown")
--description TEXT Description string
--strand-specific Flag library as strand-specific (default: R1 antisense, R2 sense)
--read1-sense Read 1 represents sense strand
--read1-antisense Read 1 represents antisense strand (default when --strand-specific)
--read2-sense Read 2 represents sense strand (default when --strand-specific)
--read2-antisense Read 2 represents antisense strand
--npolicy {drop,random,A,C,G,T}
Policy for handling 'N' nucleotides:
- drop: skip sequences containing N
- random: replace N with random base (A/C/G/T)
- A/C/G/T: replace N with specific base
Format Options:
-o, --output FILE Output file (default: stdout)
--seq-len-bytes N Bytes for sequence length: 1, 2, or 4 (default: 2)
--block-size N Block size in bytes (default: 131072)
--zstd Force Zstd compression
--uncompressed Force uncompressed
--level N Zstd compression level 1-22 (default: 3)
zna decode
Convert ZNA to FASTA format.
Usage:
zna decode [FILE] [OPTIONS]
Positional Arguments:
FILE Input ZNA file (default: stdin)
Options:
-o, --output FILE Output FASTA file. Use '#' for split R1/R2
-q, --quiet Suppress progress messages
--gzip Force gzip compression for stdout
--restore-strand Restore original strand orientation for antisense reads
zna inspect
Display ZNA file statistics.
Usage:
zna inspect FILE
input FILE Input ZNA file to inspect
---
## Performance Characteristics
### Compression Ratios
Typical compression ratios compared to raw FASTQ:
| Format | Size | Ratio | Notes |
|--------|------|-------|-------|
| FASTQ (uncompressed) | 100% | 1.0x | Baseline |
| FASTQ.gz (gzip -6) | 25-30% | 3-4x | Standard |
| ZNA (uncompressed) | 12-15% | 6-8x | 2-bit encoding only |
| ZZNA (Zstd L3) | 8-10% | 10-12x | Fast compression |
| ZZNA (Zstd L9) | 6-8% | 12-16x | High compression |
*Results vary based on sequence complexity and redundancy*
### Speed
- **Encoding**: ~5-10M reads/second (single thread)
- **Decoding**: ~8-15M reads/second (single thread)
- **Block-based**: Enables parallel processing (future)
### Memory Usage
- **Streaming I/O**: Constant memory usage
- **Default block size**: 128KB buffer
- **No index required**: Sequential scan
---
## Technical Details
### 2-Bit Encoding
DNA bases are encoded in 2 bits:
A = 00 = 0 C = 01 = 1 G = 10 = 2 T = 11 = 3
Four bases pack into one byte:
Byte: [B1][B2][B3][B4] 76 54 32 10 (bit positions)
### Lookup Tables
Pre-computed lookup tables provide O(1) encoding/decoding:
- **Encoding**: 256-element array mapping ASCII → 2-bit
- **Decoding**: 256-element tuple mapping byte → 4-character string
### Block-Based Architecture
Data is organized in independently compressed blocks:
- **Advantages**: Random access, parallel processing potential
- **Overhead**: ~12 bytes per block
- **Optimal size**: 128KB balances compression ratio and I/O
### Compression Strategy
- **Zstd**: Modern compression algorithm (Facebook)
- **Reusable compressor**: Amortizes initialization cost
- **Memoryview parsing**: Zero-copy decompression
- **Pre-sized buffers**: Eliminates reallocations
---
## Strand-Specific Libraries
ZNA supports strand-specific RNA-seq libraries by normalizing all reads to sense strand orientation during encoding. This enables consistent downstream analysis while preserving the ability to restore original strand information.
### How It Works
1. **Encoding**: Antisense reads are reverse-complemented to sense strand
2. **Storage**: All reads stored in sense orientation
3. **Decoding**: Use `--restore-strand` to recover original orientation
### Strand Flags
| Flag | Description |
|------|-------------|
| `--strand-specific` | Enable strand-specific mode (default: R1 antisense, R2 sense) |
| `--read1-sense` | Read 1 represents sense strand |
| `--read1-antisense` | Read 1 represents antisense strand |
| `--read2-sense` | Read 2 represents sense strand |
| `--read2-antisense` | Read 2 represents antisense strand |
### Common Library Protocols
| Protocol | R1 | R2 | ZNA Flags |
|----------|----|----|-----------|
| **dUTP / TruSeq Stranded** | antisense | sense | `--strand-specific` (default) |
| **Illumina Stranded mRNA** | antisense | sense | `--strand-specific` |
| **fr-firststrand** | antisense | sense | `--strand-specific` |
| **fr-secondstrand** | sense | antisense | `--strand-specific --read1-sense --read2-antisense` |
| **Ligation (ScriptSeq)** | sense | antisense | `--strand-specific --read1-sense --read2-antisense` |
### Examples
```bash
# dUTP/TruSeq protocol (most common - this is the default)
zna encode R1.fastq.gz R2.fastq.gz --strand-specific -o library.zzna
# fr-secondstrand protocol
zna encode R1.fastq.gz R2.fastq.gz \
--strand-specific --read1-sense --read2-antisense \
-o library.zzna
# Decode with sense-normalized sequences (for alignment)
zna decode library.zzna -o normalized.fasta
# Decode with original strand orientation restored
zna decode library.zzna --restore-strand -o original.fasta
Use Cases
Recommended For
- ✅ Long-term archival: High compression with fast retrieval
- ✅ Data transfer: Reduced bandwidth requirements
- ✅ Cloud storage: Lower storage costs
- ✅ Pipeline integration: Unix-friendly streaming
- ✅ Reference storage: Efficient genome/transcriptome storage
Not Recommended For
- ❌ Random access: Sequential format (no index)
- ❌ Quality scores: Sequences only (use CRAM/BAM for qualities)
- ❌ Small files: Overhead outweighs benefits (<10K reads)
- ❌ Real-time streaming: Use case requires quality scores
Comparison with Other Formats
| Feature | ZNA | FASTA | FASTQ | CRAM | FASTA.gz |
|---|---|---|---|---|---|
| Compression | Excellent | None | None | Excellent | Good |
| Speed | Fast | Fastest | Fast | Slow | Medium |
| Quality Scores | ❌ | ❌ | ✅ | ✅ | ❌ |
| Paired-End | ✅ | ❌ | ❌ | ✅ | ❌ |
| Random Access | ❌ | ✅ | ✅ | ✅ | ❌ |
| Streaming | ✅ | ✅ | ✅ | Limited | ✅ |
| Dependencies | 1 | 0 | 0 | Many | 0 |
Python API
In addition to the CLI, ZNA provides a Python API:
from zna import ZnaHeader, ZnaWriter, ZnaReader, COMPRESSION_ZSTD
# Writing
header = ZnaHeader(
read_group="Sample_01",
compression_method=COMPRESSION_ZSTD,
compression_level=5
)
with open("output.zzna", "wb") as f:
with ZnaWriter(f, header) as writer:
writer.write_record("ACGTACGT", is_paired=False,
is_read1=False, is_read2=False)
writer.write_record("TGCATGCA", is_paired=False,
is_read1=False, is_read2=False)
# Reading
with open("output.zzna", "rb") as f:
reader = ZnaReader(f)
print(f"Read Group: {reader.header.read_group}")
for seq, is_paired, is_read1, is_read2 in reader.records():
print(seq)
Development
Running Tests
# All tests
PYTHONPATH=src pytest -v
# Specific test suite
PYTHONPATH=src pytest tests/test_cli.py -v
PYTHONPATH=src pytest tests/test_core.py -v
# With coverage
PYTHONPATH=src pytest --cov=zna tests/
Code Quality
# Format code
black src/ tests/
# Type checking
mypy src/zna/
Limitations
- Sequences only: No quality scores, headers, or annotations
- Sequential access: No random access without full scan
- DNA/RNA only: A, C, G, T bases (N or IUPAC codes not supported)
- Case insensitive: Lowercase converted to uppercase
- No index: Full file scan required for record counting
Future Enhancements
- Parallel compression/decompression
- Optional index for random access
- Support for IUPAC ambiguity codes
- Memory-mapped I/O for large files
- Streaming statistics (GC content, length distribution)
License
GNU General Public License v3.0 (GPLv3)
Citation
If you use ZNA in your research, please cite:
Iyer, M. (2026). ZNA: A compressed binary format for nucleic acid sequences.
GitHub: https://github.com/mkiyer/zna
Contributing
Contributions are welcome! Please:
- Fork the repository
- Create a feature branch
- Add tests for new functionality
- Ensure all tests pass
- Submit a pull request
Contact
- Author: Matthew Iyer
- Email: mkiyer@umich.edu
- Issues: https://github.com/mkiyer/zna/issues
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file zna-0.1.2.tar.gz.
File metadata
- Download URL: zna-0.1.2.tar.gz
- Upload date:
- Size: 65.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2782d58fdcb90627d9b2808ca25c631088e80bfd1d7b0bde2cf6b916e933e334
|
|
| MD5 |
6bd6552c36a62eb56a1ed155bda3266b
|
|
| BLAKE2b-256 |
9dcf84c1e928287c9b37422968c1fddfdec452383fa89eb4eddbd0c48e70e8fc
|
File details
Details for the file zna-0.1.2-cp314-cp314-macosx_15_0_arm64.whl.
File metadata
- Download URL: zna-0.1.2-cp314-cp314-macosx_15_0_arm64.whl
- Upload date:
- Size: 66.4 kB
- Tags: CPython 3.14, macOS 15.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b2497fa91510458221fd63fc15accc208f17c1f436d35d7e4e49f49cdea12350
|
|
| MD5 |
6b1b4db31888b77fdfa7554472ffccef
|
|
| BLAKE2b-256 |
3ae0d9acf88fa8b44d6cb7ca6de86150abf8336e6be06a8d39df04e73a51b324
|