Skip to main content

High-performance compressed binary format for DNA/RNA sequences with 2-bit encoding and Zstd compression

Project description

ZNA: Compressed Nucleic Acid Format

ZNA (Compressed Z-Nucleic N-Acid A) is a high-performance binary format for storing DNA/RNA sequences with exceptional compression and I/O speed.

Performance

  • 135 MB/s roundtrip throughput (9.5x faster than Python baseline)
  • 2.8+ GB/s encoding/decoding for long reads
  • 3.7-4.0x compression ratio with Zstd
  • C++ acceleration with pure Python fallback

Features

  • High Compression: 2-bit encoding (4 bases per byte) + optional Zstd compression
  • Ultra-Fast I/O: C++ accelerated encode/decode with block-based architecture
  • Minimal Dependencies: zstandard only (C++ extension auto-compiled)
  • Flexible: Single-end, paired-end, and interleaved reads
  • Strand-Specific Support: dUTP, TruSeq, and custom strand protocols
  • Metadata Rich: Read groups, descriptions, and custom flags
  • Unix-Friendly: Pipe-compatible CLI for seamless workflow integration
  • Streaming: Memory-efficient block-based processing

Installation

# From source (recommended - includes C++ acceleration)
git clone https://github.com/mkiyer/zna.git
cd zna
pip install -e .

# Check if C++ acceleration is available
python -c "from zna.core import is_accelerated; print(f'Accelerated: {is_accelerated()}')"

Requirements:

  • Python ≥3.10
  • C++ compiler (for optimal performance)
  • CMake ≥3.15 (auto-installed via pip)

Quick Start

# Encode FASTQ to compressed ZNA
zna encode sample.fastq.gz -o sample.zzna

# Decode back to FASTA
zna decode sample.zzna -o sample.fasta

# Inspect file statistics
zna inspect sample.zzna

# Pipe-friendly workflows
cat reads.fastq | zna encode -o reads.zzna
zna decode reads.zzna | head -n 1000

Performance Benchmarks

Throughput by Read Length

Read Type Encode (MB/s) Decode (MB/s) Compression
Short (Illumina, 100-150bp) 189.5 668.8 3.68x
Medium (300-500bp) 540.5 1,280.9 3.87x
Long (PacBio, 1-5kb) 1,921.5 2,864.6 3.98x
Very Long (Nanopore, 5-15kb) 2,824.7 3,392.7 3.99x

Key Insights:

  • Performance scales dramatically with read length
  • Compression ratio remains consistent across workloads
  • C++ acceleration provides 9.5x speedup over pure Python

See PERFORMANCE.md for detailed benchmarking.


File Format Specification

Overview

ZNA files use a binary format optimized for nucleic acid sequences:

  • Magic Number: ZNA\x1A (4 bytes)
  • Version: 1 (1 byte)
  • 2-bit Encoding: A=00, C=01, G=10, T=11
  • Block Structure: Data organized in compressed/uncompressed blocks
  • Metadata: Read groups, descriptions, and custom information

File Structure

┌─────────────────────────────────────┐
│         File Header                 │
│  - Magic (4 bytes)                  │
│  - Version (1 byte)                 │
│  - Sequence length encoding (1 byte)│
│  - Flags (1 byte)                   │
│  - Compression method (1 byte)      │
│  - Compression level (1 byte)       │
│  - Metadata lengths (6 bytes)       │
│  - Variable metadata strings        │
├─────────────────────────────────────┤
│         Block 1                     │
│  - Block Header (12 bytes)          │
│    * Compressed size (4 bytes)      │
│    * Uncompressed size (4 bytes)    │
│    * Record count (4 bytes)         │
│  - Compressed/Raw Payload           │
│    * Record 1: flags, length, seq   │
│    * Record 2: flags, length, seq   │
│    * ...                            │
├─────────────────────────────────────┤
│         Block 2                     │
│  ...                                │
└─────────────────────────────────────┘

Record Format

Each record in a block contains:

  • Flags (1 byte): IS_READ1, IS_READ2, IS_PAIRED
  • Length (1-4 bytes): Sequence length (configurable)
  • Sequence (variable): 2-bit encoded bases

Compression

  • Method 0: Uncompressed (.zna)
  • Method 1: Zstd compression (.zzna, levels 1-22)
  • Block Size: Default 128KB (configurable)

Usage Guide

Encoding

Single-End Reads

# From FASTQ file
zna encode sample.fastq -o sample.zna

# From FASTA file  
zna encode sample.fasta -o sample.zna

# From gzipped input
zna encode sample.fastq.gz -o sample.zzna

# With compression
zna encode sample.fastq --zstd --level 5 -o sample.zzna

# From stdin
cat sample.fastq | zna encode -o sample.zna

# Force format (when extension detection fails)
cat data.txt | zna encode --fastq -o sample.zna

Paired-End Reads

# Separate R1/R2 files
zna encode R1.fastq.gz R2.fastq.gz -o paired.zzna

# Interleaved file (strict alternating R1/R2 pairs)
zna encode interleaved.fastq --interleaved -o paired.zzna

# Interleaved from stdin
cat interleaved.fastq | zna encode --interleaved -o paired.zzna

Mixed Paired-End and Single-End Reads (Interleaved)

The --interleaved mode intelligently detects both paired-end and single-end reads in the same file by analyzing read names. This is useful for output from tools like fastp that produce mixed merged (single) and unmerged (paired) reads.

How it works:

  • Reads with matching base names (e.g., read1/1 and read1/2) are paired
  • Reads without matching pairs are treated as single-end
  • Read names are used to determine pairing (not just alternating order)
# Mixed interleaved input (fastp output with merged + unmerged reads)
zna encode fastp_output.fastq --interleaved -o mixed.zzna

# Example input structure:
#   @read1/1         →  paired with next read
#   @read1/2
#   @merged1         →  single-end (no pair)
#   @read2/1         →  paired with next read
#   @read2/2
#   @merged2         →  single-end (no pair)

Read name formats supported:

  • /1 and /2 suffixes: read1/1, read1/2
  • No suffix: treated as single-end unless next read has matching base name
  • Comments ignored: read1/1 merged_length:150 extracts read1/1

Advanced Options

# Custom metadata
zna encode sample.fastq \
  --read-group "Sample_01" \
  --description "Experiment XYZ" \
  -o sample.zzna

# Strand-specific library (default: R1 antisense, R2 sense)
zna encode R1.fastq.gz R2.fastq.gz \
  --strand-specific \
  -o stranded.zzna

# Custom strand orientation (e.g., fr-secondstrand protocol)
zna encode R1.fastq.gz R2.fastq.gz \
  --strand-specific --read1-sense --read2-antisense \
  -o stranded.zzna

# Handle sequences with N nucleotides
zna encode sample.fastq --npolicy drop -o clean.zzna       # Skip sequences with N
zna encode sample.fastq --npolicy random -o clean.zzna     # Replace N with random base
zna encode sample.fastq --npolicy A -o clean.zzna          # Replace N with A

# Control compression
zna encode sample.fastq \
  --zstd --level 9 \
  --block-size 262144 \
  -o sample.zzna

# Sequence length encoding (max sequence length)
zna encode sample.fastq \
  --seq-len-bytes 1 \  # Max 255 bp
  -o sample.zna

zna encode sample.fastq \
  --seq-len-bytes 2 \  # Max 65,535 bp (default)
  -o sample.zna

zna encode sample.fastq \
  --seq-len-bytes 4 \  # Max 4.2 billion bp
  -o sample.zna

Decoding

Basic Decoding

# To FASTA file
zna decode sample.zzna -o output.fasta

# To gzipped FASTA
zna decode sample.zzna -o output.fasta.gz

# To stdout (pipe-friendly)
zna decode sample.zzna | head -n 1000

# From stdin
cat sample.zzna | zna decode -o output.fasta

Paired-End Decoding

# Interleaved output (default)
zna decode paired.zzna -o interleaved.fasta

# Split to R1/R2 files (use # placeholder)
zna decode paired.zzna -o reads#.fasta
# Creates: reads_1.fasta and reads_2.fasta

# Split with gzip
zna decode paired.zzna -o reads#.fasta.gz
# Creates: reads_1.fasta.gz and reads_2.fasta.gz

# Restore original strand for strand-specific libraries
zna decode stranded.zzna --restore-strand -o reads.fasta

Piping Examples

# Extract first 1M reads
zna decode large.zzna | head -n 2000000 > subset.fasta

# Count sequences
zna decode sample.zzna | grep -c "^>"

# Convert to gzipped output via pipe
zna decode sample.zzna --gzip > output.fasta.gz

# Chain operations
zna decode sample.zzna | seqtk seq -r - | gzip > reversed.fasta.gz

Inspecting Files

# Show file statistics
zna inspect sample.zzna

Example Output:

File: sample.zzna
Total Size: 45.32 MB

--- Header Metadata ---
Read Group:       Sample_01
Description:      Experiment XYZ
Seq Length:       2 bytes (Max: 65535 bp)
Strand Specific:  True
R1 Antisense:     True
R2 Antisense:     False
Compression:      ZSTD (Level 3)

--- Content Statistics ---
Total Blocks:       356
Total Records:      1000000
Compressed Payload: 42.15 MB
Uncompressed Data:  125.50 MB
Compression Ratio:  2.98x

Command Reference

zna encode

Convert FASTQ/FASTA to ZNA format.

Usage:

zna encode [FILE1] [FILE2] [OPTIONS]

Positional Arguments:
  FILE1 [FILE2]          Input files (0=stdin, 1=single/interleaved, 2=paired R1 R2)

Options:
  --interleaved          Treat input as interleaved (auto-detects mixed paired/single reads)
  --fasta                Force FASTA format (overrides extension detection)
  --fastq                Force FASTQ format (overrides extension detection)

Metadata:
  --read-group TEXT      Read group ID (default: "Unknown")
  --description TEXT     Description string
  --strand-specific      Flag library as strand-specific (default: R1 antisense, R2 sense)
  --read1-sense          Read 1 represents sense strand
  --read1-antisense      Read 1 represents antisense strand (default when --strand-specific)
  --read2-sense          Read 2 represents sense strand (default when --strand-specific)
  --read2-antisense      Read 2 represents antisense strand
  --npolicy {drop,random,A,C,G,T}
                         Policy for handling 'N' nucleotides:
                         - drop: skip sequences containing N
                         - random: replace N with random base (A/C/G/T)
                         - A/C/G/T: replace N with specific base

Format Options:
  -o, --output FILE      Output file (default: stdout)
  --seq-len-bytes N      Bytes for sequence length: 1, 2, or 4 (default: 2)
  --block-size N         Block size in bytes (default: 131072)
  --zstd                 Force Zstd compression
  --uncompressed         Force uncompressed
  --level N              Zstd compression level 1-22 (default: 3)

zna decode

Convert ZNA to FASTA format.

Usage:

zna decode [FILE] [OPTIONS]

Positional Arguments:
  FILE                   Input ZNA file (default: stdin)

Options:
  -o, --output FILE      Output FASTA file. Use '#' for split R1/R2
  -q, --quiet            Suppress progress messages
  --gzip                 Force gzip compression for stdout
  --restore-strand       Restore original strand orientation for antisense reads

zna inspect

Display ZNA file statistics.

Usage:

zna inspect FILE

input FILE Input ZNA file to inspect


---

## Performance Characteristics

### Compression Ratios

Typical compression ratios compared to raw FASTQ:

| Format | Size | Ratio | Notes |
|--------|------|-------|-------|
| FASTQ (uncompressed) | 100% | 1.0x | Baseline |
| FASTQ.gz (gzip -6) | 25-30% | 3-4x | Standard |
| ZNA (uncompressed) | 12-15% | 6-8x | 2-bit encoding only |
| ZZNA (Zstd L3) | 8-10% | 10-12x | Fast compression |
| ZZNA (Zstd L9) | 6-8% | 12-16x | High compression |

*Results vary based on sequence complexity and redundancy*

### Speed

- **Encoding**: ~5-10M reads/second (single thread)
- **Decoding**: ~8-15M reads/second (single thread)
- **Block-based**: Enables parallel processing (future)

### Memory Usage

- **Streaming I/O**: Constant memory usage
- **Default block size**: 128KB buffer
- **No index required**: Sequential scan

---

## Technical Details

### 2-Bit Encoding

DNA bases are encoded in 2 bits:

A = 00 = 0 C = 01 = 1 G = 10 = 2 T = 11 = 3


Four bases pack into one byte:

Byte: [B1][B2][B3][B4] 76 54 32 10 (bit positions)


### Lookup Tables

Pre-computed lookup tables provide O(1) encoding/decoding:

- **Encoding**: 256-element array mapping ASCII → 2-bit
- **Decoding**: 256-element tuple mapping byte → 4-character string

### Block-Based Architecture

Data is organized in independently compressed blocks:

- **Advantages**: Random access, parallel processing potential
- **Overhead**: ~12 bytes per block
- **Optimal size**: 128KB balances compression ratio and I/O

### Compression Strategy

- **Zstd**: Modern compression algorithm (Facebook)
- **Reusable compressor**: Amortizes initialization cost
- **Memoryview parsing**: Zero-copy decompression
- **Pre-sized buffers**: Eliminates reallocations

---

## Strand-Specific Libraries

ZNA supports strand-specific RNA-seq libraries by normalizing all reads to sense strand orientation during encoding. This enables consistent downstream analysis while preserving the ability to restore original strand information.

### How It Works

1. **Encoding**: Antisense reads are reverse-complemented to sense strand
2. **Storage**: All reads stored in sense orientation
3. **Decoding**: Use `--restore-strand` to recover original orientation

### Strand Flags

| Flag | Description |
|------|-------------|
| `--strand-specific` | Enable strand-specific mode (default: R1 antisense, R2 sense) |
| `--read1-sense` | Read 1 represents sense strand |
| `--read1-antisense` | Read 1 represents antisense strand |
| `--read2-sense` | Read 2 represents sense strand |
| `--read2-antisense` | Read 2 represents antisense strand |

### Common Library Protocols

| Protocol | R1 | R2 | ZNA Flags |
|----------|----|----|-----------|
| **dUTP / TruSeq Stranded** | antisense | sense | `--strand-specific` (default) |
| **Illumina Stranded mRNA** | antisense | sense | `--strand-specific` |
| **fr-firststrand** | antisense | sense | `--strand-specific` |
| **fr-secondstrand** | sense | antisense | `--strand-specific --read1-sense --read2-antisense` |
| **Ligation (ScriptSeq)** | sense | antisense | `--strand-specific --read1-sense --read2-antisense` |

### Examples

```bash
# dUTP/TruSeq protocol (most common - this is the default)
zna encode R1.fastq.gz R2.fastq.gz --strand-specific -o library.zzna

# fr-secondstrand protocol
zna encode R1.fastq.gz R2.fastq.gz \
  --strand-specific --read1-sense --read2-antisense \
  -o library.zzna

# Decode with sense-normalized sequences (for alignment)
zna decode library.zzna -o normalized.fasta

# Decode with original strand orientation restored
zna decode library.zzna --restore-strand -o original.fasta

Use Cases

Recommended For

  • Long-term archival: High compression with fast retrieval
  • Data transfer: Reduced bandwidth requirements
  • Cloud storage: Lower storage costs
  • Pipeline integration: Unix-friendly streaming
  • Reference storage: Efficient genome/transcriptome storage

Not Recommended For

  • Random access: Sequential format (no index)
  • Quality scores: Sequences only (use CRAM/BAM for qualities)
  • Small files: Overhead outweighs benefits (<10K reads)
  • Real-time streaming: Use case requires quality scores

Comparison with Other Formats

Feature ZNA FASTA FASTQ CRAM FASTA.gz
Compression Excellent None None Excellent Good
Speed Fast Fastest Fast Slow Medium
Quality Scores
Paired-End
Random Access
Streaming Limited
Dependencies 1 0 0 Many 0

Python API

In addition to the CLI, ZNA provides a Python API:

from zna import ZnaHeader, ZnaWriter, ZnaReader, COMPRESSION_ZSTD

# Writing
header = ZnaHeader(
    read_group="Sample_01",
    compression_method=COMPRESSION_ZSTD,
    compression_level=5
)

with open("output.zzna", "wb") as f:
    with ZnaWriter(f, header) as writer:
        writer.write_record("ACGTACGT", is_paired=False, 
                          is_read1=False, is_read2=False)
        writer.write_record("TGCATGCA", is_paired=False,
                          is_read1=False, is_read2=False)

# Reading
with open("output.zzna", "rb") as f:
    reader = ZnaReader(f)
    print(f"Read Group: {reader.header.read_group}")
    
    for seq, is_paired, is_read1, is_read2 in reader.records():
        print(seq)

Development

Running Tests

# All tests
PYTHONPATH=src pytest -v

# Specific test suite
PYTHONPATH=src pytest tests/test_cli.py -v
PYTHONPATH=src pytest tests/test_core.py -v

# With coverage
PYTHONPATH=src pytest --cov=zna tests/

Code Quality

# Format code
black src/ tests/

# Type checking
mypy src/zna/

Limitations

  1. Sequences only: No quality scores, headers, or annotations
  2. Sequential access: No random access without full scan
  3. DNA/RNA only: A, C, G, T bases (N or IUPAC codes not supported)
  4. Case insensitive: Lowercase converted to uppercase
  5. No index: Full file scan required for record counting

Future Enhancements

  • Parallel compression/decompression
  • Optional index for random access
  • Support for IUPAC ambiguity codes
  • Memory-mapped I/O for large files
  • Streaming statistics (GC content, length distribution)

License

GNU General Public License v3.0 (GPLv3)


Citation

If you use ZNA in your research, please cite:

Iyer, M. (2026). ZNA: A compressed binary format for nucleic acid sequences.
GitHub: https://github.com/mkiyer/zna

Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch
  3. Add tests for new functionality
  4. Ensure all tests pass
  5. Submit a pull request

Contact

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

zna-0.1.2.tar.gz (65.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

zna-0.1.2-cp314-cp314-macosx_15_0_arm64.whl (66.4 kB view details)

Uploaded CPython 3.14macOS 15.0+ ARM64

File details

Details for the file zna-0.1.2.tar.gz.

File metadata

  • Download URL: zna-0.1.2.tar.gz
  • Upload date:
  • Size: 65.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for zna-0.1.2.tar.gz
Algorithm Hash digest
SHA256 2782d58fdcb90627d9b2808ca25c631088e80bfd1d7b0bde2cf6b916e933e334
MD5 6bd6552c36a62eb56a1ed155bda3266b
BLAKE2b-256 9dcf84c1e928287c9b37422968c1fddfdec452383fa89eb4eddbd0c48e70e8fc

See more details on using hashes here.

File details

Details for the file zna-0.1.2-cp314-cp314-macosx_15_0_arm64.whl.

File metadata

  • Download URL: zna-0.1.2-cp314-cp314-macosx_15_0_arm64.whl
  • Upload date:
  • Size: 66.4 kB
  • Tags: CPython 3.14, macOS 15.0+ ARM64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for zna-0.1.2-cp314-cp314-macosx_15_0_arm64.whl
Algorithm Hash digest
SHA256 b2497fa91510458221fd63fc15accc208f17c1f436d35d7e4e49f49cdea12350
MD5 6b1b4db31888b77fdfa7554472ffccef
BLAKE2b-256 3ae0d9acf88fa8b44d6cb7ca6de86150abf8336e6be06a8d39df04e73a51b324

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page