Skip to main content

ZNA: Compressed Nucleic Acid Format

ZNA (Compressed Z-Nucleic N-Acid A) is a high-performance binary format for storing DNA/RNA sequences with exceptional compression and I/O speed.

Performance

  • 1.7 GB/s decode, 726 MB/s encode on 150 bp reads (in-memory, single core)
  • 6.6 GB/s decode on long reads
  • ~4x compression from 2-bit packing alone, more with Zstd on duplicated data
  • C++ acceleration with a pure Python fallback

Features

  • High Compression: 2-bit encoding (4 bases per byte) + optional Zstd compression
  • Ultra-Fast I/O: C++ accelerated encode/decode with block-based architecture
  • Minimal Dependencies: zstandard only (C++ extension auto-compiled)
  • Flexible: Single-end, paired-end, and interleaved reads
  • Overlap Merging: zna merge collapses overlapping pairs into full-fragment reads on one calibrated likelihood-ratio score, with a compiled kernel and byte-identical output on any platform
  • Strand-Specific Support: dUTP, TruSeq, and custom strand protocols
  • Built-in Shuffle: Memory-bounded random shuffling for training data preparation
  • Metadata Rich: Read groups, descriptions, and custom flags
  • Unix-Friendly: Pipe-compatible CLI for seamless workflow integration
  • Streaming: Memory-efficient block-based processing

Installation

# From source (recommended - includes C++ acceleration)
git clone https://github.com/mkiyer/zna.git
cd zna
pip install -e .

# Check if C++ acceleration is available
python -c "from zna.core import is_accelerated; print(f'Accelerated: {is_accelerated()}')"

Requirements:

  • Python ≥3.10
  • C++ compiler (for optimal performance)
  • CMake ≥3.15 (auto-installed via pip)

Quick Start

# Encode FASTQ to compressed ZNA (default: Zstd level 9)
zna encode sample.fastq.gz -o sample.zna

# Encode with shuffle (for ML training data)
zna encode sample.fastq.gz --shuffle -o shuffled.zna

# Encode with shuffle and explicit memory cap per bucket
zna encode sample.fastq.gz --shuffle --shuffle-buffer-size 512M -o shuffled.zna

# Shuffle an existing ZNA file
zna shuffle input.zna -o shuffled.zna

# Decode back to FASTA
zna decode sample.zna -o sample.fasta

# Inspect file statistics
zna inspect sample.zna

# Overlap-merge paired-end reads before encoding
zna merge --in1 R1.fq.gz --in2 R2.fq.gz --out merged.fq.gz
zna encode --interleaved --treat-unpaired-as-merged merged.fq.gz -o sample.zna

# Pipe-friendly workflows
cat reads.fastq | zna encode -o reads.zna
zna decode reads.zna | head -n 1000

Performance Benchmarks

Throughput by Read Length

Measured on ZNA 0.3.5, Apple Silicon, single core, in-memory (BytesIO), min-of-7. Throughput is sequence bases in/out per second; compression is bases per stored byte at Zstd level 9 on random (worst-case, incompressible) sequence.

Read Type Encode (MB/s) Decode (MB/s) Encode (rec/s) Decode (rec/s) Compression
Short (Illumina, 150 bp) 726 1,718 4.8 M 11.5 M 3.95x
Medium (300 bp) 1,073 2,443 3.6 M 8.1 M 4.00x
Long (PacBio, 1 kb) 1,775 2,924 1.8 M 2.9 M 4.00x
Very Long (5 kb) 2,610 5,770 0.5 M 1.2 M 4.00x
Ultra Long (15 kb) 2,701 6,621 0.2 M 0.4 M 4.00x

Key Insights:

  • Throughput scales with read length: per-record overhead dominates at 150 bp and vanishes by 5 kb.
  • 4x is the 2-bit packing floor. Real libraries with duplicate reads compress further; unique reads do not, because packed DNA is near-incompressible.
  • blocks() is faster still for batch consumers — see Batch Reading.

See docs/PERFORMANCE.md for detailed benchmarking.


Documentation

This README is the user manual — installation, usage, the file format, the command reference and the Python API are all below.

CHANGELOG.md what changed in each release, and why
docs/METHODS.md the algorithms: the overlap score and its two thresholds, the quality-aware consensus, fragment geometry and what the flags mean, the codec
docs/MERGE_BENCHMARK_RESULTS.md zna merge scored against known ground truth and head to head with fastp. Read this before changing a threshold
docs/PERFORMANCE.md compression ratios, throughput, and tuning
docs/ROADMAP.md what is scheduled, what is being considered, and what was tried and closed by measurement
docs/RELEASING.md publishing to PyPI and Bioconda (maintainers)
docs/MERGE_PAIRS_PLAN.md zna encode --merge-pairs — specified, not built (0.5.0)
docs/NPOLICY_PLAN.md the --npolicy design and what remains of it
docs/HANDOFF_0.4.0.md what 0.4.0 shipped, what is next, and the build traps (maintainers)

File Format Specification

Overview

ZNA files use a binary format optimized for nucleic acid sequences:

  • File Extension: .zna (for both compressed and uncompressed files)
  • Default Compression: Zstd level 9 (use --uncompressed to disable)
  • Magic Number: ZNA\x1A (4 bytes)
  • Version: 2 (1 byte)
  • 2-bit Encoding: A=00, C=01, G=10, T=11
  • Block Structure: Columnar blocks, compressed as one Zstd frame each
  • Metadata: Read groups, descriptions, and custom information

File Structure

┌─────────────────────────────────────────┐
│  File Header (15 bytes fixed)           │
│   - Magic "ZNA\x1A" (4 bytes)           │
│   - Version = 2 (1 byte)                │
│   - Sequence length width (1 byte)      │
│   - Flags (1 byte)                      │
│   - Compression method (1 byte)         │
│   - Compression level (1 byte)          │
│   - Label count (2 bytes)               │
│   - Read-group / description lens (4 B) │
│  + read group, description (variable)   │
│  + one 89-byte definition per label     │
├─────────────────────────────────────────┤
│  Block 0                                │
│   Block Header (20 bytes)               │
│    * Compressed size    (4 bytes)       │
│    * Uncompressed size  (4 bytes)       │
│    * Record count       (4 bytes)       │
│    * Flags column size  (4 bytes)       │
│    * Lengths column size(4 bytes)       │
│   Payload — ONE Zstd frame, COLUMNAR:   │
│    ┌───────────────────────────────┐    │
│    │ flags    (1 byte  per record) │    │
│    │ labels   (per schema, if any) │    │
│    │ lengths  (1/2/4 B per record) │    │
│    │ sequences (2-bit packed)      │    │
│    └───────────────────────────────┘    │
├─────────────────────────────────────────┤
│  Block 1 ...                            │
└─────────────────────────────────────────┘

The payload is columnar, not row-oriented: all flags come first, then all label values, then all lengths, then all packed sequence. That is what lets zna inspect tally flags without touching sequence, and blocks(labels=False) skip label columns.

The file header stores no record or block count. Each block header carries its own count, so totals come from walking the block chain — see block_index().

Record Format

A record's fields live in separate columns of the block, not adjacent to each other. Per record:

  • Flags (1 byte): IS_READ1 (bit 0), IS_READ2 (bit 1), IS_PAIRED (bit 2), IS_RC (bit 3 — set when strand normalization reverse-complemented this record), IS_FULL_FRAGMENT (bit 4 — the record spans its whole fragment, so both edges are true fragment boundaries). Bits 5-7 are reserved.
  • Length (1-4 bytes): Sequence length (configurable)
  • Sequence (variable): 2-bit encoded bases

Compression

  • Method 0: Uncompressed
  • Method 1: Zstd, levels 1-22 (default 9)
  • Block Size: Default 4 MiB (--block-size, accepts K/M/G suffixes)

Both use the .zna extension; compression is recorded in the header, not the filename. Smaller blocks cost compression ratio only on duplicate-rich data — on unique reads the packed sequence is incompressible, so block size is free. A batch consumer holding a decoded block at a time (see blocks()) may want --block-size 1M to bound its memory.


Usage Guide

Encoding

Single-End Reads

# From FASTQ file
zna encode sample.fastq -o sample.zna

# From FASTA file  
zna encode sample.fasta -o sample.zna

# From gzipped input
zna encode sample.fastq.gz -o sample.zna

# Lower level = faster encode, larger file (default is 9)
zna encode sample.fastq --level 5 -o sample.zna

# Uncompressed (rarely needed)
zna encode sample.fastq --uncompressed -o sample.zna

# From stdin
cat sample.fastq | zna encode -o sample.zna

# Force format (when extension detection fails)
cat data.txt | zna encode --fastq -o sample.zna

Paired-End Reads

# Separate R1/R2 files
zna encode R1.fastq.gz R2.fastq.gz -o paired.zna

# Interleaved file (strict alternating R1/R2 pairs)
zna encode interleaved.fastq --interleaved -o paired.zna

# Interleaved from stdin
cat interleaved.fastq | zna encode --interleaved -o paired.zna

Mixed Paired-End and Single-End Reads (Interleaved)

The --interleaved mode intelligently detects both paired-end and single-end reads in the same file by analyzing read names. This is useful for output from tools like fastp that produce mixed merged (single) and unmerged (paired) reads.

How it works:

  • Reads with matching base names (e.g., read1/1 and read1/2) are paired
  • Reads without matching pairs are treated as single-end
  • Read names are used to determine pairing (not just alternating order)
# Mixed interleaved input (fastp output with merged + unmerged reads)
zna encode fastp_output.fastq --interleaved -o mixed.zna

# Example input structure:
#   @read1/1         →  paired with next read
#   @read1/2
#   @merged1         →  single-end (no pair)
#   @read2/1         →  paired with next read
#   @read2/2
#   @merged2         →  single-end (no pair)

Read name formats supported:

  • /1 and /2 suffixes: read1/1, read1/2
  • No suffix: treated as single-end unless next read has matching base name
  • Comments ignored: read1/1 merged_length:150 extracts read1/1

Strand normalization of merged/single reads: single-end reads (including merged reads with no mate) are treated as read1 for strand normalization. Under --strand-specific, a single read is reverse-complemented exactly when read1 is antisense, so merged reads end up on the same strand as normalized paired R1 reads.

Advanced Options

# Custom metadata
zna encode sample.fastq \
  --read-group "Sample_01" \
  --description "Experiment XYZ" \
  -o sample.zna

# Strand-specific library (default: R1 antisense, R2 sense)
zna encode R1.fastq.gz R2.fastq.gz \
  --strand-specific \
  -o stranded.zna

# Custom strand orientation (e.g., fr-secondstrand protocol)
zna encode R1.fastq.gz R2.fastq.gz \
  --strand-specific --read1-sense --read2-antisense \
  -o stranded.zna

# Handle sequences with N nucleotides
zna encode sample.fastq --npolicy trim3 -o clean.zna      # Cut each read at its first N (default)
zna encode sample.fastq --npolicy random -o clean.zna     # Substitute N from a seeded stream

# Shuffle during encoding (for ML training data preparation)
zna encode sample.fastq --shuffle -o shuffled.zna
zna encode R1.fastq.gz R2.fastq.gz --shuffle --seed 12345 -o shuffled.zna

# Control compression
zna encode sample.fastq \
  --level 9 \
  --block-size 262144 \
  -o sample.zna

# Uncompressed (rarely needed, for maximum I/O speed)
zna encode sample.fastq --uncompressed -o sample.zna

# Sequence length encoding (max sequence length)
zna encode sample.fastq \
  --seq-len-bytes 1 \  # Max 255 bp
  -o short_reads.zna

zna encode sample.fastq \
  --seq-len-bytes 2 \  # Max 65,535 bp (default)
  -o sample.zna

zna encode sample.fastq \
  --seq-len-bytes 4 \  # Max 4.2 billion bp
  -o long_reads.zna

Decoding

Basic Decoding

# To FASTA file
zna decode sample.zna -o output.fasta

# To gzipped FASTA
zna decode sample.zna -o output.fasta.gz

# To stdout (pipe-friendly)
zna decode sample.zna | head -n 1000

# From stdin
cat sample.zna | zna decode -o output.fasta

Paired-End Decoding

# Interleaved output (default)
zna decode paired.zna -o interleaved.fasta

# Split to R1/R2 files (use # placeholder)
zna decode paired.zna -o reads#.fasta
# Creates: reads_1.fasta and reads_2.fasta

# Split with gzip
zna decode paired.zna -o reads#.fasta.gz
# Creates: reads_1.fasta.gz and reads_2.fasta.gz

# Restore original strand for strand-specific libraries
zna decode stranded.zna --restore-strand -o reads.fasta

Piping Examples

# Extract first 1M reads
zna decode large.zna | head -n 2000000 > subset.fasta

# Count sequences
zna decode sample.zna | grep -c "^>"

# Convert to gzipped output via pipe
zna decode sample.zna --gzip > output.fasta.gz

# Chain operations
zna decode sample.zna | seqtk seq -r - | gzip > reversed.fasta.gz

Batch Reading with blocks()

records() yields one tuple per record. A consumer that works a whole batch at a time — a training data loader, say — can instead take a block at a time and skip the per-record tuple entirely:

from zna import ZnaReader, FLAG_FIELDS

with open("sample.zna", "rb") as fh:
    for sequences, flags in ZnaReader(fh).blocks():
        # sequences: list[str];  flags: bytes, one per record, same order
        for seq, fl in zip(sequences, flags):
            is_paired, is_read1, is_read2 = FLAG_FIELDS[fl]
            ...

ENDS_BY_FLAG[fl] gives (has_start, has_end) from the same byte — whether each edge of the stored sequence is a true fragment boundary. Use it rather than inferring from the mate number: under unstranded normalization ZNA reverse-complements one mate per pair at random, so the boundary edge is a per-record fact, not a property of R1 versus R2.

stride/offset shard by block, and — the point — seek past the blocks this shard does not want instead of decoding and discarding them:

# Worker 3 of 8: decodes ~1/8 of the file, not all of it.
for sequences, flags in ZnaReader(fh).blocks(stride=8, offset=3):
    ...

That is worth 1.8x at 2 workers and 9.4x at 16, compared with striding over records(). Two conditions come with it:

  • Record order must already be arbitrary. Shards get contiguous runs, not an interleave, so a file grouped by anything meaningful hands each worker a biased sample. Use zna shuffle first.
  • The file needs many more blocks than shards. Shares are whole blocks, so a small file split many ways is lopsided, and past the block count some shards get nothing — which blocks() warns about rather than passing off as an empty file. The default 4 MiB block gives a few hundred blocks per GB; write with a smaller block_size if you need finer shards.

blocks() also takes restore_strand=True.

On a labeled file it needs an explicit labels=, because quietly discarding label columns is not a decision it should make for you:

reader.blocks()                 # labeled file -> raises
reader.blocks(labels=False)     # skip the columns  -> (sequences, flags)
reader.blocks(labels=True)      # -> (sequences, flags, label_columns)

label_columns holds one value-tuple per column in header order, each as long as sequences; len(label_columns) always equals len(header.labels), so an unlabeled file yields (). On a three-column file, labels=False is 3.4x faster than records() and labels=True 2.1x. (Both still inflate the label bytes — a block is one zstd frame — what is saved is unpacking them into Python objects.)

Sizing a file before reading it: block_index()

The ZNA file header stores no record or block count — only the format version, sequence-length width, strand flags, compression settings and label schema. Each block header does carry its own record count, so the totals are recovered by walking the block chain, seeking over each payload:

reader = ZnaReader(fh)
index = reader.block_index()          # list[BlockInfo]
total = sum(b.n_records for b in index)

This decompresses nothing. Measured at 2.3 µs per block — 1.4 ms for a 38 MB, 611-block, 1M-record file, against 89 ms to reach the same counts by decoding. Cheap enough to run at open time, or across a whole corpus to build a manifest.

That makes proportional subsampling straightforward: use the counts to decide how much of each file you want, then decode only those blocks.

import random

index = reader.block_index()
want = round(len(index) * target_fraction)
keep = random.sample([b.index for b in index], want)

for sequences, flags in reader.blocks(indices=keep):
    ...

indices is mutually exclusive with stride/offset. Prefer it when the fraction is not a unit fraction, or when repeated passes over one file should see different blocks — stride admits only stride distinct phases, so training several epochs at stride=4 would revisit the same four subsets.

Blocks are flushed on an estimated byte size, so record counts per block are near-uniform for fixed-length reads and vary for variable-length ones. That is why block_index() returns per-block counts rather than an average, and why sampling k of n blocks gives approximately, not exactly, k/n of the records.

Cataloguing a corpus: zna inspect --json

zna inspect sample.zna --json
zna inspect sample.zna --json --blocks     # include the per-block array
zna inspect sample.zna --json --counts     # add per-flag record tallies

Emits header fields plus n_blocks and n_records, read from block headers without decompressing. Fast enough to sweep thousands of files, so a manifest can record record counts once and weight a balanced sample later without opening any of them.

Batching alone (without sharding) is worth about 24% for a loader doing real per-record work, and it fades with read length: ~24% at 150 bp, ~8% at 1 kb, and nothing measurable at 10 kb, where the sequence dominates the record overhead.

Inspecting Files

# Show file statistics
zna inspect sample.zna

Example Output:

File: sample.zna
Total Size: 45.32 MB

--- Header Metadata ---
Read Group:       Sample_01
Description:      Experiment XYZ
Seq Length:       2 bytes (Max: 65535 bp)
Strand Specific:  True
R1 Antisense:     True
R2 Antisense:     False
Compression:      ZSTD (Level 3)

--- Content Statistics ---
Total Blocks:       356
Total Records:      1000000
Compressed Payload: 42.15 MB
Uncompressed Data:  125.50 MB
Compression Ratio:  2.98x

Command Reference

zna encode

Convert FASTQ/FASTA to ZNA format.

Usage:

zna encode [FILE1] [FILE2] [OPTIONS]

Positional Arguments:
  FILE1 [FILE2]          Input files (0=stdin, 1=single/interleaved, 2=paired R1 R2)

Options:
  --interleaved          Treat input as interleaved (auto-detects mixed paired/single reads)
  --shuffle              Shuffle records after encoding (for ML training data)
  --seed N               Random seed for --shuffle (default: 42)
  --shuffle-buffer-size N
                         Max memory per bucket for encode --shuffle (default: 1G).
                         Accepts K/M/G suffixes.
  --fasta                Force FASTA format (overrides extension detection)
  --fastq                Force FASTQ format (overrides extension detection)

Metadata:
  --read-group TEXT      Read group ID (default: "Unknown")
  --description TEXT     Description string
  --strand-specific      Flag library as strand-specific (default: R1 antisense, R2 sense)
  --strand-normalize     Enable strand normalization (RC reads to consistent strand).
                         With --strand-specific: deterministic (antisense reads RC'd).
                         Without: random RC (for unstranded data).
  --read1-sense          Read 1 represents sense strand
  --read1-antisense      Read 1 represents antisense strand (default when --strand-specific)
  --read2-sense          Read 2 represents sense strand (default when --strand-specific)
  --read2-antisense      Read 2 represents antisense strand
  --npolicy {trim3,random}
                         Policy for handling 'N' nucleotides:
                         - drop: skip sequences containing N
                         - random: replace N with random base (A/C/G/T)
                         - A/C/G/T: replace N with specific base

Format Options:
  -o, --output FILE      Output file (default: stdout)
  --seq-len-bytes N      Bytes for sequence length: 1, 2, or 4 (default: 2)
  --block-size N         Block size in bytes (default: 131072)
  --zstd                 Force Zstd compression
  --uncompressed         Force uncompressed
  --level N              Zstd compression level 1-22 (default: 3)

zna decode

Convert ZNA to FASTA format.

Usage:

zna decode [FILE] [OPTIONS]

Positional Arguments:
  FILE                   Input ZNA file (default: stdin)

Options:
  -o, --output FILE      Output FASTA file. Use '#' for split R1/R2
  -q, --quiet            Suppress progress messages
  --gzip                 Force gzip compression for stdout
  --restore-strand       Restore original strand orientation for antisense reads

zna inspect

Display ZNA file statistics.

Usage:

zna inspect FILE [--counts]

  input FILE             Input ZNA file to inspect
  --counts               Also report per-flag record counts (paired R1, paired R2,
                         single/merged, reverse-complemented). Reads block payloads,
                         so slower than the default header-only scan.

zna shuffle

Randomly shuffle records in a ZNA file with bounded memory usage. Preserves paired-end read associations.

Usage:

zna shuffle INPUT -o OUTPUT [OPTIONS]

Positional Arguments:
  INPUT                  Input ZNA file to shuffle

Options:
  -o, --output FILE      Output ZNA file (required)
  -s, --seed N           Random seed for reproducibility (default: 42)
  -b, --buffer-size SIZE Maximum memory per bucket (default: 1G)
                         Accepts K/M/G suffixes (e.g., 512M, 2G)
  --block-size SIZE      Block size for output ZNA (default: 4M)
  --tmp-dir DIR          Directory for temporary files (default: system temp)
  -q, --quiet            Suppress progress messages

Algorithm: Uses bucket shuffle with bounded memory:

  1. Randomly distributes records into K temporary bucket files on disk
  2. Shuffles each bucket in memory using Fisher-Yates algorithm
  3. Concatenates shuffled buckets to produce uniform random permutation

Examples:

# Shuffle with default settings (1GB memory, seed 42)
zna shuffle input.zna -o shuffled.zna

# Shuffle with custom seed for reproducibility
zna shuffle input.zna -o shuffled.zna --seed 12345

# Shuffle with limited memory (512MB buffer)
zna shuffle input.zna -o shuffled.zna --buffer-size 512M

# Shuffle paired-end data (pairs stay together)
zna shuffle paired.zna -o shuffled_paired.zna

Note: Paired-end reads (R1+R2) are kept together as a single shuffle unit.

zna merge

Overlap-merge paired-end reads into one mixed interleaved FASTQ, ready for zna encode --interleaved. Replaces fastp's PE-merge step.

Each pair is scored once: R1 is slid against revcomp(R2) over the single axis of candidate fragment lengths, and every shift gets a log-likelihood ratio in bits+1.99 per matching base (that is log2 4, the information in agreeing on one of four bases), -6.23 per mismatch at a 1% error rate. Both weights fall out of the error rate; neither is tuned. The best-scoring shift (argmax, not fastp's first-accept) is then read at two thresholds:

condition action
merge score ≥ --threshold-merge emit one full-fragment record
trim --threshold-trim ≤ score < merge keep both; split the redundant overlap between their 3' ends
keep score < --threshold-trim keep both, untouched

Three parameters, all with units. Both thresholds read one calibrated scale, so T bits tolerates a spurious rate of about N · 2^-T over the N ≈ 2 · readlen candidate shifts — the default 28 is one spurious merge in 10⁶ pairs against chance alignment (measured: 0 in 40,000 uniform-random pairs, at every read length from 50 to 300). It is not a bound against real sequence, where reads share genuine homology and repeat content. Trim sits far lower only because a wrong trim deletes bases from a read tail while a wrong merge invents sequence.

The overlap sits at the 3' end of both mates — each read starts at a fragment end and reads inward — so a trim splits it between them. The emitted pair tiles the fragment exactly once and comes out at equal length, and where the mates disagree both carry the consensus call.

Choosing --threshold-merge, measured against ground truth

The defaults are not a guess. On 1,000,000 simulated pairs from hg38 with the true fragment length known exactly (docs/MERGE_BENCHMARK_RESULTS.md), against fastp 1.1.0 at its own defaults:

setting chimera rate¹ sensitivity² merges that are wrong reconstructed exactly³
--threshold-merge 28 (default) 1.231% 99.83% 0.96% 86.59%
--threshold-merge 60 0.597% 92.63% 0.55% 88.87%
--threshold-merge 100 0.245% 83.57% 0.29% 91.39%
fastp defaults 0.621% 92.98% 0.65% 85.90%

¹ fraction of pairs with no true overlap that were merged anyway — the false-positive rate. ² fraction of pairs with a true overlap ≥ 15 bases that merged. ³ merged records equal to the true fragment, base for base.

If you want fastp's false-positive rate, use --threshold-merge 60. That is not a coincidence: 60 bits is 31 clean bases, which is essentially fastp's --overlap_len_require 30. At that matched operating point zna's sensitivity is the same (92.63% vs 92.98%) and its reconstruction is better (88.87% vs 85.90% exact), because the overlap consensus recovers 90.4% of recoverable overlap errors against fastp's 74.1%.

For best overall accuracy, keep the default 28. It minimises false positives plus false negatives by a wide margin — 6,603 total errors per million pairs against 44,145 for the fastp-equivalent setting. Raising the threshold trades ~10.9 extra missed merges for every wrong merge it prevents at 28→34, worsening to 15.7 at 28→60, so it only pays if a chimera costs you more than ~11x what a missed merge does. A missed merge is not lost data: the pair is still emitted, correctly bounded and with its redundant overlap trimmed.

Tuning cannot reach zero. At 100 bits — 3.6x the default — 1,403 wrong merges per million remain. Every one is a fragment whose two ends are genuinely homologous (median 88% identity over 79 bases, hotspots entirely pericentromeric), and the scan never picks a lower-scoring alignment than the true one. That residue is a property of the genome, not of the threshold.

Usage:

zna merge --in1 R1.fq --in2 R2.fq --out OUT.fq [OPTIONS]

Required:
  --in1 FILE             R1 FASTQ (optionally .gz)
  --in2 FILE             R2 FASTQ (optionally .gz), positionally synced with --in1
  --out FILE             Output mixed interleaved FASTQ (.gz to gzip)

Options:
  --json FILE            Write run statistics as JSON (counts, histograms, provenance)
  --threshold-merge BITS Score at or above this merges the pair (default: 28.0)
  --threshold-trim BITS  Score at or above this (but below merge) trims R2 (default: 8.0)
  --min-read-length N    Drop emitted reads shorter than this (default: 40)
  --threads N            Merge worker threads (default: min(4, cpu count))
  --io-threads N         pigz threads for the gzipped output (default: 4)
  --chunk-size N         Read pairs per work unit (default: 2000)
  --compress-level N     pigz level for --out (default: 1 — it is an intermediate)
  --backend NAME         auto (default), accel, or python
  --no-sync-check        Skip the per-pair R1/R2 read-name consistency check
  --allow-empty          Exit 0 on an input with no read pairs
  -q, --quiet            Suppress progress logging

Speed. The merge kernel is compiled C++ and releases the GIL, so --threads are real worker threads. It is not usually the bottleneck — gzip is — so 2 threads saturate and more does nothing:

µs/pair
--threads 1 2.78
--threads 2 1.40
--threads 4 1.43

With gzip removed from both ends the tool runs at 0.42 µs/pair, so on compressed input it is I/O bound. pigz is used when it is on PATH, falling back to stdlib gzip.

Determinism. The score is computed in fixed-point integers and the argmax has a specified tie-break, so a given FASTQ produces byte-identical output on any platform, compiler and thread count. --backend python selects the pure-Python reference implementation, which exists as an oracle for the compiled one; it is ~50x slower and is never chosen for you.

Output is one stream mixing both shapes: merged reads as single records with the /1,/2 suffix stripped, unmerged pairs as adjacent /1,/2 records. Feed it to zna encode --interleaved --treat-unpaired-as-merged, which is exact here because merged records span their fragment and unmerged pairs are emitted all-or-nothing — never a lone mate.

Per-record provenance

The run summary says what happened to a library. These say what happened to a read. Existing header fields are always passed through untouched — provenance is appended, never substituted — so --label reads the same tags off an emitted record that it would have read off the input.

A record that nothing happened to is emitted unchanged, so on a clean library this costs nothing.

@SRR1.7  ZI:i:42 ZN:i:6 trim3_12 rescued_1 merged_90_0
          ^tag    ^bits  ^bases cut  ^no-calls   ^fastp-style,
          yours          by trim3    recovered    stays last
                                     from the mate
token meaning
trim3_<n> / subn_<n> bases removed by --npolicy trim3, or substituted by --npolicy random
rescued_<n> no-calls this record recovered from its mate, which cost nothing
merged_<n1>_<n2> bases contributed by R1 and R2 to a merged record

The word tokens are for reading; ZN:i:<bits> is the one that survives encoding. ZNA does not store headers, so it is the only per-record provenance that reaches a corpus:

zna encode --interleaved --treat-unpaired-as-merged \
           --label provenance:C:ZN -o reads.zna merged.fq.gz

That is an ordinary label column — declare it and you get one byte per record, omit it and nothing changes. There is no provenance-specific code in the encoder.

bit set when
1 trimmed the pair's redundant overlap was split between its mates
2 rescued ≥1 no-call was recovered from the mate
4 N-trimmed ≥1 base was removed by --npolicy trim3
8 N-substituted ≥1 base was invented by --npolicy random

There is deliberately no "merged" bit: that fact already has two homes, the merged_ token here and IS_FULL_FRAGMENT in the corpus. Every bit above is one with nowhere else to live — a trimmed pair in particular is emitted as an ordinary pair, and nothing in the ZNA flag byte distinguishes it from one kept whole.

Examples:

# Defaults suit 2x150 bp data; you normally set nothing
zna merge --in1 R1.fq.gz --in2 R2.fq.gz --out merged.fq.gz

# With run statistics for a pipeline to collect
zna merge --in1 R1.fq.gz --in2 R2.fq.gz --out merged.fq.gz --json merge.json

# Straight into a training corpus
zna merge --in1 R1.fq.gz --in2 R2.fq.gz --out merged.fq.gz
zna encode --interleaved --treat-unpaired-as-merged --strand-normalize \
           --shuffle merged.fq.gz -o reads.zna

Boundary guarantee. Base 0 of every emitted read is a true fragment boundary — nothing is ever removed from a read's 5' end, and trimming only ever cuts 3' ends. A merged record is the tool's assertion of its fragment. This is what makes IS_RC and IS_FULL_FRAGMENT honest for merged input; verified at 0 violations over 1,416,630 records against genome truth. See docs/METHODS.md for the derivation and docs/MERGE_BENCHMARK_RESULTS.md for the verification.


Performance Characteristics

Compression Ratios

Typical compression ratios compared to raw FASTQ:

Format Size Ratio Notes
FASTQ (uncompressed) 100% 1.0x Baseline
FASTQ.gz (gzip -6) 25-30% 3-4x Standard
ZNA (uncompressed) 12-15% 6-8x 2-bit encoding only
ZNA (Zstd L3) 8-10% 10-12x Fast compression (--level 3)
ZNA (Zstd L9) 6-8% 12-16x Default (DEFAULT_ZSTD_LEVEL = 9)

Results vary based on sequence complexity and redundancy

Speed

  • Encoding: ~4.8M reads/second at 150 bp (single thread)
  • Decoding: ~11.5M reads/second at 150 bp (single thread)
  • Block-based: shards and subsamples without decoding what it skips

Memory Usage

  • Streaming I/O: constant memory for records(); the writer buffers one block at a time
  • Default block size: 4 MiB (--block-size)
  • blocks() holds one decoded block per open reader, so block size sets the memory of a batch consumer: at 150 bp, a 4 MiB block is ~100k records (~20 MB of Python strings) against ~25k (~5 MB) for 1 MiB
  • No stored index required: block_index() walks block headers in ~2.3 µs per block, so counts and offsets are recovered without one

Technical Details

2-Bit Encoding

DNA bases are encoded in 2 bits:

A = 00 = 0
C = 01 = 1
G = 10 = 2
T = 11 = 3

Four bases pack into one byte:

Byte: [B1][B2][B3][B4]
      76 54 32 10  (bit positions)

Lookup Tables

Pre-computed lookup tables provide O(1) encoding/decoding:

  • Encoding: 256-element array mapping ASCII → 2-bit
  • Decoding: 256-element tuple mapping byte → 4-character string

Block-Based Architecture

Data is organized in independently compressed blocks:

  • Advantages: block-granular sharding and sampling; per-block counts without decompression
  • Overhead: 20 bytes of header per block
  • Choosing a size: 4 MiB (the default) maximises compression on duplicate-rich data. On unique reads, packed sequence is incompressible and block size costs nothing, so prefer smaller blocks (1 MiB) when a consumer reads with blocks() or shards by block

Compression Strategy

  • Zstd: Modern compression algorithm (Facebook)
  • Reusable compressor: Amortizes initialization cost
  • Memoryview parsing: Zero-copy decompression
  • Pre-sized buffers: Eliminates reallocations

Strand-Specific Libraries

ZNA supports strand-specific RNA-seq libraries by normalizing all reads to sense strand orientation during encoding. This enables consistent downstream analysis while preserving the ability to restore original strand information.

How It Works

  1. Encoding: reads are reverse-complemented into one common frame. Which reads, and by what rule, depends on the mode — see Strand Normalization below: with --strand-specific the protocol decides, without it one mate per pair is chosen at random.
  2. Storage: each record's IS_RC flag records whether it was flipped. That flag is the only record of it — it cannot be recovered from the sequence.
  3. Decoding: --restore-strand consumes IS_RC to recover the original orientation.

Strand Normalization

The --strand-normalize flag controls whether reads are reverse-complemented to a consistent strand during encoding:

  • With --strand-specific: Deterministic normalization — antisense reads are reverse-complemented to sense orientation based on the library protocol. Each read's IS_RC flag records whether it was flipped.
  • Without --strand-specific: Random reverse-complementing for unstranded data. Useful for data augmentation in ML training.
  • Without --strand-normalize: Reads are stored in their original orientation (no reverse-complementing).
# Strand-normalized encoding (most common for stranded RNA-seq)
zna encode R1.fq.gz R2.fq.gz --strand-specific --strand-normalize -o lib.zna

# Decode with original strand orientation restored
zna decode lib.zna --restore-strand -o original.fasta

# Decode with sense-normalized sequences (for alignment)
zna decode lib.zna -o normalized.fasta

Unstranded Normalization and Fragment Geometry

Unstranded normalization does more than augment the data: it carries information about the molecule that cannot be reconstructed afterwards.

A fastp-style FR pair covers the two ends of one fragment, pointing inward:

    fragment, length L
    |------------------------------------------------|
    |>>>>>>>>>>>|                        |<<<<<<<<<<<<|
     R1 as sequenced                      R2 as sequenced
     = F[0:l1]                            = revcomp(F[L-l2:L])

As sequenced the mates are in opposite frames. Normalization reverse-complements exactly one of them so both land in one common frame, and records which one in that record's IS_RC flag:

    common frame after normalization
    |------------------------------------------------|
    |<<<<<<<<<<<|                        |<<<<<<<<<<<<|
     not RC'd                             RC'd
     LEFT edge  = real fragment boundary  RIGHT edge = real fragment boundary
     right edge = read-length cutoff      left edge  = read-length cutoff

The invariant: whichever mate was reverse-complemented ends up at the right of the common frame, so its right edge is the real fragment boundary and its left edge is a read-length cutoff. For the other mate it is the mirror image.

IS_RC is the only thing that distinguishes the two cases, and it cannot be recovered from the sequence. Reverse-complementing the right-hand mate reproduces the fragment-frame sequence exactly, because that mate was stored reverse-complemented to begin with — there is no residue in the bases to test. The coin is also independent of the mate number, so is_read1 is not a substitute for it.

Reading the geometry. Use records(with_ends=True), which answers the question directly instead of making you re-derive it:

with open("lib.zna", "rb") as f:
    reader = ZnaReader(f)
    for seq, is_paired, is_read1, is_read2, has_start, has_end in \
            reader.records(with_ends=True):
        # has_start: the LEFT edge of seq is a true fragment boundary
        # has_end:   the RIGHT edge is
        ...

records(with_rc=True) exposes the raw IS_RC flag instead, if you want the orientation itself rather than the boundary geometry.

A record can have two real ends. When the insert is at or below the read length — every overlap-merged read, and any pair after adapter trimming — the record spans the whole fragment and both edges are true boundaries. IS_RC names only one edge, so that case is carried by a separate flag, IS_FULL_FRAGMENT, which with_ends folds in for you. Full-overlap pairs are detected automatically at encode time (mates covering the same interval are exact reverse complements); for unpaired records the encoder cannot tell a merged read from a genuine single-end read, so declare it:

# reads from an overlap merger: unpaired records span their whole fragment
zna encode --interleaved --treat-unpaired-as-merged -o out.zna merged.fq.gz

Without the flag an unpaired record is assumed to have one real edge, which is the safe reading — a tool marking fragment ends will under-label rather than place a marker at an interior position.

restore_strand=True is not a substitute: it consumes the flag to undo the reverse-complement and hand back original-orientation reads. A caller that wants the normalized frame and the boundary geometry needs with_rc, and the two options are mutually exclusive.

Normalization happens once, at encode time, and is not idempotent. Applying it a second time returns the data to an un-normalized state while the header still reports strand_normalized. So anything that copies records between ZNA files — zna encode on a .zna input, zna shuffle — copies the existing orientation rather than re-deriving it.

A view is for reading; the flag byte is for copying. records() returns views — each of them chosen for a consumer, and none of them able to carry the whole flag byte back to a writer. Copying uses copy_records():

# A lossless ZNA -> ZNA copy.
with open("in.zna", "rb") as fin, open("out.zna", "wb") as fout:
    reader = ZnaReader(fin)
    with ZnaWriter(fout, reader.header, preserve_normalization=True) as writer:
        for rec in reader.copy_records():
            writer.write_copy(rec)

copy_records() yields ZnaRecord(seq, flags, labels) — the stored ZnaRecordFlags byte, verbatim — so a copy carries every bit, including ones this version does not interpret. This example used to read records(with_ends=True) into write_records(), and that was wrong: (has_start, has_end) has three states where (IS_RC, IS_FULL_FRAGMENT) has four, so every full-fragment record came out of the copy with IS_RC cleared. write_records() now refuses that shape rather than accepting it.

Strand Flags

Flag Description
--strand-specific Enable strand-specific mode (default: R1 antisense, R2 sense)
--read1-sense Read 1 represents sense strand
--read1-antisense Read 1 represents antisense strand
--read2-sense Read 2 represents sense strand
--read2-antisense Read 2 represents antisense strand

Common Library Protocols

Protocol R1 R2 ZNA Flags
dUTP / TruSeq Stranded antisense sense --strand-specific (default)
Illumina Stranded mRNA antisense sense --strand-specific
fr-firststrand antisense sense --strand-specific
fr-secondstrand sense antisense --strand-specific --read1-sense --read2-antisense
Ligation (ScriptSeq) sense antisense --strand-specific --read1-sense --read2-antisense

Examples

# dUTP/TruSeq protocol (most common - this is the default)
zna encode R1.fastq.gz R2.fastq.gz --strand-specific -o library.zna

# fr-secondstrand protocol
zna encode R1.fastq.gz R2.fastq.gz \
  --strand-specific --read1-sense --read2-antisense \
  -o library.zna

# Decode with sense-normalized sequences (for alignment)
zna decode library.zna -o normalized.fasta

# Decode with original strand orientation restored
zna decode library.zna --restore-strand -o original.fasta

Per-Sequence Labels

ZNA can store numeric metadata as compact columnar label columns alongside each sequence. Labels are parsed from key-value tags in FASTQ headers (e.g. output from samtools fastq -T), but the tag format is not limited to SAM — any KEY:TYPE:VALUE field in the header will work, and keys can be any length.

Defining Labels on the CLI

# Two labels with descriptions
zna encode reads.fq.gz -o reads.zna \
  --label NH:C --label AS:i \
  --label-desc NH:"Number of hits" --label-desc AS:"Alignment score"

The --label format is NAME:TYPE where TYPE is one of: A, c, C, s, S, i, I, f, d, q, Q. Use the smallest type that fits your data to minimize file size.

Decoupled Name and Tag

By default, the label name (stored in the ZNA header) is also used as the tag to parse from input. You can decouple these with the 3-part format NAME:TYPE:TAG:

# Store as "edit_dist" in ZNA, but parse "NM" tag from input headers
zna encode reads.fq.gz -o reads.zna \
  --label edit_dist:C:NM --label aln_score:i:AS

# Custom long-form tags (not SAM) work too
zna encode reads.fq.gz -o reads.zna \
  --label score:i:alignment_score --label edits:C:edit_distance

The tag is only used at encode time and is not stored in the ZNA file. When decoding, the label name is used in the output.

Defining Labels with a YAML File

For many labels, define them in a YAML file instead of many CLI flags:

zna encode reads.fq.gz -o reads.zna --label-defs labels.yaml
# labels.yaml
labels:
  - name: NM
    type: C
    description: Edit distance
    missing: 255
  - name: aln_score
    type: i
    tag: AS                    # parse "AS" from input, store as "aln_score"
    description: Alignment score
    missing: -1

CLI flags --label and --label-desc override values from the YAML file, so you can keep a YAML base and tweak individual labels per run.

See examples/labels.yaml for a fully-commented template.

Decoding Labeled Files

# Include labels as SAM-style tags in the output
zna decode reads.zna --labels > output.fq

# Inspect to see label definitions
zna inspect reads.zna

Python API with Labels

from zna.core import ZnaHeader, ZnaWriter, ZnaReader
from zna.dtypes import LabelDef, parse_dtype

defs = (
    LabelDef(0, "NM", "Edit distance", parse_dtype("C"), missing=255),
    LabelDef(1, "AS", "Alignment score", parse_dtype("i"), missing=-1),
)
header = ZnaHeader(read_group="sample", labels=defs)

with open("out.zna", "wb") as f:
    with ZnaWriter(f, header) as w:
        w.write_record("ACGT", is_paired=False,
                        is_read1=False, is_read2=False,
                        labels=(3, 280))

with open("out.zna", "rb") as f:
    reader = ZnaReader(f)
    for seq, is_paired, is_r1, is_r2, labels in reader.records():
        print(seq, labels)  # ACGT (3, 280)

Labeled files yield a 5-tuple ending in labels. With with_rc=True the is_rc flag is inserted before it — (seq, is_paired, is_read1, is_read2, is_rc, labels) — so that the unlabeled and labeled tuples agree on where is_rc lives.


Use Cases

Recommended For

  • Long-term archival: High compression with fast retrieval
  • Data transfer: Reduced bandwidth requirements
  • Cloud storage: Lower storage costs
  • Pipeline integration: Unix-friendly streaming
  • Reference storage: Efficient genome/transcriptome storage

Not Recommended For

  • ⚠️ Record-level random access: there is no record index. Access is block granular — block_index() gives per-block offsets and counts without decompressing, and blocks(indices=...) reads only the blocks you ask for. That is what sharded training uses; seeking to record n is what is not supported.
  • Quality scores: Sequences only (use CRAM/BAM for qualities)
  • Small files: Overhead outweighs benefits (<10K reads)
  • Real-time streaming: Use case requires quality scores

Comparison with Other Formats

Feature ZNA FASTA FASTQ CRAM FASTA.gz
Compression Excellent None None Excellent Good
Speed Fast Fastest Fast Slow Medium
Quality Scores
Paired-End
Random Access
Streaming Limited
Dependencies 1 0 0 Many 0

Python API

In addition to the CLI, ZNA provides a Python API:

from zna import ZnaHeader, ZnaWriter, ZnaReader, COMPRESSION_ZSTD

# Writing
header = ZnaHeader(
    read_group="Sample_01",
    compression_method=COMPRESSION_ZSTD,
    compression_level=5
)

with open("output.zna", "wb") as f:
    with ZnaWriter(f, header) as writer:
        writer.write_record("ACGTACGT", is_paired=False, 
                          is_read1=False, is_read2=False)
        writer.write_record("TGCATGCA", is_paired=False,
                          is_read1=False, is_read2=False)

# Reading
with open("output.zna", "rb") as f:
    reader = ZnaReader(f)
    print(f"Read Group: {reader.header.read_group}")
    
    for seq, is_paired, is_read1, is_read2 in reader.records():
        print(seq)

# Copying to another ZNA file: carry the flag byte, not a view of it
with open("output.zna", "rb") as fin, open("copy.zna", "wb") as fout:
    reader = ZnaReader(fin)
    with ZnaWriter(fout, reader.header, preserve_normalization=True) as writer:
        for rec in reader.copy_records():
            writer.write_copy(rec)

records() yields a 4-tuple, or a 5-tuple ending in labels for labeled files. Two options change what it yields:

Option Yields Purpose
(default) (seq, is_paired, is_read1, is_read2) stored orientation
restore_strand=True same 4-tuple undoes strand normalization, returning original-orientation reads
with_rc=True (seq, is_paired, is_read1, is_read2, is_rc) stored orientation plus the per-record IS_RC flag
with_ends=True (seq, is_paired, is_read1, is_read2, has_start, has_end) which edges are true fragment boundaries

The options are mutually exclusive: restore_strand consumes the orientation, with_rc returns it raw, and with_ends returns what it means.

All three are views, for consumers. None of them round-trips back into a writer — with_ends in particular folds IS_RC and IS_FULL_FRAGMENT into two booleans that cannot distinguish all four reachable states. To copy records between ZNA files use copy_records() / write_copy(), which carry the flag byte itself. See Unstranded Normalization and Fragment Geometry for what is_rc means and why it cannot be derived from the sequence.


Development

Running Tests

# All tests
PYTHONPATH=src pytest -v

# Specific test suite
PYTHONPATH=src pytest tests/test_cli.py -v
PYTHONPATH=src pytest tests/test_core.py -v

# With coverage
PYTHONPATH=src pytest --cov=zna tests/

Code Quality

# Format code
black src/ tests/

# Type checking
mypy src/zna/

Limitations

  1. Sequences only: No quality scores, headers, or annotations
  2. Sequential access: No random access without full scan
  3. DNA/RNA only: A, C, G, T bases (N or IUPAC codes not supported)
  4. Case insensitive: Lowercase converted to uppercase
  5. No index: Full file scan required for record counting

Future Enhancements

See docs/ROADMAP.md — what is scheduled, what is under consideration, and what has already been tried and closed by measurement.


License

GNU General Public License v3.0 or later (GPL-3.0-or-later). See LICENSE.


Citation

If you use ZNA in your research, please cite:

Iyer, M. (2026). ZNA: A compressed binary format for nucleic acid sequences.
GitHub: https://github.com/mkiyer/zna

Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch
  3. Add tests for new functionality
  4. Ensure all tests pass
  5. Submit a pull request

Contact

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

zna-0.4.0.tar.gz (445.5 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

zna-0.4.0-cp314-cp314-win_amd64.whl (475.6 kB view details)

Uploaded CPython 3.14Windows x86-64

zna-0.4.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (330.7 kB view details)

Uploaded CPython 3.14manylinux: glibc 2.17+ x86-64

zna-0.4.0-cp314-cp314-macosx_11_0_x86_64.whl (278.3 kB view details)

Uploaded CPython 3.14macOS 11.0+ x86-64

zna-0.4.0-cp314-cp314-macosx_11_0_arm64.whl (271.6 kB view details)

Uploaded CPython 3.14macOS 11.0+ ARM64

zna-0.4.0-cp313-cp313-win_amd64.whl (465.3 kB view details)

Uploaded CPython 3.13Windows x86-64

zna-0.4.0-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (330.5 kB view details)

Uploaded CPython 3.13manylinux: glibc 2.17+ x86-64

zna-0.4.0-cp313-cp313-macosx_11_0_x86_64.whl (278.5 kB view details)

Uploaded CPython 3.13macOS 11.0+ x86-64

zna-0.4.0-cp313-cp313-macosx_11_0_arm64.whl (271.6 kB view details)

Uploaded CPython 3.13macOS 11.0+ ARM64

zna-0.4.0-cp312-cp312-win_amd64.whl (465.3 kB view details)

Uploaded CPython 3.12Windows x86-64

zna-0.4.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (330.5 kB view details)

Uploaded CPython 3.12manylinux: glibc 2.17+ x86-64

zna-0.4.0-cp312-cp312-macosx_11_0_x86_64.whl (278.5 kB view details)

Uploaded CPython 3.12macOS 11.0+ x86-64

zna-0.4.0-cp312-cp312-macosx_11_0_arm64.whl (271.5 kB view details)

Uploaded CPython 3.12macOS 11.0+ ARM64

zna-0.4.0-cp311-cp311-win_amd64.whl (466.3 kB view details)

Uploaded CPython 3.11Windows x86-64

zna-0.4.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (332.8 kB view details)

Uploaded CPython 3.11manylinux: glibc 2.17+ x86-64

zna-0.4.0-cp311-cp311-macosx_11_0_x86_64.whl (279.4 kB view details)

Uploaded CPython 3.11macOS 11.0+ x86-64

zna-0.4.0-cp311-cp311-macosx_11_0_arm64.whl (272.9 kB view details)

Uploaded CPython 3.11macOS 11.0+ ARM64

zna-0.4.0-cp310-cp310-win_amd64.whl (466.3 kB view details)

Uploaded CPython 3.10Windows x86-64

zna-0.4.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (332.8 kB view details)

Uploaded CPython 3.10manylinux: glibc 2.17+ x86-64

zna-0.4.0-cp310-cp310-macosx_11_0_x86_64.whl (279.3 kB view details)

Uploaded CPython 3.10macOS 11.0+ x86-64

zna-0.4.0-cp310-cp310-macosx_11_0_arm64.whl (273.0 kB view details)

Uploaded CPython 3.10macOS 11.0+ ARM64

File details

Details for the file zna-0.4.0.tar.gz.

File metadata

  • Download URL: zna-0.4.0.tar.gz
  • Upload date:
  • Size: 445.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0.tar.gz
Algorithm Hash digest
SHA256 778aa1d6ad3050289409e79eccb5a9f29da5a5de2167664d9689fc65db03411c
MD5 683deb9492f3c937b8a94851d702415f
BLAKE2b-256 f83e70a6b89f2527203df9b19c8e0f62a8a06723ec3532c8965d26122711a7bb

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0.tar.gz:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp314-cp314-win_amd64.whl.

File metadata

  • Download URL: zna-0.4.0-cp314-cp314-win_amd64.whl
  • Upload date:
  • Size: 475.6 kB
  • Tags: CPython 3.14, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp314-cp314-win_amd64.whl
Algorithm Hash digest
SHA256 4f0030572e23038318abcd0b31c8143536515af290678f1ac682f382451e5d72
MD5 28998a17cc916ef990223f7deb6f7ab7
BLAKE2b-256 3dfc7b46f17c00e4ad4f2453946f43197c808d20fdcc08334662ef0295693301

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp314-cp314-win_amd64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.

File metadata

File hashes

Hashes for zna-0.4.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl
Algorithm Hash digest
SHA256 0fc99be649bcea4e2022ca0451ede9a78883f8626ecb57bed1abb3d34a9ea494
MD5 88369701b76dd912cbb295a8a3da7406
BLAKE2b-256 3ccb2b15907facf1b2d65e8cee3b9e7c1d8e0667acb4ce41bc8d107c466b4e33

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp314-cp314-macosx_11_0_x86_64.whl.

File metadata

  • Download URL: zna-0.4.0-cp314-cp314-macosx_11_0_x86_64.whl
  • Upload date:
  • Size: 278.3 kB
  • Tags: CPython 3.14, macOS 11.0+ x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp314-cp314-macosx_11_0_x86_64.whl
Algorithm Hash digest
SHA256 d2c16ed35f68ec6adc71573212e7106efeecc4ce326bb13de6c874ef64bd75d8
MD5 04126c436d248b5f19d214cf742363fe
BLAKE2b-256 f008998fb01c0c1c4b5bc2f92ba4ac56b62ddbb94aca4fab40cb918fe3500e7a

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp314-cp314-macosx_11_0_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp314-cp314-macosx_11_0_arm64.whl.

File metadata

  • Download URL: zna-0.4.0-cp314-cp314-macosx_11_0_arm64.whl
  • Upload date:
  • Size: 271.6 kB
  • Tags: CPython 3.14, macOS 11.0+ ARM64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp314-cp314-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 674017d8a6a83a4cda092dd9f412c2d065e792d4d4511a3284f8e6b328a4fc58
MD5 c80e7d54d067becb16b891fa439f93a7
BLAKE2b-256 04e80519c80f0dbae6a771faa2bd7a5401aeeaff56cc30cdd55431b28a9bfc4f

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp314-cp314-macosx_11_0_arm64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp313-cp313-win_amd64.whl.

File metadata

  • Download URL: zna-0.4.0-cp313-cp313-win_amd64.whl
  • Upload date:
  • Size: 465.3 kB
  • Tags: CPython 3.13, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp313-cp313-win_amd64.whl
Algorithm Hash digest
SHA256 7145fb7ab2d3fc524178f73058c7db11edd872161dcef6acd8df7c5dd559dfa6
MD5 d682a7403a2174d609c6a9dd2bc35c0e
BLAKE2b-256 6a543952e1315965710885b10ba205c87eab44e8265566f23858c92a70738ef0

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp313-cp313-win_amd64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.

File metadata

File hashes

Hashes for zna-0.4.0-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.whl
Algorithm Hash digest
SHA256 5cdf571f389e880783749795469444ab59b5f2c10055d8b0111f905b6eec9bcc
MD5 6ce6fc5a30b05964d41a8aab21bacff7
BLAKE2b-256 3565bbddcc0a977d5207661c8b955857b6605457598b59e3ed6b236e4055fe69

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp313-cp313-macosx_11_0_x86_64.whl.

File metadata

  • Download URL: zna-0.4.0-cp313-cp313-macosx_11_0_x86_64.whl
  • Upload date:
  • Size: 278.5 kB
  • Tags: CPython 3.13, macOS 11.0+ x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp313-cp313-macosx_11_0_x86_64.whl
Algorithm Hash digest
SHA256 333faa596ed17f98df2062c3902a19e750410256986de3b01337d72626a6eef1
MD5 be02310d88486e77a986f4f5e7f67a02
BLAKE2b-256 03690c891a495bfd852fc2d3e6c63e76a3b8252233ac5a2312f202c152386f62

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp313-cp313-macosx_11_0_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp313-cp313-macosx_11_0_arm64.whl.

File metadata

  • Download URL: zna-0.4.0-cp313-cp313-macosx_11_0_arm64.whl
  • Upload date:
  • Size: 271.6 kB
  • Tags: CPython 3.13, macOS 11.0+ ARM64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp313-cp313-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 fb125754b2113089815aff3f7268b2ffe28aadb4ab01bbae94bf15b32cec8589
MD5 7087fc9315a41bee281aed971c3a244e
BLAKE2b-256 d0ac5a8f8c1536592f602d1b7cfab50c09be8a0db999d101eb29244438c89830

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp313-cp313-macosx_11_0_arm64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp312-cp312-win_amd64.whl.

File metadata

  • Download URL: zna-0.4.0-cp312-cp312-win_amd64.whl
  • Upload date:
  • Size: 465.3 kB
  • Tags: CPython 3.12, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp312-cp312-win_amd64.whl
Algorithm Hash digest
SHA256 8b4b69949cce8fd0e28eb6483f1b4257fddc1d668776117941e0d81d4c9e94df
MD5 efe35e8a9561f10af192513997db40d2
BLAKE2b-256 f2c9ae9bafa114afcd2c7916f8786dbe8b37300c7ce30dc5d956140f0de04151

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp312-cp312-win_amd64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.

File metadata

File hashes

Hashes for zna-0.4.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl
Algorithm Hash digest
SHA256 744fc559131103aeaa5ec02091f6955220013063f0113cfa6f3a32ba10611f93
MD5 0fab75e557e91a17561d6e7dabef1815
BLAKE2b-256 4ebafd79e11ea11f8259a58d5c9c3ca46ea6cb72843d1deaa751b16e05e2247c

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp312-cp312-macosx_11_0_x86_64.whl.

File metadata

  • Download URL: zna-0.4.0-cp312-cp312-macosx_11_0_x86_64.whl
  • Upload date:
  • Size: 278.5 kB
  • Tags: CPython 3.12, macOS 11.0+ x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp312-cp312-macosx_11_0_x86_64.whl
Algorithm Hash digest
SHA256 8ee42df0da86115373ed3f9ca527083cdc563a192cc198ae5fb1dae2563b1cd6
MD5 55f067ea23666ee1eb91851160661cdb
BLAKE2b-256 5f5720f36992d686d20c61972b8418bd2f88ace6ff45d504d6bdab6b9595af67

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp312-cp312-macosx_11_0_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp312-cp312-macosx_11_0_arm64.whl.

File metadata

  • Download URL: zna-0.4.0-cp312-cp312-macosx_11_0_arm64.whl
  • Upload date:
  • Size: 271.5 kB
  • Tags: CPython 3.12, macOS 11.0+ ARM64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp312-cp312-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 fc8471d7e2c64d547fd118fb5e68c8562d36604bab57b9b28e83ac94d8037cf5
MD5 64ac727ef5d53577d43cc331e8f241bf
BLAKE2b-256 51d2656bc0a66be2d37ae0b3f9c560b23b23869246735efad02604b572d9c412

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp312-cp312-macosx_11_0_arm64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp311-cp311-win_amd64.whl.

File metadata

  • Download URL: zna-0.4.0-cp311-cp311-win_amd64.whl
  • Upload date:
  • Size: 466.3 kB
  • Tags: CPython 3.11, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp311-cp311-win_amd64.whl
Algorithm Hash digest
SHA256 b5bbe3011694464aa07a45aa7b445113c4fec5c796c65ba0e0d93ec207a882d7
MD5 24e44a6825eeb04e40e0c4432381cf7b
BLAKE2b-256 f82b11104815e062705fbcd313d9c3cfbff4d75a23d8993b62f4e37843a4507b

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp311-cp311-win_amd64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.

File metadata

File hashes

Hashes for zna-0.4.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl
Algorithm Hash digest
SHA256 2d302895a99d8a9140ce5e3267f30f89a96d6b8ecb4505ce76a7cb773a3b1047
MD5 5eb8cec1fa20c72f5bb20957eb4f5804
BLAKE2b-256 30ff95bb6a2a5299ef912164ff1f3fbc6cd5f2521e78ef72cfec89976ca7ba52

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp311-cp311-macosx_11_0_x86_64.whl.

File metadata

  • Download URL: zna-0.4.0-cp311-cp311-macosx_11_0_x86_64.whl
  • Upload date:
  • Size: 279.4 kB
  • Tags: CPython 3.11, macOS 11.0+ x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp311-cp311-macosx_11_0_x86_64.whl
Algorithm Hash digest
SHA256 0532b049d2738fae8a607cc83c54a2743eab8233d150a996d6c86dd8ff8ab045
MD5 81dfb89729fca6504e57aafb2c53fc29
BLAKE2b-256 5beaf50e2ff32b329698edcfd727899376c71da45bb31f2ed425f77a5edde0ea

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp311-cp311-macosx_11_0_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp311-cp311-macosx_11_0_arm64.whl.

File metadata

  • Download URL: zna-0.4.0-cp311-cp311-macosx_11_0_arm64.whl
  • Upload date:
  • Size: 272.9 kB
  • Tags: CPython 3.11, macOS 11.0+ ARM64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp311-cp311-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 0d7a77476cc94d7f5132e65b45f957df164dea8907e6dbafded6722a2606ab2f
MD5 3b63387aa25d3274853d4c20fbe8ee40
BLAKE2b-256 eea2efe4a6c8c97e5fc4b3e5439f7bb9753b3e110891ab91a5642cab7df8af6e

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp311-cp311-macosx_11_0_arm64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp310-cp310-win_amd64.whl.

File metadata

  • Download URL: zna-0.4.0-cp310-cp310-win_amd64.whl
  • Upload date:
  • Size: 466.3 kB
  • Tags: CPython 3.10, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp310-cp310-win_amd64.whl
Algorithm Hash digest
SHA256 d3d7ad84b8de9306503f42aa333d933e7202bac243c185d05182f49164c51b70
MD5 e6dc28207130c96c8303d8eb4c7d79b6
BLAKE2b-256 466b293b2596fa580f8c7390924b07e380f81046b9dab74d3092d438461a429c

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp310-cp310-win_amd64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.

File metadata

File hashes

Hashes for zna-0.4.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl
Algorithm Hash digest
SHA256 02eaedfed6a31b5833087176d31bd0873b1b541527984e4a9610676bd150312c
MD5 b366530ac4adeddf32793e6b4eccbade
BLAKE2b-256 b5f690c5aa93202f2d549f397636682b077e5a87581d900e4205e057e8771ead

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp310-cp310-macosx_11_0_x86_64.whl.

File metadata

  • Download URL: zna-0.4.0-cp310-cp310-macosx_11_0_x86_64.whl
  • Upload date:
  • Size: 279.3 kB
  • Tags: CPython 3.10, macOS 11.0+ x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp310-cp310-macosx_11_0_x86_64.whl
Algorithm Hash digest
SHA256 52daf48ea1fcb02e2b88a16f3eb107bb603af4f9930e28b3098cac97c42a357e
MD5 0285e1721ab4c6fbfc3a93b4566fcf6e
BLAKE2b-256 6fe73d6119c5c8ccc603b336734033a8bc93b8cd1459a59ec488f3082c953179

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp310-cp310-macosx_11_0_x86_64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file zna-0.4.0-cp310-cp310-macosx_11_0_arm64.whl.

File metadata

  • Download URL: zna-0.4.0-cp310-cp310-macosx_11_0_arm64.whl
  • Upload date:
  • Size: 273.0 kB
  • Tags: CPython 3.10, macOS 11.0+ ARM64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for zna-0.4.0-cp310-cp310-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 2c95f619f6707a090a991445376ebd4e6688f7ae072138b189cf8719ec907679
MD5 cea17293b1db4910ac1f61a5b78126c0
BLAKE2b-256 aa4848acc75b5b21b225bf8c2ef4796ee656b9a054177281c429b80f6b774788

See more details on using hashes here.

Provenance

The following attestation bundles were made for zna-0.4.0-cp310-cp310-macosx_11_0_arm64.whl:

Publisher: publish.yml on mkiyer/zna

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.5.3

21 files

0.5.1

21 files

0.5.0

21 files

0.4.1

21 files

This release

0.4.0 This release

21 files

0.3.5

21 files

0.3.4

17 files

0.3.3

17 files

0.3.1

17 files

0.3.0

17 files

0.2.0

17 files

0.1.8

17 files

0.1.2

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page