Skip to main content

HGVS variant normalizer - Python bindings for the ferro bioinformatics toolkit

Project description

CI Nightly reference-aware tests Codecov Crates.io Documentation License install with bioconda DOI

ferro-hgvs

A high-performance HGVS variant nomenclature parser and normalizer written in Rust.

WARNING: ALPHA SOFTWARE - USE AT YOUR OWN RISK

This software is currently in ALPHA. While we have extensively tested it across a wide variety of HGVS patterns, no guarantees are made regarding correctness or stability.

Fulcrum Genomics

Features

  • Full HGVS Parsing: All coordinate systems (g/c/n/r/p/m/o) and edit types
  • Variant Normalization: 3'/5' shifting per HGVS specification
  • High Performance: ~5M variants/sec single-threaded parsing (>12M/s parallel), zero-copy with nom
  • Type-Safe: Leverages Rust's type system for correctness

Installation

Python

pip install ferro-hgvs

Pre-built wheels are available for Linux (x86_64, aarch64), macOS (x86_64, Apple Silicon), and Windows (x86_64) on Python 3.10+.

Rust

Add to your Cargo.toml:

[dependencies]
ferro-hgvs = "0.1"

Or install the CLI:

cargo install ferro-hgvs

Quick Start

CLI

# Parse a variant
ferro parse "NM_000088.3:c.459A>G"

# Parse from file
ferro parse -i variants.txt -f json

# Prepare reference data (downloads RefSeq, genome, cdot — RefSeq-only by default)
ferro prepare --output-dir ferro-reference

# Verify reference data is ready
ferro check --reference ferro-reference

# (Optional) pre-build the on-disk cdot cache as a setup step, so the one-time
# cache build doesn't slow the start of a real (or timed/benchmarked) run.
ferro check --reference ferro-reference --build-cache

# Normalize with reference
ferro normalize "NM_000088.3:c.459del" --reference ferro-reference/

Throughput tip: when normalizing many variants, feed them sorted by transcript accession (or by genomic position). ferro caches each resolved transcript, so consecutive variants on the same transcript skip the (dominant) cost of re-reading and re-building it from the reference. Sorted input keeps the relevant transcripts resident in the cache and is markedly faster on large batches — see Performance Comparison.

Optional reference data

A bare ferro prepare builds a RefSeq-only reference (accessions NM_/NR_/NP_/NG_). Two opt-in flags provision additional data — pass them at prepare time; they are what a fully-provisioned ("blessed") reference is built with:

# Add Ensembl support (accessions ENST/ENSG/ENSP). Downloads the Ensembl cdot
# metadata and cDNA FASTAs (~1 GB+); off by default. Without it, an ENST/ENSG/ENSP
# input reports "Reference not found" and the message points back at this flag.
ferro prepare --output-dir ferro-reference --ensembl

# Derive version-independent NG_ placements and the NG_→transcript-version map
# (ng_hosted_transcripts) for a curated list of RefSeqGene accessions. Required to
# resolve legacy gene-symbol selectors (NG_(GENE):c.…) and bare-NG_ hosted lookups.
ferro prepare --output-dir ferro-reference \
  --derive-ng-placements path/to/ng_accessions.txt

# A fully-provisioned reference combines both in one run:
ferro prepare --output-dir ferro-reference --ensembl \
  --derive-ng-placements path/to/ng_accessions.txt

Both flags are incremental: re-running ferro prepare over an existing reference adds the requested data and preserves already-provisioned artifacts.

Library

use ferro_hgvs::{parse_hgvs, HgvsVariant};

fn main() -> Result<(), ferro_hgvs::FerroError> {
    let variant = parse_hgvs("NM_000088.3:c.459A>G")?;

    match &variant {
        HgvsVariant::Cds(v) => println!("CDS variant: {}", v),
        HgvsVariant::Genome(v) => println!("Genomic variant: {}", v),
        _ => println!("Other: {}", variant),
    }

    Ok(())
}

Python

import ferro_hgvs

# Parse a variant
variant = ferro_hgvs.parse("NM_000088.3:c.459A>G")
print(variant.variant_type)  # "coding"
print(variant.reference)     # "NM_000088.3"
print(str(variant))          # "NM_000088.3:c.459A>G"

# Normalize with reference data
normalizer = ferro_hgvs.Normalizer(reference_json="ferro-reference/cdot.json")
normalized = normalizer.normalize("NM_000088.3:c.459del")

Supported HGVS Syntax

Type Prefix Example
Genomic g. NC_000001.11:g.12345A>G
Coding DNA c. NM_000088.3:c.459A>G
Non-coding n. NR_000001.1:n.100A>G
RNA r. NM_000088.3:r.459a>g
Protein p. NP_000079.2:p.Val600Glu
Mitochondrial m. NC_012920.1:m.3243A>G

Edit Types

  • Substitution: A>G, Val600Glu
  • Deletion: del, 100_200del
  • Insertion: 100_101insATG
  • Deletion-Insertion: 100_102delinsATG
  • Duplication: 100_102dup
  • Inversion: 100_200inv
  • Repeat: 100CAG[20]

CLI Commands

The ferro CLI provides commands beyond parsing and normalization:

Command Description
prepare Download and prepare reference data for normalization
check Verify reference data setup
parse Parse and validate HGVS variants
normalize Normalize HGVS variants (3'/5' shifting)
explain Explain error/warning codes (e.g., ferro explain W1001)
annotate-vcf Annotate VCF files with HGVS notation
vcf-to-hgvs Convert VCF records to HGVS
hgvs-to-vcf Convert HGVS to VCF format
liftover Liftover coordinates between genome builds
describe Generate HGVS from reference/observed sequences
effect Predict protein effect from variant
backtranslate Reverse translate protein to DNA variants
convert-gff Convert GFF3/GTF to transcripts.json
generate Generate HGVS descriptions from components
extract-hgvs Extract HGVS from VEP-annotated VCFs

Error Handling

ferro-hgvs provides configurable error handling with three modes:

Mode Behavior
strict Reject non-conformant input (default)
lenient Auto-correct with warnings
silent Auto-correct silently
# Use lenient mode to auto-correct common issues
ferro parse --error-mode lenient "p.val600glu"  # Corrects to p.Val600Glu

# Ignore specific warnings
ferro parse --ignore W1001,W2001 "p.val600glu"

# Get help on any error/warning code
ferro explain W1001
ferro explain --list

Configuration File

Create .ferro.toml in your project directory:

[error-handling]
mode = "lenient"
ignore = ["W1001", "W2001"]  # Silently correct these
reject = ["W3003"]           # Always reject these

Why ferro-hgvs?

ferro-hgvs provides the most comprehensive HGVS variant normalization across all pattern types, with performance orders of magnitude faster than alternatives.

Normalization Capabilities Comparison

Pattern Type ferro mutalyzer biocommons hgvs-rs
Genomic (g.)
Coding (c.) exonic
Coding (c.) intronic ✓**
Non-coding (n.)
RNA (r.)
Protein (p.) Net*

* mutalyzer protein normalization requires network access for NP_→NM_ lookups (cannot be cached locally). ** mutalyzer intronic support is enabled by default via genomic-context rewriting; disable with --no-rewrite-intronic.

Performance Comparison

All tools are benchmarked in ferro's offline configuration — best case for every tool. Reference data is preloaded locally (a local UTA database and SeqRepo) and the network is disabled, so the figures below measure parse/normalize compute, not I/O. Out of the box, hgvs-rs, biocommons/hgvs, and mutalyzer resolve each variant against a remote UTA/SeqRepo or the Mutalyzer web API — a network round-trip per variant (~100–1000 ms), i.e. roughly 1–10 variants/sec, hundreds to thousands of times slower than shown here (an order-of-magnitude estimate from per-call network latency, not separately benchmarked). That local, offline setup is exactly what ferro's prepare command builds; ferro needs no external service.

Median patterns/sec over 5 reps on an Apple M2 Max, local/offline. All tools draw from one stratified ClinVar population; per-tool sample sizes are calibrated so each tool is measured over a meaningful interval — fast cells (e.g. ferro/hgvs-rs parse) draw from millions of patterns, while slower cells (e.g. the per-tool normalize columns) draw from as few as tens to thousands. All tools exclude process/interpreter startup from the timed region — the mutalyzer/biocommons Python subprocesses are timed by their own internal startup-excluded timer, matching ferro/hgvs-rs. Only ferro parallelizes natively (rayon); the other tools are single-threaded libraries, so their normalize @8 workers figures come from the benchmark harness running 8 independent instances in parallel, while parsing is not sharded for them — hence the single-threaded label in their parse @8 workers column (mutalyzer normalize likewise shows no gain at 8 workers: per-call cache and IPC overhead dominate, so sharding does not help). Every tool runs fully offline against local reference data — a local UTA database and SeqRepo, with mutalyzer's network lookups disabled — the configuration ferro's prepare command enables; the figures therefore reflect compute throughput, not per-variant network latency. Reference-data load is excluded for all tools. ferro full-population peak: parse 20.0M/s, normalize 77.0k/s. See docs/BENCHMARK_RUNBOOK.md for the full method.

Parse

Tool Throughput @ 1 worker Throughput @ 8 workers ferro speedup @ 8w
ferro 5.1M/s 12.2M/s
mutalyzer 352/s single-threaded 35,000×
biocommons 3.9k/s single-threaded 3,100×
hgvs-rs 3.6M/s single-threaded

Normalize

Tool Throughput @ 1 worker Throughput @ 8 workers ferro speedup @ 8w
ferro 78.1k/s 260.2k/s
mutalyzer 4/s 4/s 73,000×
biocommons 368/s 818/s 320×
hgvs-rs 195/s 1.3k/s 200×

ferro thread scaling

Threads 1 2 4 8
ferro parse 5.1M/s 9.4M/s 16.0M/s 12.0M/s

Input ordering matters for batch throughput. Resolving a transcript (reading its full sequence from the reference and rebuilding its CDS/exon metadata) dominates per-variant cost. ferro memoizes resolved transcripts in a bounded in-memory cache, so repeated lookups of the same transcript are near-free. Providing variants sorted by transcript accession — or by genomic position, which clusters variants onto the same transcripts — maximizes the cache hit rate and can speed up large batches by an order of magnitude versus randomly-ordered input. Ordering matters most when the number of distinct transcripts in the run exceeds the cache capacity (very large or genome-wide inputs); below that, the working set stays resident regardless of order.

Reference Data: What ferro Prepares

The ferro prepare command downloads and organizes all reference data needed for comprehensive normalization. This data is then shared with other tools (mutalyzer, biocommons, hgvs-rs) to enable their local operation.

Data Type Source Size Enables
RefSeq transcripts NCBI ~1GB NM_/NR_/XM_ normalization
cdot metadata MANE ~200MB Transcript-to-genome mappings
GRCh38 + GRCh37 genomes NCBI ~4GB NC_ genomic normalization
RefSeqGene (sequences + genome alignments) NCBI ~600MB NG_ gene-region normalization; projecting c./n. variants into an NG_ parent's own g. frame (via the RefSeqGene→genome alignment GFF3)
LRG sequences + XML EBI ~50MB LRG_ stable-reference normalization; projecting c./n. variants into an LRG_ parent's own g. frame (via the LRG XML genomic mapping)
Protein sequences Derived from CDS ~200MB NP_/XP_ protein normalization
Legacy transcript versions NCBI ~50MB Historical ClinVar variants

Key insight: Without ferro's reference preparation, other tools require network access for each variant lookup (adding 100-1000ms latency per variant). With ferro's cached reference data, all tools can operate fully offline with consistent, reproducible results.

Deriving version-independent NG_ placements (#728)

ferro prepare --derive-ng-placements <accessions.txt> derives genomic placements for the listed NG_ versions (one exact accession per line, e.g. NG_012337.3; blank lines and # comments ignored), writing derived_refseqgene_placements.json into the reference directory and wiring the manifest's derived_refseqgene_placements field. This fills version gaps the archived RefSeqGene→genome GFF3 snapshots do not cover. It needs cdot + the genome in the same prepare run and uses NCBI EFetch per accession; accessions that cannot be validated are skipped with a warning. The field is preserved across subsequent prepare runs.

Benchmark: Reference Data & Tool Comparison

The main ferro binary includes commands to prepare reference data (ferro prepare) and check its status (ferro check). The ferro-benchmark tool (build with --features benchmark) extends this for tool comparison benchmarks.

Command Description
prepare <tool> Prepare reference data for a tool
check <tool> Verify tool configuration and dependencies
parse <tool> Parse HGVS patterns with specified tool
normalize <tool> Normalize HGVS patterns with specified tool
compare results Compare parse/normalize results between tools
extract Extract patterns from ClinVar, VCFs, or create samples
setup Set up UTA database, SeqRepo, and other services
generate Generate summary reports and configs
collate Aggregate sharded results

Quick Start

# Prepare ferro reference (main binary - no special features needed)
ferro prepare --output-dir data/ferro

# Check reference data
ferro check --reference data/ferro

# Normalize with ferro
ferro normalize -i patterns.txt --reference data/ferro

# For tool comparison, build with benchmark support
cargo build --release --features benchmark

# Prepare other tools (uses ferro reference for transcript data)
ferro-benchmark prepare mutalyzer --ferro-reference data/ferro --output-dir data/mutalyzer
ferro-benchmark prepare biocommons --seqrepo-dir data/seqrepo --uta-dump uta_20210129b.pgd.gz --ferro-reference data/ferro

# Compare results between tools
ferro-benchmark normalize mutalyzer -i patterns.txt -o mutalyzer.json --mutalyzer-settings data/mutalyzer/mutalyzer_settings.conf
ferro-benchmark compare results normalize ferro.json mutalyzer.json -o comparison.json

Supported tools: ferro-hgvs, mutalyzer, biocommons/hgvs, hgvs-rs

Note: The pixi.toml and pixi.lock files in this repository define a pixi environment for the Python-based external tools (mutalyzer, biocommons/hgvs, seqrepo) used in benchmarking. Run pixi shell to activate it.

See docs/BENCHMARK_GUIDE.md for detailed usage.

Development

cargo build
cargo test
cargo clippy -- -D warnings

License

Licensed under the MIT License. See LICENSE for details.

Disclaimer

This software is under active development. While we make a best effort to test this software and to fix issues as they are reported, this software is provided as-is without any warranty (see the license for details). Please submit an issue, and better yet a pull request as well, if you discover a bug or identify a missing feature. Please contact Fulcrum Genomics if you are considering using this software or are interested in sponsoring its development.

Contributing

See CONTRIBUTING.md for guidelines.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ferro_hgvs-0.7.1.tar.gz (3.0 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

ferro_hgvs-0.7.1-cp310-abi3-win_amd64.whl (3.3 MB view details)

Uploaded CPython 3.10+Windows x86-64

ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_x86_64.whl (3.6 MB view details)

Uploaded CPython 3.10+musllinux: musl 1.2+ x86-64

ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_aarch64.whl (3.4 MB view details)

Uploaded CPython 3.10+musllinux: musl 1.2+ ARM64

ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.4 MB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ x86-64

ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (3.3 MB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ ARM64

ferro_hgvs-0.7.1-cp310-abi3-macosx_11_0_arm64.whl (3.1 MB view details)

Uploaded CPython 3.10+macOS 11.0+ ARM64

ferro_hgvs-0.7.1-cp310-abi3-macosx_10_13_x86_64.whl (3.3 MB view details)

Uploaded CPython 3.10+macOS 10.13+ x86-64

File details

Details for the file ferro_hgvs-0.7.1.tar.gz.

File metadata

  • Download URL: ferro_hgvs-0.7.1.tar.gz
  • Upload date:
  • Size: 3.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for ferro_hgvs-0.7.1.tar.gz
Algorithm Hash digest
SHA256 71c6fbca26028b73ace5ac906e9c1799c0c5038197ca0c2c27f7687b5225b081
MD5 e6b2e3b6c5f13dcf7c0481ccb8c52568
BLAKE2b-256 0c6187440aab9c08a82265f2556e82992e56f90286e96fbfe53bdd5886588677

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1.tar.gz:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.7.1-cp310-abi3-win_amd64.whl.

File metadata

  • Download URL: ferro_hgvs-0.7.1-cp310-abi3-win_amd64.whl
  • Upload date:
  • Size: 3.3 MB
  • Tags: CPython 3.10+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for ferro_hgvs-0.7.1-cp310-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 deca15d3dfa43ebfd3729a4b2235233b221de7ca739e9be0dc0d0666a8983d95
MD5 cc111583d6eae2b97f7215545a94a570
BLAKE2b-256 c6449c4caab5cde4f333f99390c5273d8ee8800fcf0cd277cb9b898113b4c41e

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1-cp310-abi3-win_amd64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_x86_64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_x86_64.whl
Algorithm Hash digest
SHA256 664909f220b0d093cb47b1457961dd76baa18283366874f80cb11899e5c21a36
MD5 9d4fd5c1e3043d5ac270a7f529282cac
BLAKE2b-256 f8410509b84a10596534e15d4f3834f82553de08225676c5c351535b721679d2

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_x86_64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_aarch64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_aarch64.whl
Algorithm Hash digest
SHA256 cc8b8e080921c484b097a7191e046e9ec87887d65d3667ebda1dd69302b08569
MD5 0d91534644388cd7a12646e900e6113f
BLAKE2b-256 b51943c1767bb99de109a84a6b037e2f0856129aa2bdbd6828f61f9bd1d610b3

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1-cp310-abi3-musllinux_1_2_aarch64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 17a82120fa2b2c22d36b168e6174218ed94fa49f92c9a1a0eb82e7db758bc45b
MD5 639e457667c617aaa1757ac7928a4262
BLAKE2b-256 906afba2276da855ee49afb0cd309e760387c5ef58a8ae9423a48a5cf743b46a

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 87e39674970e2202b0dadddfec23b4eb175f4499cc739fb440934d258e4511fd
MD5 058c0b903c6d7529c6fcdb51422e1f97
BLAKE2b-256 53941c33aff2e94d6f28fdee14db7af8786e2a921a69d39966117100b75e90da

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.7.1-cp310-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.7.1-cp310-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 4517127e8dab2f66dfa0c66d64c4406d91d8dabc779de2e055f9913a92339fbb
MD5 05a750a37e4c68341b492bb1b1b304f2
BLAKE2b-256 c598b8f6a0f8ab1c66958b735fc9882ea59ed12ba4b1ee0c4447cbf35f2d3432

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1-cp310-abi3-macosx_11_0_arm64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.7.1-cp310-abi3-macosx_10_13_x86_64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.7.1-cp310-abi3-macosx_10_13_x86_64.whl
Algorithm Hash digest
SHA256 f81fdfc5cb47f9185820ef418321cd06e096a6facc106a39bda44b34cfd85385
MD5 975f69c4a8057c1c2c1bead522a43f72
BLAKE2b-256 335e9fd821bab391e0235b4195b898bb2744db6f2ea3ad4650429f1b7a60c14d

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.7.1-cp310-abi3-macosx_10_13_x86_64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page