Skip to main content

CI Nightly reference-aware tests Codecov Crates.io PyPI Python versions Documentation License install with bioconda DOI

ferro-hgvs

A high-performance HGVS variant nomenclature parser and normalizer written in Rust.

WARNING: ALPHA SOFTWARE - USE AT YOUR OWN RISK

This software is currently in ALPHA. While we have extensively tested it across a wide variety of HGVS patterns, no guarantees are made regarding correctness or stability.

Fulcrum Genomics

Features

  • Full HGVS Parsing: All coordinate systems (g/c/n/r/p/m/o) and edit types
  • Variant Normalization: 3'/5' shifting per HGVS specification
  • High Performance: ~5M variants/sec single-threaded parsing (>12M/s parallel), zero-copy with nom
  • Type-Safe: Leverages Rust's type system for correctness

Installation

Python

pip install ferro-hgvs

Pre-built wheels are available for Linux (x86_64, aarch64), macOS (x86_64, Apple Silicon), and Windows (x86_64) on Python 3.10+.

Rust

Add to your Cargo.toml:

[dependencies]
ferro-hgvs = "0.1"

Or install the CLI:

cargo install ferro-hgvs

Quick Start

CLI

# Parse a variant
ferro parse "NM_000088.3:c.459A>G"

# Parse from file
ferro parse -i variants.txt -f json

# Prepare reference data (downloads RefSeq, genome, cdot — RefSeq-only by default)
ferro prepare --output-dir ferro-reference

# Verify reference data is ready
ferro check --reference ferro-reference

# (Optional) pre-build the on-disk cdot cache as a setup step, so the one-time
# cache build doesn't slow the start of a real (or timed/benchmarked) run.
ferro check --reference ferro-reference --build-cache

# Normalize with reference
ferro normalize "NM_000088.3:c.459del" --reference ferro-reference/

Throughput tip: when normalizing many variants, feed them sorted by transcript accession (or by genomic position). ferro caches each resolved transcript, so consecutive variants on the same transcript skip the (dominant) cost of re-reading and re-building it from the reference. Sorted input keeps the relevant transcripts resident in the cache and is markedly faster on large batches — see Performance Comparison.

Optional reference data

A bare ferro prepare builds a RefSeq-only reference (accessions NM_/NR_/NP_/NG_). Two opt-in flags provision additional data — pass them at prepare time; they are what a fully-provisioned ("blessed") reference is built with:

# Add Ensembl support (accessions ENST/ENSG/ENSP). Downloads the Ensembl cdot
# metadata and cDNA FASTAs (~1 GB+); off by default. Without it, an ENST/ENSG/ENSP
# input reports "Reference not found" and the message points back at this flag.
ferro prepare --output-dir ferro-reference --ensembl

# Derive version-independent NG_ placements and the NG_→transcript-version map
# (ng_hosted_transcripts) for a curated list of RefSeqGene accessions. Required to
# resolve legacy gene-symbol selectors (NG_(GENE):c.…) and bare-NG_ hosted lookups.
ferro prepare --output-dir ferro-reference \
  --derive-ng-placements path/to/ng_accessions.txt

# A fully-provisioned reference combines both in one run:
ferro prepare --output-dir ferro-reference --ensembl \
  --derive-ng-placements path/to/ng_accessions.txt

Both flags are incremental: re-running ferro prepare over an existing reference adds the requested data and preserves already-provisioned artifacts.

Library

use ferro_hgvs::{parse_hgvs, HgvsVariant};

fn main() -> Result<(), ferro_hgvs::FerroError> {
    let variant = parse_hgvs("NM_000088.3:c.459A>G")?;

    match &variant {
        HgvsVariant::Cds(v) => println!("CDS variant: {}", v),
        HgvsVariant::Genome(v) => println!("Genomic variant: {}", v),
        _ => println!("Other: {}", variant),
    }

    Ok(())
}

Python

import ferro_hgvs

# Parse a variant
variant = ferro_hgvs.parse("NM_000088.3:c.459A>G")
print(variant.variant_type)  # "coding"
print(variant.reference)     # "NM_000088.3"
print(str(variant))          # "NM_000088.3:c.459A>G"

# Normalize with reference data
normalizer = ferro_hgvs.Normalizer(reference_json="ferro-reference/cdot.json")
normalized = normalizer.normalize("NM_000088.3:c.459del")

Supported HGVS Syntax

Type Prefix Example
Genomic g. NC_000001.11:g.12345A>G
Coding DNA c. NM_000088.3:c.459A>G
Non-coding n. NR_000001.1:n.100A>G
RNA r. NM_000088.3:r.459a>g
Protein p. NP_000079.2:p.Val600Glu
Mitochondrial m. NC_012920.1:m.3243A>G

Edit Types

  • Substitution: A>G, Val600Glu
  • Deletion: del, 100_200del
  • Insertion: 100_101insATG
  • Deletion-Insertion: 100_102delinsATG
  • Duplication: 100_102dup
  • Inversion: 100_200inv
  • Repeat: 100CAG[20]

CLI Commands

The ferro CLI provides commands beyond parsing and normalization:

Command Description
prepare Download and prepare reference data for normalization
check Verify reference data setup
parse Parse and validate HGVS variants
normalize Normalize HGVS variants (3'/5' shifting)
explain Explain error/warning codes (e.g., ferro explain W1001)
annotate-vcf Annotate VCF files with HGVS notation
vcf-to-hgvs Convert VCF records to HGVS
hgvs-to-vcf Convert HGVS to VCF format
liftover Liftover coordinates between genome builds
describe Generate HGVS from reference/observed sequences
effect Predict protein effect from variant
backtranslate Reverse translate protein to DNA variants
convert-gff Convert GFF3/GTF to transcripts.json
generate Generate HGVS descriptions from components
extract-hgvs Extract HGVS from VEP-annotated VCFs

Error Handling

ferro-hgvs provides configurable error handling with three modes:

Mode Behavior
strict Reject non-conformant input (default)
lenient Auto-correct with warnings
silent Auto-correct silently
# Use lenient mode to auto-correct common issues
ferro parse --error-mode lenient "p.val600glu"  # Corrects to p.Val600Glu

# Ignore specific warnings
ferro parse --ignore W1001,W2001 "p.val600glu"

# Get help on any error/warning code
ferro explain W1001
ferro explain --list

Configuration File

Create .ferro.toml in your project directory:

[error-handling]
mode = "lenient"
ignore = ["W1001", "W2001"]  # Silently correct these
reject = ["W3003"]           # Always reject these

Why ferro-hgvs?

ferro-hgvs provides the most comprehensive HGVS variant normalization across all pattern types, with performance orders of magnitude faster than alternatives.

Normalization Capabilities Comparison

Pattern Type ferro mutalyzer biocommons hgvs-rs
Genomic (g.)
Coding (c.) exonic
Coding (c.) intronic ✓**
Non-coding (n.)
RNA (r.)
Protein (p.) Net*

* mutalyzer protein normalization requires network access for NP_→NM_ lookups (cannot be cached locally). ** mutalyzer intronic support is enabled by default via genomic-context rewriting; disable with --no-rewrite-intronic.

Performance Comparison

All tools are benchmarked in ferro's offline configuration — best case for every tool. Reference data is preloaded locally (a local UTA database and SeqRepo) and the network is disabled, so the figures below measure parse/normalize compute, not I/O. Out of the box, hgvs-rs, biocommons/hgvs, and mutalyzer resolve each variant against a remote UTA/SeqRepo or the Mutalyzer web API — a network round-trip per variant (~100–1000 ms), i.e. roughly 1–10 variants/sec, hundreds to thousands of times slower than shown here (an order-of-magnitude estimate from per-call network latency, not separately benchmarked). That local, offline setup is exactly what ferro's prepare command builds; ferro needs no external service.

Median patterns/sec over 5 reps on an Apple M2 Max, local/offline. All tools draw from one stratified ClinVar population; per-tool sample sizes are calibrated so each tool is measured over a meaningful interval — fast cells (e.g. ferro/hgvs-rs parse) draw from millions of patterns, while slower cells (e.g. the per-tool normalize columns) draw from as few as tens to thousands. All tools exclude process/interpreter startup from the timed region — the mutalyzer/biocommons Python subprocesses are timed by their own internal startup-excluded timer, matching ferro/hgvs-rs. Only ferro parallelizes natively (rayon); the other tools are single-threaded libraries, so their normalize @8 workers figures come from the benchmark harness running 8 independent instances in parallel, while parsing is not sharded for them — hence the single-threaded label in their parse @8 workers column (mutalyzer normalize likewise shows no gain at 8 workers: per-call cache and IPC overhead dominate, so sharding does not help). Every tool runs fully offline against local reference data — a local UTA database and SeqRepo, with mutalyzer's network lookups disabled — the configuration ferro's prepare command enables; the figures therefore reflect compute throughput, not per-variant network latency. Reference-data load is excluded for all tools. ferro full-population peak: parse 20.0M/s, normalize 77.0k/s. See docs/BENCHMARK_RUNBOOK.md for the full method.

Parse

Tool Throughput @ 1 worker Throughput @ 8 workers ferro speedup @ 8w
ferro 5.1M/s 12.2M/s
mutalyzer 352/s single-threaded 35,000×
biocommons 3.9k/s single-threaded 3,100×
hgvs-rs 3.6M/s single-threaded

Normalize

Tool Throughput @ 1 worker Throughput @ 8 workers ferro speedup @ 8w
ferro 78.1k/s 260.2k/s
mutalyzer 4/s 4/s 73,000×
biocommons 368/s 818/s 320×
hgvs-rs 195/s 1.3k/s 200×

ferro thread scaling

Threads 1 2 4 8
ferro parse 5.1M/s 9.4M/s 16.0M/s 12.0M/s

Input ordering matters for batch throughput. Resolving a transcript (reading its full sequence from the reference and rebuilding its CDS/exon metadata) dominates per-variant cost. ferro memoizes resolved transcripts in a bounded in-memory cache, so repeated lookups of the same transcript are near-free. Providing variants sorted by transcript accession — or by genomic position, which clusters variants onto the same transcripts — maximizes the cache hit rate and can speed up large batches by an order of magnitude versus randomly-ordered input. Ordering matters most when the number of distinct transcripts in the run exceeds the cache capacity (very large or genome-wide inputs); below that, the working set stays resident regardless of order.

Reference Data: What ferro Prepares

The ferro prepare command downloads and organizes all reference data needed for comprehensive normalization. This data is then shared with other tools (mutalyzer, biocommons, hgvs-rs) to enable their local operation.

Data Type Source Size Enables
RefSeq transcripts NCBI ~1GB NM_/NR_/XM_ normalization
cdot metadata MANE ~200MB Transcript-to-genome mappings
GRCh38 + GRCh37 genomes NCBI ~4GB NC_ genomic normalization
RefSeqGene (sequences + genome alignments) NCBI ~600MB NG_ gene-region normalization; projecting c./n. variants into an NG_ parent's own g. frame (via the RefSeqGene→genome alignment GFF3)
LRG sequences + XML EBI ~50MB LRG_ stable-reference normalization; projecting c./n. variants into an LRG_ parent's own g. frame (via the LRG XML genomic mapping)
Protein sequences Derived from CDS ~200MB NP_/XP_ protein normalization
Legacy transcript versions NCBI ~50MB Historical ClinVar variants

Key insight: Without ferro's reference preparation, other tools require network access for each variant lookup (adding 100-1000ms latency per variant). With ferro's cached reference data, all tools can operate fully offline with consistent, reproducible results.

Deriving version-independent NG_ placements (#728)

ferro prepare --derive-ng-placements <accessions.txt> derives genomic placements for the listed NG_ versions (one exact accession per line, e.g. NG_012337.3; blank lines and # comments ignored), writing derived_refseqgene_placements.json into the reference directory and wiring the manifest's derived_refseqgene_placements field. This fills version gaps the archived RefSeqGene→genome GFF3 snapshots do not cover. It needs cdot + the genome in the same prepare run and uses NCBI EFetch per accession; accessions that cannot be validated are skipped with a warning. The field is preserved across subsequent prepare runs.

Benchmark: Reference Data & Tool Comparison

The main ferro binary includes commands to prepare reference data (ferro prepare) and check its status (ferro check). The ferro-benchmark tool (build with --features benchmark) extends this for tool comparison benchmarks.

Command Description
prepare <tool> Prepare reference data for a tool
check <tool> Verify tool configuration and dependencies
parse <tool> Parse HGVS patterns with specified tool
normalize <tool> Normalize HGVS patterns with specified tool
compare results Compare parse/normalize results between tools
extract Extract patterns from ClinVar, VCFs, or create samples
setup Set up UTA database, SeqRepo, and other services
generate Generate summary reports and configs
collate Aggregate sharded results

Quick Start

# Prepare ferro reference (main binary - no special features needed)
ferro prepare --output-dir data/ferro

# Check reference data
ferro check --reference data/ferro

# Normalize with ferro
ferro normalize -i patterns.txt --reference data/ferro

# For tool comparison, build with benchmark support
cargo build --release --features benchmark

# Prepare other tools (uses ferro reference for transcript data)
ferro-benchmark prepare mutalyzer --ferro-reference data/ferro --output-dir data/mutalyzer
ferro-benchmark prepare biocommons --seqrepo-dir data/seqrepo --uta-dump uta_20210129b.pgd.gz --ferro-reference data/ferro

# Compare results between tools
ferro-benchmark normalize mutalyzer -i patterns.txt -o mutalyzer.json --mutalyzer-settings data/mutalyzer/mutalyzer_settings.conf
ferro-benchmark compare results normalize ferro.json mutalyzer.json -o comparison.json

Supported tools: ferro-hgvs, mutalyzer, biocommons/hgvs, hgvs-rs

Note: The pixi.toml and pixi.lock files in this repository define a pixi environment for the Python-based external tools (mutalyzer, biocommons/hgvs, seqrepo) used in benchmarking. Run pixi shell to activate it.

See docs/BENCHMARK_GUIDE.md for detailed usage.

Development

cargo build
cargo test
cargo clippy -- -D warnings

License

Licensed under the MIT License. See LICENSE for details.

Disclaimer

This software is under active development. While we make a best effort to test this software and to fix issues as they are reported, this software is provided as-is without any warranty (see the license for details). Please submit an issue, and better yet a pull request as well, if you discover a bug or identify a missing feature. Please contact Fulcrum Genomics if you are considering using this software or are interested in sponsoring its development.

Contributing

See CONTRIBUTING.md for guidelines.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ferro_hgvs-0.12.0.tar.gz (4.0 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

ferro_hgvs-0.12.0-cp310-abi3-win_amd64.whl (3.7 MB view details)

Uploaded CPython 3.10+Windows x86-64

ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_x86_64.whl (4.0 MB view details)

Uploaded CPython 3.10+musllinux: musl 1.2+ x86-64

ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_aarch64.whl (3.8 MB view details)

Uploaded CPython 3.10+musllinux: musl 1.2+ ARM64

ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.8 MB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ x86-64

ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (3.6 MB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ ARM64

ferro_hgvs-0.12.0-cp310-abi3-macosx_11_0_arm64.whl (3.5 MB view details)

Uploaded CPython 3.10+macOS 11.0+ ARM64

ferro_hgvs-0.12.0-cp310-abi3-macosx_10_13_x86_64.whl (3.7 MB view details)

Uploaded CPython 3.10+macOS 10.13+ x86-64

File details

Details for the file ferro_hgvs-0.12.0.tar.gz.

File metadata

  • Download URL: ferro_hgvs-0.12.0.tar.gz
  • Upload date:
  • Size: 4.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for ferro_hgvs-0.12.0.tar.gz
Algorithm Hash digest
SHA256 8725a10cc044f92221785078cd741454ae43ecb18d07306804d0d661b091295c
MD5 a26641c4f87db2aed8c70c05b4f9e91e
BLAKE2b-256 b731736cb610cf549cad5303f82757bf261103cb64023282f096b19c57b67638

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0.tar.gz:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.12.0-cp310-abi3-win_amd64.whl.

File metadata

  • Download URL: ferro_hgvs-0.12.0-cp310-abi3-win_amd64.whl
  • Upload date:
  • Size: 3.7 MB
  • Tags: CPython 3.10+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for ferro_hgvs-0.12.0-cp310-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 d28e4ab2ce6c7cb92ca4782cd1bb1f751df4914f515a24acff304bc2bffffce2
MD5 01915b78bd8ae639dc1df66322275504
BLAKE2b-256 c57cd0fac590bad510d72f0bd9b3992690800860328a8c0a6c4c37f9f46f7946

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0-cp310-abi3-win_amd64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_x86_64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_x86_64.whl
Algorithm Hash digest
SHA256 2535cb353f41b7107a442783a3e2785b44da0384d1b50df8ab04339461501902
MD5 e1746bce8b1c667b35af48c0cb4ad187
BLAKE2b-256 0bdac3ba4214afd618fa2bdbde5d49f2e9f6a6e18d8bb1a675a8d2cce8044608

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_x86_64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_aarch64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_aarch64.whl
Algorithm Hash digest
SHA256 8282ed89eacb8776ab494928e7726a432e3053ebc58ad7f23d9003ca39698ea4
MD5 d79400e45b3ec83459e583ffeca68a18
BLAKE2b-256 2708d9bbd02c21545aa906070ca8959af70470709fcbfc1bc9b0a38c36c0d30f

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0-cp310-abi3-musllinux_1_2_aarch64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 c34e7b1778d7b175564592c70ef2aedd4fb0836f3745ee2bed67ce6a20611a5d
MD5 c7269cbdd614b82a5f3bafd747db76b1
BLAKE2b-256 9fb5f6d23c7fa71238413f817bae2bc0e6aacfe5c142ed097e7141d56d6c3686

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 4cf47b961c516253e45af36df431b1c510a4547e41b4ebd9e04244ff28371784
MD5 964a905f5816d514f86d907bedd41680
BLAKE2b-256 939886893b3d533efb4f9638bd4940bb9fd43781423c06ba0f8a9f9cc77a5f39

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.12.0-cp310-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.12.0-cp310-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 7f9c9cc3c02d0a9cb9b32c91094d211f466499c61843256f3c6841ba24b98694
MD5 b9c7f6e2067764e8c5c99dd5d80315da
BLAKE2b-256 f13e934bf0f275050d2c35ac96dd051dd71ecfb61074b69bfa23c0ce37766b2a

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0-cp310-abi3-macosx_11_0_arm64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ferro_hgvs-0.12.0-cp310-abi3-macosx_10_13_x86_64.whl.

File metadata

File hashes

Hashes for ferro_hgvs-0.12.0-cp310-abi3-macosx_10_13_x86_64.whl
Algorithm Hash digest
SHA256 9d83157729e35c469f166e77cd10705533868879e214ea67f158d310abfd5a0c
MD5 6bc7c0ae7ea679cfe2666292115e0f25
BLAKE2b-256 a1250676297320f809c04bacb85de6331b61a0f26aab2111365e2738e776f929

See more details on using hashes here.

Provenance

The following attestation bundles were made for ferro_hgvs-0.12.0-cp310-abi3-macosx_10_13_x86_64.whl:

Publisher: release-wheels.yml on fulcrumgenomics/ferro-hgvs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.0.0

8 files

0.17.2

8 files

0.17.1

8 files

0.17.0

8 files

0.16.0

8 files

0.15.0

8 files

0.14.0

8 files

0.13.1

8 files

0.13.0

8 files

This release

0.12.0 This release

8 files

0.11.0

8 files

0.10.1

8 files

0.10.0

8 files

0.9.1

8 files

0.9.0

8 files

0.8.1

8 files

0.8.0

8 files

0.7.1

8 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page