vartriage
Clinical variant interpretation library for gene panels and whole genomes. VCF in, ACMG-classified report out.
pip install vartriage[all]
vartriage --vcf patient.vcf.gz --output report.html --output-format clinical-html \
--patient-id PAT-001 --panel-name "Cardiac Panel v3" --use-bundles
What it does: quality filtering, consequence annotation (GENCODE, with codon-level resolution via reference FASTA), population frequency lookup (gnomAD, population-specific via local files, remote tabix, or API), pathogenicity scoring (CADD/REVEL/SpliceAI with ClinGen-calibrated thresholds), gene-disease linkage (OMIM/ClinGen/HPO/gnomAD constraint), phenotype-driven prioritization, ACMG/AMP classification (12 criteria with strength modulation: PVS1, PS1, PM1, PM2, PM4, PM5, PP3, PP5, BA1, BS1, BP4, BP7), Bayesian-adapted combining rules (Tavtigian et al. 2018), trio inheritance analysis, multi-sample cohort analysis (recurrence, gene burden), ACMG Secondary Findings screening, structural variant triage (ClinGen 2020 framework), mitochondrial variant analysis (mtDNA-specific classification with heteroplasmy, MITOMAP, and HelixMTdb), remote tabix scoring (CADD/gnomAD via HTTP byte-range, no 80 GB download), VCF quality control (Ti/Tv, het/hom, variant count sanity checks with strict-gate support), and clinical report generation with audit trail and computational-only disclaimer.
Why use it:
- Single Python package, no Java/Perl/Spark dependencies
- Streams 4M+ variant WGS files under 2 GB RAM
- Codon-level consequence calling with reference FASTA (correct missense vs synonymous)
- Benign + pathogenic ACMG criteria (12 criteria, ClinGen-calibrated): classifies variants across all 5 tiers
- Bayesian-adapted combining rules (Tavtigian et al. 2018): 2 Moderate = LP, 1 Moderate + 4 Supporting = LP
- Gene-disease linkage: OMIM, ClinGen validity, HPO phenotype matching, gnomAD constraint, actionability
- Phenotype-driven:
--hpo-termsboosts variants in genes matching patient symptoms - Trio-aware: de novo, dominant, recessive, compound het, X-linked
- ACMG Secondary Findings (SF v3.2): screens 71 medically actionable genes
- Score bundle downloader:
vartriage bundle download --bundle clinvarfetches and prepares reference files - API mode: annotate gene panels via Ensembl VEP + ClinVar + gnomAD API with zero local files
- Outputs: JSON, CSV, PDF, HTML clinical reports, IGV-loadable annotated VCF
- Structural variant triage: ClinGen 2020 framework for DEL/DUP/INV/INS/BND/CNV
- Mitochondrial variant analysis: automatic chrM detection, heteroplasmy quantification, MITOMAP/HelixMTdb annotation, mtDNA-specific classification
- Remote tabix scoring: query CADD (80 GB) and gnomAD (15+ GB) via HTTP byte-range without local downloads (
--cadd-remote cadd-v1.7-grch38) - Typed API with Protocol-based backends
Benchmarks:
| Workload | Variants | Wall time | Peak RSS |
|---|---|---|---|
| WGS QC only | 4M | 156 s | 122 MB |
| chr22 full annotation (GIAB) | 50K | 30 s | ~670 MB |
| chr22 annotation (100K gnomAD) | 130K | 19.5 s | 453 MB |
Reference files are cached after first parse. Subsequent runs load from cache in seconds.
Validation (eRepo run with SpliceAI SQLite + remote gnomAD, v0.17.5):
| Benchmark | Variants | Pathogenic Sensitivity | Specificity | PPV |
|---|---|---|---|---|
| GIAB chr22 | 50,284 | n/a | 71.8% | n/a |
| ClinVar eRepo | 21,506 | 70.5% | n/a | 99.2% |
Splice-site sensitivity improved from 9.8% to 55.9% once the SpliceAI SQLite backend (--spliceai-db) landed in v0.17.5.
Known limitations:
- BS2 is defined as an evidence tag but is not emitted by the classifier (it needs gnomAD homozygote-count data that is not parsed yet).
- BP1, BP3, BP6 benign criteria are not implemented.
- Benign sensitivity is low (6.0%) because of the missing benign criteria; VUS is the default when evidence is absent.
Install
pip install vartriage
Optional extras:
pip install vartriage[accelerated] # polars + pyranges backends
pip install vartriage[pdf] # reportlab PDF reports
pip install vartriage[clinical] # weasyprint + python-docx for clinical HTML/PDF/DOCX reports
pip install vartriage[api] # httpx for API annotation mode
pip install vartriage[all] # everything
CLI
vartriage --vcf sample.vcf.gz --output candidates.json
Score bundles
Download reference files automatically:
# See available bundles
vartriage bundle list
# Download ClinVar + gnomAD for chr22
vartriage bundle download --bundle clinvar
vartriage bundle download --bundle gnomad-exomes-chr22
# Run with auto-resolved reference paths
vartriage --vcf sample.vcf.gz --output results.json --use-bundles
API mode
Annotate variants via remote APIs with zero local reference files:
# Gene panel, no downloads needed
vartriage --vcf panel.vcf --output results.json --mode api
# Hybrid: local gnomAD + API for ClinVar/CADD
vartriage --vcf panel.vcf --output results.json --mode hybrid --gnomad gnomad.tsv
# With NCBI API key for faster ClinVar queries
vartriage --vcf panel.vcf --output results.json --mode api --api-key YOUR_KEY
Queries Ensembl VEP, ClinVar, CADD, and SpliceAI. Responses are cached in SQLite for instant re-runs. See API Mode Guide for configuration and performance details.
Cohort analysis
Analyze multiple samples together to find shared variants, compute recurrence frequencies, and generate per-gene burden reports:
# Multiple VCF files directly
vartriage cohort --vcf sample1.vcf.gz sample2.vcf.gz sample3.vcf.gz \
--output cohort_results/ --cohort-name "cardiac_cohort"
# From a manifest file (one VCF path per line, optional tab-separated labels)
vartriage cohort --manifest samples.tsv --output cohort_results/ \
--cohort-name "cardiac_cohort" --output-format csv
# With annotation references and gene filtering
vartriage cohort --manifest samples.tsv --output cohort_results/ \
--gene-annotation gencode.v44.gtf --gnomad gnomad.v4.sites.tsv \
--clinvar clinvar.tsv --cadd-scores cadd.tsv --revel-scores revel.tsv \
--gene-list cardiac_panel.txt --use-bundles
# Parallel processing with custom thresholds
vartriage cohort --manifest samples.tsv --output cohort_results/ \
--parallel --max-workers 8 --min-recurrence 3 --max-af 0.01 --no-singletons
Manifest format - plain text, one VCF path per line. Optional tab-separated second column for sample labels:
# Cardiac cohort 2026
/data/vcfs/patient_001.vcf.gz Patient 001
/data/vcfs/patient_002.vcf.gz Patient 002
/data/vcfs/patient_003.vcf.gz Patient 003
Output files - three files per run:
| File | Contents |
|---|---|
{cohort_name}_variants.json |
All cohort variants with recurrence counts, per-sample classifications, and evidence tags |
{cohort_name}_gene_burden.json |
Per-gene statistics: variant count, pathogenic count, penetrance, samples affected |
{cohort_name}_summary.json |
Top-level metrics: total variants, shared/singleton/universal counts, top recurrent genes |
Key options:
| Flag | Default | Description |
|---|---|---|
--min-recurrence |
2 | Exclude variants appearing in fewer than this many samples |
--max-af |
0.05 | Exclude variants above this population frequency |
--no-singletons |
false | Drop variants seen in only one sample |
--parallel |
false | Process samples concurrently |
--max-workers |
4 | Thread pool size for parallel mode |
Parallel mode uses ThreadPoolExecutor. Per-sample work is I/O-bound (pysam releases the GIL during C-level VCF parsing), so threads are effective here without needing multiprocessing.
Gene-disease linkage
Connect variants to clinical context with phenotype-driven prioritization:
# Phenotype-driven: boost epilepsy-related genes for a seizure patient
vartriage --vcf patient.vcf.gz --output results.json \
--gene-annotation gencode.gtf --gnomad gnomad.tsv \
--hpo-terms HP:0001250,HP:0001249,HP:0002197
# Filter to autosomal recessive genes
vartriage --vcf patient.vcf.gz --output results.json \
--gene-annotation gencode.gtf --gnomad gnomad.tsv \
--inheritance-mode AR
# Only actionable findings (ClinGen curations)
vartriage --vcf patient.vcf.gz --output results.json \
--gene-annotation gencode.gtf --gnomad gnomad.tsv \
--flag-actionable
# Combine all three
vartriage --vcf patient.vcf.gz --output results.json \
--gene-annotation gencode.gtf --gnomad gnomad.tsv \
--hpo-terms HP:0001250,HP:0001249 --inheritance-mode AD --flag-actionable
Output includes per-variant: disease associations (with MIM numbers), ClinGen validity level, gnomAD constraint metrics (pLI/LOEUF/mis_z), actionability status, and phenotype match score.
See Gene-Disease Linkage Guide for data file formats, Python API usage, and validation details.
Structural variant triage
Classify structural variants (DEL, DUP, INV, INS, BND, CNV) using the ClinGen 2020 technical standards:
# Standalone SV analysis
vartriage sv --sv-vcf sv_calls.vcf.gz --output sv_report.json \
--gene-annotation gencode.v46.gtf \
--dosage-sensitivity clingen_dosage.tsv \
--gnomad-sv gnomad_sv.bed \
--pathogenic-regions clingen_pathogenic_regions.bed
# Combined SNV + SV analysis
vartriage --vcf snv.vcf.gz --sv-vcf sv_calls.vcf.gz \
--output report.json --gene-annotation gencode.v46.gtf
Pipeline: SVParser (streams from VCF) -> SVAnnotator (gene overlap, dosage sensitivity, gnomAD-SV frequency) -> SVScorer (composite pathogenicity score) -> SVClassifier (ClinGen evidence sections 1-4, 5-tier classification).
Supports Manta, DELLY, GATK-SV, GRIDSS, and LUMPY. See Structural Variants Guide for the full CLI reference and Python API.
Mitochondrial variant analysis
Automatic detection and classification of mitochondrial DNA variants using mtDNA-specific criteria:
# Automatic: chrM/MT variants in the VCF are detected and classified separately
vartriage --vcf wgs.vcf.gz --output results.json
# Custom heteroplasmy threshold (default: 1%)
vartriage --vcf wgs.vcf.gz --output results.json --mt-min-heteroplasmy 5.0
# Skip mitochondrial analysis (targeted panels without mtDNA capture)
vartriage --vcf panel.vcf.gz --output results.json --skip-mito
The mitochondrial pipeline uses the vertebrate mitochondrial genetic code for amino acid prediction, extracts heteroplasmy levels from AD/AF fields, queries MITOMAP for disease associations, checks HelixMTdb for population frequency, and classifies variants independently of the nuclear ACMG criteria. Results appear in a dedicated "Mitochondrial Findings" section in output reports.
See Mitochondrial Variants Guide for heteroplasmy thresholds, classification rules, and data update instructions.
Quality control
Compute sample-level QC metrics (Ti/Tv, het/hom, variant count, ins/del) and validate them against expected ranges before annotation runs. Catches contamination, sample swaps, and caller artifacts early:
# Standalone QC check, no annotation
vartriage qc --vcf sample.vcf.gz --sample SAMPLE1 --assay-type wes
# Write a machine-readable QC report
vartriage qc --vcf sample.vcf.gz --sample SAMPLE1 --assay-type wgs \
--output-json qc_report.json
# Gate the full pipeline: halt before annotation on any FAIL
vartriage --vcf sample.vcf.gz --output results.json \
--gene-annotation gencode.gtf --gnomad gnomad.tsv \
--assay-type wgs --strict-qc
# Skip QC for a pre-validated file
vartriage --vcf sample.vcf.gz --output results.json \
--gene-annotation gencode.gtf --gnomad gnomad.tsv --skip-qc
Each metric is validated against assay-specific ranges (wgs, wes, panel) and reported as PASS, WARN, or FAIL with an overall verdict printed to stderr. Without --strict-qc, WARN and FAIL are logged and the pipeline proceeds; with --strict-qc, a FAIL halts before annotation and exits with code 3. Warn ranges are overridable via --expected-titv, --expected-het-hom, or a [qc] section in ~/.vartriage/config.toml.
See Quality Control Guide for the metric definitions, thresholds, and Python API.
Remote tabix scoring
Query CADD and gnomAD scores from public servers via HTTP byte-range requests. No local download of the 80+ GB CADD file or 15+ GB gnomAD required:
# CADD scores via remote tabix (named preset)
vartriage --vcf panel.vcf --output results.json \
--gene-annotation gencode.gtf --cadd-remote cadd-v1.7-grch38
# gnomAD frequencies via remote tabix
vartriage --vcf panel.vcf --output results.json \
--gene-annotation gencode.gtf --gnomad-remote gnomad-exomes-v4-grch38
# Both together with local ClinVar
vartriage --vcf patient.vcf.gz --output report.json \
--gene-annotation gencode.gtf --clinvar clinvar.tsv \
--cadd-remote cadd-v1.7-grch38 --gnomad-remote gnomad-exomes-v4-grch38
# Pinned cache for clinical reproducibility (scores never expire)
vartriage --vcf panel.vcf --output results.json \
--cadd-remote cadd-v1.7-grch38 --remote-cache-ttl -1
# List available presets
vartriage remote list-presets
Scores are cached locally in SQLite (~/.vartriage/remote_cache.db) with a 30-day TTL by default. Local files always take priority over remote. A circuit breaker prevents the pipeline from stalling on network failures.
See Remote Tabix Guide for presets, caching, Python API, and configuration details.
Full options
vartriage \
--vcf sample.vcf.gz \
--output report.json \
--output-format json \
--reference-fasta GRCh38.fa \
--gene-annotation gencode.v44.gtf \
--gnomad gnomad.v4.sites.tsv \
--clinvar clinvar_20240101.tsv \
--cadd-scores cadd_scores.tsv \
--revel-scores revel_scores.tsv \
--spliceai-scores spliceai_scores.tsv \
--gene-list my_panel.txt \
--regions target_regions.bed \
--sample PROBAND_01 \
--min-gq 20 \
--secondary-findings
Clinical report options:
vartriage \
--vcf sample.vcf.gz \
--output clinical_report.html \
--output-format clinical-html \
--patient-id PAT-2026-001 \
--panel-name "Cardiac Panel v3" \
--gene-annotation gencode.v44.gtf \
--gnomad gnomad.v4.sites.tsv \
--clinvar clinvar_20240101.tsv \
--cadd-scores cadd_scores.tsv \
--revel-scores revel_scores.tsv \
--spliceai-scores spliceai_scores.tsv \
--gene-list cardiac_panel.txt
Formats: clinical-html (self-contained HTML), clinical-pdf (requires weasyprint), clinical-docx (requires python-docx). Both --patient-id and --panel-name are required for clinical formats.
Run vartriage --help for the complete list.
Python API
Run the whole pipeline:
from pathlib import Path
from vartriage import (
Pipeline, PipelineConfig, AnnotationConfig,
PrioritizationConfig, QualityFilterConfig, ReportConfig,
)
config = PipelineConfig(
vcf_path=Path("sample.vcf.gz"),
output_path=Path("candidates.json"),
quality_filter=QualityFilterConfig(min_qual=30.0),
annotation=AnnotationConfig(
gene_annotation_path=Path("gencode.v44.gtf"),
gnomad_path=Path("gnomad.v4.sites.tsv"),
clinvar_path=Path("clinvar_20240101.tsv"),
),
prioritization=PrioritizationConfig(
max_allele_frequency=0.01,
cadd_scores_path=Path("cadd_scores.tsv"),
revel_scores_path=Path("revel_scores.tsv"),
),
report=ReportConfig(output_format="json"),
)
pipeline = Pipeline(config)
pipeline.run()
Or use stages individually:
from vartriage import VCFParser, QualityFilter, QualityFilterConfig
with VCFParser(Path("input.vcf.gz")) as parser:
qf = QualityFilter(QualityFilterConfig(min_qual=30.0))
for variant in qf.apply(iter(parser)):
print(f"{variant.chrom}:{variant.pos} {variant.ref}>{variant.alt}")
API mode (no local reference files needed):
from pathlib import Path
from vartriage import Pipeline, PipelineConfig, ReportConfig
from vartriage.api.config import APIConfig
api_config = APIConfig.load(
mode="api",
genome_build="grch38",
ncbi_api_key="your-key-here", # or set NCBI_API_KEY env var
)
config = PipelineConfig(
vcf_path=Path("panel.vcf"),
output_path=Path("results.json"),
report=ReportConfig(output_format="json"),
api=api_config,
)
pipeline = Pipeline(config)
pipeline.run()
Cohort analysis across multiple samples:
from pathlib import Path
from vartriage import CohortPipeline, CohortConfig, PipelineConfig, AnnotationConfig, PrioritizationConfig
cohort_config = CohortConfig(
sample_vcfs=[
Path("patient_001.vcf.gz"),
Path("patient_002.vcf.gz"),
Path("patient_003.vcf.gz"),
],
output_path=Path("cohort_results/"),
cohort_name="cardiac_cohort",
min_recurrence=2,
max_af_threshold=0.05,
output_format="json",
)
# Optional: shared annotation config applied to all samples
annotation = AnnotationConfig(
gene_annotation_path=Path("gencode.v44.gtf"),
gnomad_path=Path("gnomad.v4.sites.tsv"),
)
pipeline = CohortPipeline(
cohort_config=cohort_config,
annotation_config=annotation,
)
report_paths = pipeline.run()
# Access results programmatically
for variant in pipeline.variants:
if variant.sample_count >= 2:
print(f"{variant.gene_name}: {variant.chrom}:{variant.pos} "
f"in {variant.sample_count}/{variant.total_samples} samples")
for burden in pipeline.gene_burdens:
if burden.pathogenic_count > 0:
print(f"{burden.gene_name}: {burden.pathogenic_count} pathogenic, "
f"penetrance={burden.penetrance:.0%}")
Structural variant triage:
from pathlib import Path
from vartriage.structural import SVTriagePipeline, SVTriageConfig
config = SVTriageConfig(
vcf_path=Path("sv_calls.vcf.gz"),
output_path=Path("sv_report.json"),
gene_annotation_path=Path("gencode.v46.gtf"),
dosage_sensitivity_path=Path("clingen_dosage.tsv"),
gnomad_sv_path=Path("gnomad_sv.bed"),
pathogenic_regions_path=Path("pathogenic_regions.bed"),
)
pipeline = SVTriagePipeline(config)
output = pipeline.run()
Pipeline stages
[QCPreflight] → VCFParser → [SampleExtractor] → [RegionFilter] → QualityFilter → AnnotationEngine → [GeneFilter] → [SecondaryFindingsFilter] → [GeneKnowledgeAnnotator] → PrioritizationEngine → [PhenotypeBoost] → ACMGClassifier → ReportGenerator
Stages in brackets are optional and activate based on config.
QC pre-flight (--assay-type, --strict-qc, --skip-qc) - Streams the VCF once before annotation to compute Ti/Tv, het/hom, variant count, and ins/del ratios, then validates them against assay-specific ranges. Prints a PASS/WARN/FAIL table to stderr. With --strict-qc, a FAIL halts the pipeline (exit code 3) before annotation runs. Disabled with --skip-qc.
Sample extraction (--sample) - Pulls a single sample from multi-sample VCFs. Only variants where the named sample carries an alternate allele are kept. Optional --min-gq threshold drops low-confidence genotype calls.
Region filtering (--regions) - Restricts to variants overlapping intervals in a BED file. Useful for gene panel target regions.
Quality filtering - Drops variants where FILTER isn't PASS/., QUAL is below threshold (default 20), or QUAL is missing entirely.
Annotation - Adds functional consequence (from GTF gene models), population frequency (gnomAD), and ClinVar significance. Multiple-transcript conflicts resolve to the most damaging consequence. Consequence severity: Frameshift > Nonsense > Splice_Site > Missense > In_Frame_Insertion > In_Frame_Deletion > Synonymous > Intergenic.
Gene filtering (--gene-list) - After annotation, restricts to variants in genes from a user-supplied text file. Case-insensitive matching. Logs a warning for any panel genes with zero hits (catches typos).
Prioritization - Two phases. First: frequency gate drops variants with AF above the threshold (default 0.01); unknown-frequency variants always pass. Second: composite scoring from normalized CADD Phred, REVEL, and SpliceAI:
composite = (REVEL × 0.5) + (CADD_normalized × 0.3) + (SpliceAI × 0.2)
CADD normalization: Phred score divided by 99.0 (CADD_MAX_PHRED), capped at 1.0. The separate prioritization_score field uses the same Phred / 99.0 normalization for non-missense variants. REVEL and SpliceAI are already bounded 0.0–1.0 and used directly without rescaling.
When only two scores are available, weights redistribute proportionally. Single available score is used directly. Falls back to the legacy two-score formula (0.6/0.4) when SpliceAI is not configured.
ACMG classification - Tags evidence per ACMG/AMP 2015 guidelines with ClinGen SVI-calibrated thresholds (Pejaver et al. 2022):
| Tag | Strength | Condition |
|---|---|---|
| PVS1 | Very Strong | Nonsense, Frameshift, or Splice_Site + SpliceAI > 0.8 (strength modulated by pLI/LOEUF) |
| PVS1_Strong | Strong | PVS1 downgraded when gene constraint is moderate (0.5 < pLI < 0.9) |
| PS1 | Strong | Same amino acid change as ClinVar Pathogenic via different nucleotide (requires protein index + reference FASTA) |
| PM1 | Moderate | Missense in a critical functional domain (missense constraint region, gnomAD mis_z > 3.09) |
| PM2 | Moderate | All population AFs < 0.0001, or absent from gnomAD (population-aware) |
| PM4 | Moderate | In-frame insertion/deletion or stop-loss variant in a non-repetitive region |
| PM5 | Moderate | Novel missense at amino acid position with known pathogenic missense in ClinVar (requires protein index) |
| PP3 | Supporting | REVEL > 0.644 or SpliceAI > 0.5 on splice-adjacent |
| PP3_MODERATE | Moderate | REVEL > 0.773 (ClinGen-calibrated) |
| PP5 | Supporting | ClinVar Pathogenic without conflicting Benign |
| BA1 | Standalone | Any population AF > 5% (standalone Benign, overrides all pathogenic evidence) |
| BS1 | Strong | Any population AF > 1% (strong benign) |
| BP4 | Supporting | REVEL < 0.290 (missense) or CADD < 10 (non-missense) |
| BP4_MODERATE | Moderate | REVEL < 0.183 (ClinGen-calibrated) |
| BP7 | Supporting | Synonymous + SpliceAI < 0.1 |
Tags combine into Pathogenic, Likely_Pathogenic, VUS, Likely_Benign, or Benign using Bayesian-adapted combining rules (Tavtigian et al. 2018). Relaxed LP rules: 2 Moderate = LP, 1 Moderate + 4 Supporting = LP. BA1 is standalone and overrides all conflicting pathogenic evidence. Missing data sources mean the tag is simply omitted.
BS2 (strong benign, observed in healthy adults) exists in the evidence-tag enum and the combining rules but has no evaluator; it is never emitted until gnomAD homozygote-count parsing is added.
Report output - JSON and CSV stream directly from the iterator (no buffering). PDF materializes for page layout. VCF re-reads the source file, injects VARTRIAGE_* INFO fields for classified variants, and writes bgzipped output with a tabix index. Clinical formats (clinical-html, clinical-pdf, clinical-docx) produce structured reports with a computational-only disclaimer (citing ACMG/AMP 2015), per-variant evidence narratives, an executive summary, a Sample Quality Control section (when QC runs), findings table, evidence cards, limitations, methodology, and sign-off sections. A JSON audit trail sidecar (.audit.json) is written alongside each clinical report. Output fields: chromosome, position, ref/alt alleles, gene_name, functional consequence, allele frequency, revel_score, composite rank, prioritization_score, ClinVar assertion, ACMG classification, evidence tags, disease_associations, clingen_validity, gene_constraint, is_actionable, phenotype_match_score.
Configuration
QualityFilterConfig
| Field | Type | Default | Range |
|---|---|---|---|
| min_qual | float | 20.0 | 0-1,000,000 |
AnnotationConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| gene_annotation_path | Path | required | GTF/GFF |
| gnomad_path | Path | required | TSV or tabix VCF (.vcf.bgz/.vcf.gz) |
| clinvar_path | Path | None | TSV |
| reference_fasta_path | Path | None | Indexed FASTA (.fa + .fai) for codon-level consequence calling |
| batch_size | int | 10,000 | 1,000-100,000 |
PrioritizationConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| max_allele_frequency | float | 0.01 | 0.0-1.0 |
| cadd_scores_path | Path | None | CADD Phred TSV |
| revel_scores_path | Path | None | REVEL TSV |
| spliceai_scores_path | Path | None | SpliceAI TSV |
| batch_size | int | 10,000 | 1,000-100,000 |
ReportConfig
| Field | Type | Default | Options |
|---|---|---|---|
| output_format | str | "json" | "json", "csv", "pdf", "vcf", "clinical-html", "clinical-pdf", "clinical-docx" |
ClinicalReportConfig
| Field | Type | Default | Options |
|---|---|---|---|
| patient_id | str | required | Patient identifier |
| panel_name | str | required | Gene panel name |
| output_format | str | required | "clinical-pdf", "clinical-html", "clinical-docx" |
| report_template | str | "standard" | Template name |
Constructed automatically when --output-format is a clinical-* value. Requires --patient-id and --panel-name.
GeneFilterConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| gene_list_path | Path | required | Plain text, one gene symbol per line |
RegionFilterConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| bed_path | Path | required | BED file with target intervals |
SampleConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| sample_name | str | required | Sample name from VCF header |
| min_gq | int | None | Genotype quality threshold (0-99) |
CohortConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| sample_vcfs | list[Path] | required | At least 2 VCF file paths |
| output_path | Path | required | Output directory for reports |
| cohort_name | str | "cohort" | Identifier for output filenames |
| min_recurrence | int | 2 | Minimum samples for recurrence |
| output_format | str | "json" | "json" or "csv" |
| max_af_threshold | float | 0.05 | Max population AF for inclusion (0.0-1.0) |
| include_singletons | bool | True | Include variants in only 1 sample |
| sample_labels | dict | None | Map file stems to display labels |
| parallel | bool | False | Process samples concurrently |
| max_workers | int | 4 | Thread pool size (>= 1) |
KnowledgeBaseConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| data_dir | Path | None | None | Custom TSV directory (defaults to bundled package data) |
| hpo_terms | frozenset[str] | empty | Patient HPO terms (HP:NNNNNNN format) |
| inheritance_mode | str | None | None | Filter: AD, AR, XL, XLD, XLR, MT |
| flag_actionable | bool | False | Filter to ClinGen actionable genes only |
Any non-default field activates the gene-disease linkage pipeline stage.
MitoConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| enabled | bool | True | Enable/disable mitochondrial analysis |
| min_heteroplasmy | float | 1.0 | Minimum heteroplasmy % for reporting (0.0-100.0) |
| gene_map_path | Path | None | None | Custom mt_gene_map.tsv (defaults to bundled) |
| mitomap_path | Path | None | None | Custom mitomap_pathogenic.tsv (defaults to bundled) |
| helixmtdb_path | Path | None | None | Custom helixmtdb_frequency.tsv (defaults to bundled) |
Mitochondrial analysis is auto-enabled when chrM/MT variants are present in the VCF. Use --skip-mito (CLI) or MitoConfig(enabled=False) to disable.
QCConfig
| Field | Type | Default | Notes |
|---|---|---|---|
| assay_type | str | "wes" | wgs, wes, or panel (selects threshold preset) |
| strict | bool | False | FAIL halts the pipeline with exit code 3 |
| skip | bool | False | Bypass QC entirely |
| expected_ti_tv | tuple | None | None | Override Ti/Tv warn range (min, max) |
| expected_het_hom | tuple | None | None | Override het/hom warn range (min, max) |
| sample_id | str | None | None | Sample for het/hom (single-sample VCFs auto-detect) |
QC runs before annotation. Use --assay-type, --strict-qc, and --skip-qc (CLI) or set PipelineConfig(qc=QCConfig(...)). Warn ranges are also overridable via a [qc] section in ~/.vartriage/config.toml.
Reference file formats
All TSV with a header row. Tab-separated.
Gene list - Plain text file, one gene symbol per line. Comment lines starting with # are ignored. Blank lines are skipped. Symbols are matched case-insensitively.
# Cardiac panel v2
BRCA1
BRCA2
TP53
MLH1
gnomAD (TSV) - columns: chrom, pos, ref, alt, af. The value '.' in the af column is treated as null (gnomAD compatibility).
gnomAD (tabix VCF) - bgzipped VCF with a .tbi index (.vcf.bgz or .vcf.gz). When you point gnomad_path at a tabix-indexed file, vartriage queries it on the fly with zero memory overhead for the reference. Useful when your gnomAD file is too large to fit in RAM as a dict.
ClinVar - columns: chrom, pos, ref, alt, clinical_significance. Values: Pathogenic, Likely pathogenic, Uncertain significance, Likely benign, Benign.
CADD / REVEL / SpliceAI - columns: chrom, pos, ref, alt, score. Lines starting with # are skipped. All three use the same TSV format.
Missing data handling
Variants absent from gnomAD are never dropped; they get frequency_unknown=True and pass the frequency filter. Same for ClinVar: no match means clinvar_unknown=True.
A MissingDataWarning fires per lookup miss. After a run:
acc = pipeline.warning_accumulator
print(f"{acc.total_count} missing data events across {acc.sources}")
Warning hierarchy
All warnings inherit from VarTriageWarning (a UserWarning subclass). Silence everything at once:
import warnings
from vartriage import VarTriageWarning
warnings.filterwarnings("ignore", category=VarTriageWarning)
Dependencies
| Package | Required | Extra | Purpose |
|---|---|---|---|
| pysam >=0.22,<1.0 | yes | - | VCF streaming via htslib |
| numpy >=1.24,<3.0 | yes | - | Score normalization |
| polars >=0.20,<2.0 | no | [accelerated] | Batch frequency/ClinVar joins |
| pyranges >=0.1,<1.0 | no | [accelerated] | Interval overlap queries |
| reportlab >=4.0,<5.0 | no | [pdf] | PDF report rendering |
| weasyprint >=69.0,<71.0 | no | [clinical] | Clinical HTML/PDF rendering |
| pydyf >=0.11,<0.13 | no | [clinical] | PDF writer (matches weasyprint) |
| python-docx >=1.0,<2.0 | no | [clinical] | Clinical DOCX rendering |
| httpx >=0.27,<1.0 | no | [api] | Remote API annotation |
Without optional extras, the library uses pure-Python fallbacks (dict lookups, bisect-based interval tree). Same output either way; the accelerated path is faster on large reference files.
Caching
Reference files (GTF gene models, CADD scores, REVEL scores, SpliceAI scores) are parsed once and cached as pickle files adjacent to the source (with a .vartriage.cache suffix). On subsequent runs, the cache loads in seconds instead of re-parsing.
Cache invalidation is automatic: if the source file's mtime changes or the vartriage version changes, the cache rebuilds. Writes are atomic (temp file + rename), so a crash mid-write won't corrupt anything.
To force a fresh parse, delete the .vartriage.cache file next to your reference.
Type checking
The package ships a py.typed marker (PEP 561). All protocol return types are fully typed, with no Any in the annotation engine interfaces.
mypy --strict vartriage/
Tests
pytest tests/ # full suite
pytest tests/ -m "not slow" # skip benchmarks
CI
GitHub Actions runs on Python 3.10, 3.11, and 3.12. PyPI publishing uses trusted publisher (no token in secrets).
Project layout
vartriage/
cli.py # CLI entry point
pipeline.py # Orchestrator
protocols.py # Protocol interfaces (IntervalIndex, FrequencyDatabase, etc.)
io/ # VCF parsing
filter/ # Quality, region, sample, and gene filtering
annotation/ # Consequence, frequency, ClinVar lookups
knowledge/ # Gene-disease linkage (OMIM, ClinGen, HPO, constraint)
prioritization/ # AF gating + CADD/REVEL/SpliceAI scoring (ScoreLoader)
classification/ # ACMG evidence tagging
structural/ # SV triage: parser, annotator, scorer, classifier, report
mito/ # Mitochondrial: genetic code, heteroplasmy, MITOMAP, classifier
reporting/ # JSON, CSV, PDF, VCF (streaming writers)
clinical/ # Clinical report generation (HTML/PDF/DOCX + audit trail)
cohort/ # Multi-sample cohort analysis
models/ # Dataclasses, enums, configs, warnings
bundle/ # Score bundle downloader and transformer
api/ # Remote API annotation (VEP, ClinVar, CADD, SpliceAI)
data/knowledge/ # Bundled gene-disease TSV files
data/mito/ # Bundled mtDNA reference data (gene map, MITOMAP, HelixMTdb)
_internal/ # Batch utils, interval tree, caching, vectorized ops
py.typed # PEP 561 marker
Contributing
See CONTRIBUTING.md.
License
MIT
Metadata
Release files for vartriage 0.18.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vartriage-0.18.3.tar.gz | 737.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vartriage-0.18.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.1 MB
Release files / vartriage-0.18.3.tar.gz
| Download URL | vartriage-0.18.3.tar.gz |
|---|---|
| Size | 737.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
70856070fb6deaa8e1fd11d50ef656196b58ed2c97c786b874c2d5fbd9f67b06
|
|
BLAKE2b-256 checksum How to use checksums |
dc18a60f0d0004e4c4a5a89ec58acdb84bb8056d5e974cb4f687f193962ce769
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency logRelease files / vartriage-0.18.3-py3-none-any.whl
| Download URL | vartriage-0.18.3-py3-none-any.whl |
|---|---|
| Size | 357.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c5c0cb92c56ab9362d23dad7074446ccd644135432195d33d7d018e060cae1df
|
|
BLAKE2b-256 checksum How to use checksums |
cb7e074670d79d5535496ea8a21206e919705c8af62b3e9e99d447b48c02731a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency log