Bioinformatics
vartriage: clinical variant interpretation
A pure-Python library that turns a raw VCF into an ACMG-classified clinical variant report for gene panels and whole genomes.
Research software. It is not validated for clinical use, and its output should not guide diagnosis or treatment without review by a qualified clinician.
Sources for the numbers on this page: vartriage README: benchmarks and validation.
Problem
A sequencing run produces a VCF with millions of variant calls. Turning that file into a clinical answer, which handful of variants are pathogenic and worth a clinician’s attention, normally means chaining several heavy tools together. Annotators, frequency databases, pathogenicity scorers, and ACMG classifiers each arrive as separate programs with their own Java, Perl, or Spark runtimes and their own file conventions.
That stack is awkward to install, hard to reproduce, and heavy on memory. A whole-genome file with four million variants does not fit comfortably in RAM when a tool loads everything at once. vartriage exists to do the whole path in one Python package: VCF in, ACMG-classified report out, with no Java, Perl, or Spark dependency and a memory footprint that stays bounded on genome-scale input.
How it works
vartriage is a streaming pipeline. Variants flow through a chain of stages, each an iterator over the last, so the process holds only a small window of records in memory at a time. Optional stages activate based on config, so a targeted panel skips work that only a whole genome needs.
Reference lookups have two backends behind Protocol interfaces. Without optional extras the library uses pure-Python fallbacks: dict lookups and a bisect-based interval tree. With the accelerated extra it swaps in polars for batch frequency and ClinVar joins and pyranges for interval overlap queries. The output is the same either way. Parsed reference files are cached as pickle files next to the source, and the cache rebuilds automatically when the source changes or the vartriage version changes.
Scoring and classification follow published clinical standards. Composite prioritization combines normalized CADD, REVEL, and SpliceAI (REVEL 0.5, CADD 0.3, SpliceAI 0.2, redistributing when a score is missing). ACMG/AMP evidence tagging uses ClinGen SVI-calibrated thresholds, and tags combine into the five tiers via Bayesian-adapted combining rules from Tavtigian et al. 2018. Beyond nuclear SNVs, the package carries separate pipelines for structural variants (ClinGen 2020 framework) and mitochondrial DNA with heteroplasmy, MITOMAP, and HelixMTdb.
[QCPreflight] -> VCFParser -> [SampleExtractor] -> [RegionFilter]
-> QualityFilter -> AnnotationEngine -> [GeneFilter]
-> [SecondaryFindingsFilter] -> [GeneKnowledgeAnnotator]
-> PrioritizationEngine -> [PhenotypeBoost]
-> ACMGClassifier -> ReportGenerator
(stages in brackets activate based on config)Hard parts
- Genome-scale memory: streaming iterators keep a WGS QC pass over 4M variants under 2 GB RAM, with a measured peak of 122 MB on the QC-only workload.
- Correct missense vs synonymous calls: consequences are resolved at the codon level against an indexed reference FASTA.
- No 80 GB downloads: remote tabix scoring queries CADD and gnomAD over HTTP byte-range requests, cached in a local SQLite database (30-day default TTL, pinnable for reproducibility) with a circuit breaker so a network failure does not stall the pipeline.
- Missing data as a first-class state: variants absent from gnomAD or ClinVar are never dropped; they are flagged, pass the frequency gate, and raise a tracked warning, so no variant disappears silently.
- Multiple-transcript conflicts resolve deterministically to the most damaging consequence via a fixed severity ordering.
- A typed public API with a py.typed marker that passes mypy --strict, with no Any in the annotation engine interfaces.
Results
The repository records these performance and validation figures. Benchmarks: WGS QC-only over 4M variants ran in 156 s at 122 MB peak RSS; chr22 full annotation on GIAB (50K variants) ran in 30 s at about 670 MB; chr22 annotation against 100K gnomAD (130K variants) ran in 19.5 s at 453 MB.
Validation (v0.17.5, SpliceAI SQLite plus remote gnomAD): GIAB chr22 over 50,284 variants reported 71.8% specificity; the ClinVar eRepo benchmark over 21,506 variants reported 70.5% pathogenic sensitivity and 99.2% PPV. Splice-site sensitivity improved from 9.8% to 55.9% once the SpliceAI SQLite backend landed.
These validation figures come from the v0.17.5 run and predate the point-based classifier introduced in v0.18.3; they have not been re-measured.
- 12 ACMG/AMP criteria implemented with strength modulation, plus Secondary Findings screening across 71 medically actionable genes (SF v3.2).
- Test coverage gate at 75%, CI across Python 3.10, 3.11, and 3.12.
- Limitations stated in the README: BS2 has no evaluator; BP1/BP3/BP6 are not implemented; benign sensitivity is low (6.0%), so VUS is the default when evidence is absent.
Artifacts
- Package: vartriage on PyPI, MIT licensed, with extras [accelerated], [pdf], [clinical], [api], and [all].
- CLI covering single-sample runs, cohort analysis, standalone SV triage, QC pre-flight, remote-tabix presets, and score-bundle downloads.
- Report formats: JSON, CSV, PDF, IGV-loadable annotated VCF, and clinical HTML/PDF/DOCX, each clinical report paired with a JSON audit-trail sidecar and a computational-only disclaimer.
- Source on GitHub with a MkDocs site, GitHub Actions CI, CodeQL scanning, and trusted-publisher PyPI releases.