Validation¶
This document describes how to validate vartriage against benchmark datasets. Two complementary validation approaches are used:
- GIAB chr22 — genomic-scale validation measuring specificity and runtime characteristics
- ClinVar eRepo — expert-panel curated variants measuring pathogenic sensitivity and PPV
Summary (v1.0.0 full-genome release run)¶
The v1.0.0 release figures come from a single full-genome run over an expert-curated truth set of 18,166 pinned-ClinVar variants (GRCh38), using genome-wide REVEL, gnomAD exomes v4.1.1 frequencies (remote tabix, pinned cache), and GENCODE v46. All metrics carry a 95% Wilson confidence interval.
| Metric | Value | 95% CI | n |
|---|---|---|---|
| PPV (P/LP call) | 0.999 | [0.997, 0.999] | 8,062 |
| Pathogenic sensitivity | 0.642 | [0.633, 0.650] | 12,545 |
| Specificity | 0.990 | [0.982, 0.994] | 1,187 |
| NPV | 0.207 | [0.197, 0.218] | 5,670 |
The thresholds were pre-registered before the figures existed (.kiro/specs/validation-benchmarks/accuracy-decision-rule.md). The PPV lower bound (0.997) clears the pre-registered clinical bar; sensitivity (0.642) is below the 0.70 clinical gate, so v1.0.0 is positioned as a research-grade computational-triage tool: high precision on positive calls, with uncertain variants held at VUS rather than overcalled.
Confusion counts: 8,050 true positive, 12 false positive, 1,175 true negative, 4,495 false negative. Of the P/LP variants not recovered, 4,486 were classified VUS rather than mis-called benign. Specificity recovered from 0.0 (frequency-blind REVEL-only run) to 0.990 once gnomAD frequency was available, confirming the benign-evidence path depends on frequency data rather than any classifier defect.
Coverage boundaries (these bound the figures): frequency is from gnomAD exomes v4.1.1 only (genomes not queried); the 127 mitochondrial truth variants received no gnomAD frequency; CADD and SpliceAI were not used, so REVEL was the only in-silico score. Each plausibly suppresses sensitivity and is the subject of the sensitivity-evidence-coverage work scoped for v1.2.0.
Reproducibility: the gnomAD remote cache is pinned (entries never expire), the truth set is built deterministically from a pinned ClinVar release, and reference versions are fixed, so a re-run reproduces these figures.
Path to clinical-grade¶
Clinical-grade status has two distinct requirements, and both must hold.
- Clear the pre-registered gate honestly. The only unmet threshold is sensitivity (0.642, needs ≥ 0.70). It must rise by giving the classifier evidence it currently lacks — gnomAD genomes, CADD, SpliceAI, and chrM frequency let PP3, splice, and rarity criteria fire on variants presently held at VUS — never by fitting thresholds to the truth set, which the pre-registration forbids. Whether sensitivity reaches 0.70 is then measured on a re-run under the same frozen rule. This is scoped in the sensitivity-evidence-coverage spec for v1.2.0.
- Validate beyond ClinVar. The current benchmark measures concordance against ClinVar, and the classifier itself uses ClinVar-derived evidence (PP5/BP6, review status), so the study carries a circularity risk stated plainly here. Clinical-grade validation requires an independent, orthogonal truth set the classifier did not learn from, per-gene and per-disease breakdowns rather than one global number, and concordance against the full ACMG five-tier rather than the collapsed P/VUS/B this run produced. Formal CAP/CLIA validation and regulatory review are a separate organizational track required before any diagnostic use. This is a later milestone.
Historical run (v0.17.5 eRepo, superseded)¶
The earlier eRepo figures below predate both the ClinGen SVI Bayesian point engine (the combining path since v0.18.3) and the v1.0.0 full-genome benchmark. They are retained for history and do not describe the current release.
| Benchmark | Variants | Pathogenic Sensitivity | Specificity | PPV | Runtime |
|---|---|---|---|---|---|
| GIAB chr22 | 50,284 | n/a | 71.8% | n/a | 30 s |
| ClinVar eRepo | 21,506 | 70.5% | n/a | 99.2% | n/a |
Known limitations:
- Pathogenic sensitivity 0.642: uncertain P/LP variants are held at VUS rather than escalated, pending the additional in-silico and frequency evidence scoped for v1.2.0 (gnomAD genomes, CADD, SpliceAI, chrM frequency). Sensitivity must improve through added evidence, never by fitting thresholds to the truth set.
- NPV 0.207: a non-pathogenic call is weak evidence of benignity, a consequence of the VUS-heavy behavior.
- Validation is concordance against ClinVar, which also supplies some classifier evidence (PP5/BP6, review status); an independent, orthogonal truth set is required for clinical-grade validation.
- BS2 is emitted for dominant-disorder genes when gnomAD homozygote counts are available; it does not fire for recessive genes or without homozygote-count data.
GIAB Validation¶
This document describes how to validate vartriage against Genome in a Bottle (GIAB) benchmark data, the gold standard for germline variant calling evaluation.
Overview¶
The validation uses GIAB sample HG002 (Ashkenazi Jewish trio son) with the v4.2.1 benchmark set on GRCh38. We run vartriage on the benchmark VCF and compare classification results against ClinVar assertions to measure concordance.
This is not a variant calling benchmark (vartriage does not call variants). It validates the annotation, prioritization, and ACMG classification stages against known-truth clinical significance.
What we measure¶
| Metric | Definition |
|---|---|
| Sensitivity vs ClinVar | Fraction of ClinVar Pathogenic/Likely Pathogenic variants that vartriage also classifies as Pathogenic or Likely Pathogenic |
| Classification distribution | Breakdown of P / LP / VUS across the variant set |
| Evidence tag coverage | How often each ACMG criterion (PVS1, PM2, PP3, PP5) fires |
| Missing data rate | Fraction of variants lacking gnomAD, ClinVar, or predictor scores |
Quick start¶
# Install dependencies
pip install vartriage[all]
brew install bcftools htslib wget # macOS
# or: apt-get install bcftools tabix wget # Ubuntu/Debian
# Run the validation script
./scripts/validate_giab.sh --output-dir validation_results
The script handles downloading, extraction, running the pipeline, and computing metrics. Expect ~20 GB of downloads on first run (gnomAD chr22 is the largest file). Subsequent runs reuse cached downloads.
Data sources¶
| Resource | Size | Source |
|---|---|---|
| GIAB HG002 benchmark VCF (GRCh38) | ~250 MB | NIST FTP |
| GIAB high-confidence regions BED | ~2 MB | Same FTP |
| ClinVar VCF (GRCh38) | ~80 MB | NCBI FTP |
| gnomAD v4.1.1 exomes chr22 | ~4.7 GB | gnomAD Downloads |
The script downloads chr22 only for a faster validation cycle. For full-genome validation, modify the script to download the complete gnomAD exomes VCF and remove the -r chr22 filter.
Expected results (chr22)¶
Based on GIAB HG002 chr22 with ClinVar and gnomAD annotation:
- ~85,000 variants pass quality filters in the high-confidence regions
- 5-15 variants have ClinVar Pathogenic or Likely Pathogenic assertions
- Pipeline sensitivity vs ClinVar should be >90% for P/LP concordance
- Most variants classify as VUS (expected for a healthy individual's germline)
Interpreting validation output¶
The script produces four output files:
validation_metrics.json¶
{
"validation_summary": {
"total_variants_processed": 84231,
"classification_distribution": {
"VUS": 84210,
"Likely_Pathogenic": 14,
"Pathogenic": 7
}
},
"concordance": {
"clinvar_actionable_count": 12,
"pipeline_actionable_count": 21,
"concordant_count": 11,
"sensitivity_vs_clinvar": 0.917
}
}
Key considerations¶
Why sensitivity might be below 100%:
- ClinVar assertions based on functional studies that the pipeline cannot replicate computationally
- Variants where CADD/REVEL scores are below threshold despite clinical evidence
- Variants in regions not covered by the predictor score files
Why pipeline might flag variants ClinVar does not:
- PM2 fires for rare variants absent from gnomAD, even without ClinVar annotation
- PP3 fires on high predictor scores regardless of ClinVar status
- These are candidates for clinical review, not false positives per se
Full-genome validation¶
For full validation across all chromosomes:
# Modify the script to use full gnomAD and CADD/REVEL scores
# Warning: requires ~500 GB disk space and several hours
export FULL_GENOME=true
./scripts/validate_giab.sh --output-dir full_validation
You will also need:
- CADD whole-genome pre-scored file (~350 GB uncompressed)
- REVEL scores file (~2 GB)
- SpliceAI pre-computed scores (~30 GB)
- GENCODE v44 GTF (~1.5 GB)
Trio validation¶
To validate inheritance pattern classification, use the full Ashkenazi trio:
vartriage \
--vcf giab_trio_merged.vcf.gz \
--proband HG002 \
--mother HG004 \
--father HG003 \
--output trio_validation.json \
--output-format json \
--gnomad refs/gnomad.v4.exomes.vcf.bgz \
--clinvar refs/clinvar.tsv \
--cadd-scores refs/cadd_v1.7.tsv \
--revel-scores refs/revel_v1.3.tsv
Expected: de novo variants in HG002 should be correctly identified when comparing against parental genotypes.
CI integration¶
The validation script can run in CI with a cached reference data volume:
# .github/workflows/validation.yml (manual trigger)
name: GIAB Validation
on: workflow_dispatch
jobs:
validate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- run: pip install -e .[all]
- run: sudo apt-get install -y bcftools tabix
- run: ./scripts/validate_giab.sh
- uses: actions/upload-artifact@v4
with:
name: validation-results
path: validation_results/
This runs on-demand rather than on every PR (too slow and too much data).
ClinVar Expert-Panel (eRepo) Validation¶
Historical (v0.17.5), superseded by the v1.0.0 full-genome run at the top of this document. The figures in this section predate the ClinGen SVI Bayesian point engine (the combining path since v0.18.3) and the v1.0.0 benchmark. They are retained to document the eRepo methodology and the pre-point-engine behavior, and do not describe the current release.
The eRepo validation uses ClinVar submissions from Expert Panels only (review status 3+ stars). These are the highest-confidence clinical variant interpretations available, representing consensus multi-lab assessments.
eRepo Overview¶
Unlike GIAB (which tests runtime and specificity on a healthy genome), the eRepo benchmark directly measures whether vartriage's ACMG criteria produce correct classifications for variants with established clinical significance.
Data source¶
- ClinVar VCF filtered to
CLNREVSTATcontaining "reviewed_by_expert_panel" - 21,928 variants across all chromosomes
- Split: 4,116 Pathogenic/Likely Pathogenic + 17,812 Benign/Likely Benign
Metrics (v0.17.5, SpliceAI SQLite + remote gnomAD)¶
| Metric | v0.17.3 (no SpliceAI) | v0.17.5 (SpliceAI SQLite) |
|---|---|---|
| Classified variants | 21,614 | 21,506 |
| Pathogenic sensitivity | 65.6% | 70.5% |
| Splice-site sensitivity | 9.8% | 55.9% |
| Benign sensitivity | 6.0% | 6.0% |
| Positive predictive value (PPV) | 99.2% | 99.2% |
Stratified by consequence type (v0.17.5)¶
| Consequence | Sensitivity |
|---|---|
| Missense | 46.9% |
| Frameshift | 99.7% |
| Splice_Site | 55.9% |
| Synonymous | 0.0% |
Interpretation¶
High PPV (99.2%) means when vartriage calls something Pathogenic or Likely Pathogenic, it is almost always correct. The tool is conservative and precise — it does not over-call.
Moderate sensitivity (70.5%) means vartriage misses about 30% of known pathogenic variants. This is expected given:
- Missense sensitivity is 46.9%: many pathogenic missense variants don't exceed the ClinGen-calibrated REVEL threshold (0.644)
- Splice-site sensitivity is 55.9%: splice calls depend on precomputed SpliceAI delta scores being available for the variant
- Some pathogenic variants require functional evidence the tool cannot assess computationally
Low benign sensitivity (6.0%) means most benign variants are classified as VUS rather than Benign/Likely Benign. This is because:
- BP1 (missense in a gene where truncating variants cause disease) is not yet implemented
- BP3 (in-frame in repetitive region without known function) is not yet implemented
- BP6 (reputable source reports benign) is not yet implemented
- Without benign criteria firing, variants default to VUS
Running the eRepo validation¶
The validation script lives in the paper reproducibility repository (vartriage-streaming-acmg):
This produces stratified results in results/erepo_stratified.json and summary metrics in the console.
Comparison to existing tools¶
| Tool | Pathogenic Sensitivity | PPV | Architecture |
|---|---|---|---|
| VarTriage v0.17.5 | 70.5% | 99.2% | Streaming, zero-dependency |
| BIAS-2015 (reference) | ~85% | ~75% | Batch, requires InterVar |
| AutoACMG | ~78% | ~82% | Web service |
VarTriage is positioned on high precision (PPV) and operational simplicity (single pip install, streaming architecture, no Java/Perl), not on raw sensitivity. Further sensitivity gains depend on additional benign criteria (BP1/BP3/BP6) and functional evidence the tool cannot assess computationally.
Provenance¶
Validation results are tracked with full provenance: