Changelog¶
All notable changes to vartriage are documented here.
Format follows Keep a Changelog. Versioning follows Semantic Versioning.
[Unreleased]¶
Nothing yet. The most recent release is 1.0.1, below.
[1.0.1] - 2026-10-05¶
Fixed¶
- Reference-cache compatibility across backend upgrades: the on-disk cache for parsed reference data (notably the GENCODE gene-model used for consequence assignment) now stamps and checks the versions of its serialization-backend libraries (pandas, pyranges, polars) alongside the vartriage version, Python version, and source mtime it already validated. A cache written under one backend version is rejected and transparently rebuilt rather than deserialized into an object a newer backend mishandles. No behaviour change when the cache is fresh or absent.
- Reference-cache namespacing per backend: the pyranges consequence backend and the pure-Python interval-tree backend serialize incompatible payloads but previously shared one
.vartriage.cachefilename per source, so whichever ran last could leave a cache the other could not read. Each backend now tags its cache file, so both can coexist for the same source and switching backends between runs no longer corrupts the read path. - Vectorized consequence path degrades instead of aborting: if the pyranges vectorized overlap raises at run time, the affected batch now falls back to the pure-Python consequence annotator (same rules, same result) for that batch rather than failing the whole classification run. Makes the optional accelerated path (
vartriage[accelerated]) robust to backend hiccups.
[1.0.0] - 2026-10-04¶
Added¶
- Quantified full-genome accuracy validation: v1.0.0 is the first release with release-grade accuracy figures measured against an expert-curated truth set of 18,166 pinned-ClinVar variants (12,545 Pathogenic/Likely Pathogenic, 5,621 Benign/Likely Benign) on GRCh38. On the Pathogenic/Likely Pathogenic call the classifier reaches a PPV of 0.999 (95% CI [0.997, 0.999], n=8,062) with specificity 0.990 (95% CI [0.982, 0.994]); pathogenic sensitivity is 0.642 (95% CI [0.633, 0.650]). The thresholds and their interpretation were pre-registered before any figure existed. Judged against that rule, v1.0.0 is a research-grade computational-triage tool: the PPV lower bound clears the clinical bar, while sensitivity below 0.70 keeps it short of clinical-grade. Full detail, coverage boundaries, and the path to clinical-grade are in the validation guide (docs/validation.md).
- Remote gnomAD per-population frequency cache: the ancestry-aware population lookup backing BA1, BS1, BS2 and PM2 is now cached on disk with the same TTL and clinical-pinning semantics as the scalar score cache, storing each variant's global and per-population allele frequencies and homozygote count. The first run over a cohort populates and pins the cache; subsequent runs read from it rather than re-querying the remote server. Results persist per query-group, so an interrupted run resumes from the last completed group.
- Benchmark runner and truth-set builder:
vartriage/validation/benchmark.pycomputes PPV, sensitivity, specificity, NPV, Cohen's kappa and a five-class confusion matrix, each binomial metric with a Wilson confidence interval;scripts/build_truth_set.pybuilds the normalized truth set deterministically from a pinned ClinVar release.
Changed¶
- Remote tabix queries are robust to transient failures: a remote gnomAD range query now treats a transient read failure (a timeout, a socket error, or a truncated block) as a retryable event and degrades to no frequency for that one range after exhausting retries, and the index open is bounded by a connect timeout. A single stalled or failed range read can no longer abort a genome-wide run.
Internal¶
- The test suite redirects its temporary base to a filesystem that can host a WAL SQLite database when the inherited
TMPDIRcannot, so cache-backed tests exercise the code rather than the environment's filesystem limits.
[0.19.0] - 2026-10-03¶
Added¶
- BS2 evaluator (observed in healthy controls): the classifier now fires BS2 at Strong benign strength when a variant is observed homozygous in gnomAD (homozygote count above zero) and the gene is associated with a dominant disorder in the gene-disease knowledge base. A homozygous observation of a dominant-pathogenic candidate in a healthy-control database is strong evidence against pathogenicity. BS2 does not fire for recessive disorders, where homozygous carriers are expected, and is applicable only when a dominant association is known. Homozygote counts are read from the gnomAD
nhomalt/AC_Homfield; when a dominant gene lacks that count,gnomAD_homozygotesis recorded as a missing source. - PP5 strength tracks ClinVar review status: when the ClinVar reference supplies a review-status column, a Pathogenic assertion reviewed by an expert panel contributes PP5 at Strong (PP5_Strong), while single-submitter and multiple-submitter assertions contribute PP5 at Supporting, and an assertion with no criteria or with conflicting interpretations does not contribute PP5. A reference in the basic format without a review-status column fires PP5 at Supporting for every Pathogenic assertion, matching prior behavior.
PP5_Strongis a new evidence tag scored at the Strong tier by the SVI point engine. - Disease-specific frequency thresholds for PM2, BA1, and BS1: with disease-aware thresholds enabled, the benign frequency gates are selected per variant from the gene's inheritance mode. A dominant disorder uses stricter thresholds (BA1 at 0.001, BS1 at 0.0003, PM2 at 0.00001) than a recessive one (BA1 at 0.05, BS1 at 0.01, PM2 at 0.0001). A gene with no known inheritance mode, or a run with the feature disabled, uses the fixed defaults, so existing output is unchanged until a disease context is supplied.
Changed¶
- PVS1 is downgraded to Strong for variants that escape nonsense-mediated decay: when the transcript CDS structure is available, a Very Strong PVS1 for a nonsense or frameshift variant becomes Strong (PVS1_Strong) if the variant falls in the last exon, within 50 nt of the final exon-exon junction, or in a single-exon gene, because such a truncating variant may still produce a partially functional protein. The transcript structure is derived from the GTF annotation already loaded for consequence calling, so no new reference file is required. The downgrade only ever lowers strength and never affects a non-null variant; when no transcript structure is available the strength is unchanged and
transcript_structureis recorded as a missing source for a Very Strong PVS1. - PVS1 without gene-constraint data resolves to Strong rather than Very Strong: a null variant whose gene has no pLI constraint record is assigned PVS1 at Strong, since loss of function cannot be confirmed as the disease mechanism from absence of data. The gene-list and pLI routes to Very Strong are unchanged. (Documentation for this behavior, shipped in the 0.17.x line, is corrected in this release.)
- The trio inheritance filter recognizes consanguineous recessive transmission: a homozygous-recessive variant in the proband now classifies as recessive when at least one parent is heterozygous and the other parent is heterozygous, homozygous-alt, or unavailable, rather than requiring both parents heterozygous. A parent confirmed homozygous-reference still rejects the call.
- The trio inheritance filter integrates gene-disease inheritance mode: when a gene-inheritance map is supplied, inheritance calls incompatible with the gene's expected mode are dropped — a dominant-only gene keeps dominant and de novo calls, a recessive gene keeps recessive and compound-het calls, an X-linked gene keeps X-linked calls. A gene absent from the map, or no map at all, leaves all patterns in place.
Internal¶
AnnotatedVariantcarries an optionalclinvar_review_status;PopulationFrequenciescarries an optionalhom_count. The annotation engine populates both when the backend exposes them, through optional backend methods that leave existing backends unchanged.
[0.18.5] - 2026-09-21¶
Changed¶
- Structural-variant classification thresholds are calibrated to the evidence scale: the copy-number classifier maps its accumulated ClinGen Section 1-4 evidence score to the five-tier verdict at thresholds that match the additive strength of the evidence lines it evaluates (a strong line contributes about 0.45, moderate 0.25, supporting 0.10). Pathogenic is reached by two independent strong lines, Likely Pathogenic by one; the benign side is symmetric, so one strong benign line reaches Likely Benign and two reach Benign. This makes every tier, including the benign tiers, reachable from evidence the pipeline accumulates.
- A clearly common structural variant is strong benign evidence on its own: a gnomAD-SV population frequency at or above 5 percent now contributes a strong benign score line, enough to reach Likely Benign without any other evidence, and enough to reach Benign when combined with a benign-region overlap. The 1 percent, 0.5 percent, and 0.1 percent bands are unchanged.
- Missing structural-variant annotation resolves to neutral evidence: an SV with no gnomAD-SV frequency match now scores neutral for the frequency component rather than the maximum pathogenic-direction value, and a gene overlap with no ClinGen dosage record scores neutral rather than a partial-pathogenic default. Absence of a lookup result is treated as absence of evidence, not as evidence of rarity or dosage sensitivity, matching the point-variant path.
[0.18.4] - 2026-09-20¶
Performance¶
- Splice-site consequence check is vectorized over exons: the pyranges consequence index evaluates the donor and acceptor splice-site windows as a NumPy boolean mask across a chromosome's exon boundaries, rather than iterating exon rows in Python once per variant. Annotation cost stays flat as exon count grows; a full chr22 run that previously took over an hour now completes in minutes, with identical splice-site calls.
[0.18.3] - 2026-09-18¶
Added¶
- ClinGen SVI Bayesian point engine (
vartriage.classification.points): a pure scoring function that assigns each ACMG evidence criterion a signed point value by strength (Supporting 1, Moderate 2, Strong 4, Very Strong 8; benign criteria negated), sums them, and maps the total through the Tavtigian threshold ladder (Pathogenic at 10, Likely Pathogenic 6 to 9, Likely Benign minus 6 to minus 1, Benign at minus 7), with BA1 as a stand-alone Benign override.combine_evidencedelegates to this engine. - PP3_Strong for high-confidence missense (REVEL > 0.932): the PP3 computational-evidence ladder now escalates to Strong at the ClinGen-calibrated REVEL threshold of 0.932 (Pejaver et al. 2022), above the existing Moderate (> 0.773) and Supporting (> 0.644) tiers. Under the SVI point system a missense variant absent from gnomAD (PM2, Moderate) with a REVEL score above 0.932 (PP3_Strong) now reaches Likely Pathogenic on computational and frequency evidence alone, which the Moderate ceiling could not express.
- Tri-state lookup resolution type (
vartriage.models.Resolution): a value resolved from an external source (allele frequency, CADD score, gene constraint) now carries one of three states,FOUND,CONFIRMED_ABSENT, orLOOKUP_FAILED, so a source that was consulted and returned nothing is distinguishable from a source that could not be consulted. The type is additive; existing lookups are unchanged until later work routes them through it.
Changed¶
- Evidence combining uses the ClinGen SVI point system:
combine_evidencenow sums the signed points for a variant's evidence tags and maps the total through the Tavtigian threshold ladder, replacing the combinatorial rule table. Opposing evidence resolves by arithmetic (a pathogenic and a benign criterion net to their difference) rather than by forcing VUS whenever both signs are present;has_conflicting_evidenceremains as a reporting-only annotation. This corrects tier assignments the rule table produced, most notably that two Moderate criteria sum to 4 points and classify as VUS rather than Likely Pathogenic, and that a single Very Strong criterion (8 points) reaches Likely Pathogenic. Across the 1,023 pathogenic-side tag combinations, 31 percent change tier under the point system. - PM2 requires consulted population data: the "absent from or rare in controls" criterion now fires only when the population database was consulted, either an observed allele frequency below the rarity threshold in every available population, or a confirmed gnomAD miss recorded during annotation. A variant with no frequency data and no record of a completed lookup is treated as missing data rather than as evidence of rarity.
- PVS1 strength reflects the loss-of-function mechanism evidence: a null variant (nonsense or frameshift) reaches Very Strong only when loss of function is an established mechanism for the gene, shown by membership of a supplied LoF gene list or by gnomAD constraint (pLI greater than 0.9). When no such evidence is available the mechanism is unknown, and the criterion fires at Strong rather than Very Strong. The gene-list and constraint routes to Very Strong are unchanged.
- BP4 has a CADD fallback for missense variants: when a missense variant carries no REVEL score, a low CADD Phred (below 10) now supports BP4, mirroring the CADD path already used for other consequence classes and the SpliceAI fallback the pathogenic PP3 criterion uses. A missense variant with neither REVEL nor CADD records REVEL as a missing source. The REVEL-driven BP4 and BP4_Moderate calls are unchanged and still take precedence when REVEL is present.
- Coding single-base variants are classified from their resolved codon: a substitution inside a coding region is refined to NONSENSE (stop gained), STOP_LOSS (stop removed), or SYNONYMOUS (same amino acid) using the codon the annotation engine resolves, instead of the interval backend's base "coding SNV maps to missense" call. The refinement runs at the engine level, so the pyranges and pure-Python backends now return the same amino-acid-level consequence for the same variant. Variants whose codon cannot be resolved keep their base call.
- Multi-allelic records are split into one variant per ALT: a VCF line carrying several comma-separated ALT alleles now yields one variant per allele, each sharing the record's CHROM, POS, REF, QUAL, FILTER, and INFO, so every alternate allele is triaged instead of only the first.
- ClinVar significance parsing handles compound and qualified terms: significance strings are normalized (whitespace, case) and parsed for compound assertions joined by "/" or "|" (the more clinically significant component wins), a trailing qualifier after a comma, and the "Conflicting interpretations of pathogenicity" category (treated as uncertain). A recognized assertion such as "Pathogenic/Likely pathogenic" is no longer discarded as unknown. Both the pure-Python and polars ClinVar loaders share one parser.
- Per-population gnomAD frequencies are threaded through annotation: when the frequency backend can return ancestry-group frequencies, the annotation engine now populates the variant's per-population frequencies (the global AF plus each
AF_<pop>) rather than only the single global number. BA1, BS1, and PM2 can then reason about a variant that is common in one ancestry but globally rare. Backends that expose only the global lookup are unchanged and leave the per-population frequencies unset. - Remote score cache is isolated by reference build, version, and dataset: cached gnomAD and CADD scores are now keyed by a source id derived from the resolved preset (which encodes build, version, and dataset kind, for example
gnomad-genomes-v4-grch38) or, for a raw URL, a stable hash of that URL. A GRCh37 run and a GRCh38 run, or an exomes run and a genomes run, no longer share cache rows. - Installed bundles are checksum-verified on load: when a bundle manifest records a checksum for its transformed file, the on-disk file is verified before the pipeline uses it and a mismatch raises rather than being used silently. A manifest with no recorded checksum skips verification unchanged.
- Knowledge-library loaders validate their TSV header on load: the OMIM, HPO, ClinGen validity, actionability, and gnomAD constraint loaders now check that their required columns are present and raise an error naming the file and the missing columns, instead of a renamed or missing column loading zero rows with only an informational log.
- QC warns instead of passing when a metric cannot be evaluated: a sample with no genotype calls (the het/hom ratio) or no indels (the ins/del ratio) is now reported as a warning rather than a pass, so an empty or sites-only input no longer clears strict QC on those axes silently.
- Cohort config rejects the singleton recurrence inversion:
include_singletons=Truewithmin_recurrencegreater than 2 is refused, because it would keep variants seen in one sample while dropping variants seen in two. Useinclude_singletons=False, or amin_recurrenceof 1 or 2. - Mitochondrial routing recognizes chrMT and the RefSeq mtDNA accession: the mitochondrial check now matches the "chr" prefix as a literal and recognizes chrM, chrMT, MT, M, and NC_012920, so mtDNA is routed to the mitochondrial classifier and an unrelated contig is not misrouted.
- Heteroplasmy data is available in the default mitochondrial run: per-sample allele-depth and allele-fraction fields are now extracted whenever mitochondrial analysis is active, not only when an inheritance or single-sample selection is configured, so the mitochondrial classifier's heteroplasmy axis is populated in the common single-sample run.
Fixed¶
- A maternal no-call no longer produces a de novo mtDNA call: maternal inheritance now carries the three-state parental genotype, so an all-missing maternal genotype resolves to unknown rather than being read as absence. A de novo call requires a confirmed reference genotype in the mother.
- A systematic client error trips the API circuit breaker: a non-retryable 4xx response (a revoked key, a moved endpoint, a schema mismatch) now records a circuit-breaker failure instead of a success, so a run of them opens the breaker rather than leaving a misconfigured upstream looking healthy while every call fails. Successful (2xx and 3xx) responses still reset the breaker as before.
- INFO extraction tolerates a single malformed field: reading a per-record INFO field whose stored value is inconsistent with the VCF header (which pysam signals by raising) now skips that one field instead of dropping the record's entire INFO dictionary, so the remaining INFO fields are preserved.
- SV parser tolerates VCFs that declare only the INFO fields they use: the structural variant parser probes caller-specific END/SVLEN/copy-number/mate fields (END2, CHR2_POS, INSLEN, HOMLEN, CN, ...) to support multiple SV callers. It now checks the VCF header before reading each field, so a file that declares only standard fields parses cleanly instead of aborting. This matches pysam 0.24's behaviour of raising on access to an undeclared INFO key.
Performance¶
- gnomAD absence is cached across runs: when the gnomAD GraphQL API reports a variant as absent (a "Variant not found" response), the remote client now records that absence in the response cache alongside frequency hits. A repeated lookup for an absent variant is served from the cache instead of issuing a fresh rate-limited request, so a validation pass over tens of thousands of variants pays the per-request cost once rather than on every run. Other GraphQL errors (rate limit, timeout, schema drift) remain uncached and are retried on the next run.
[0.18.2] - 2026-09-15¶
Added¶
- Per-population gnomAD allele frequencies (remote backend):
RemoteTabixGnomAD.lookup_batch_populationsreturns the globalAFplus each ancestry-group subfield (AF_afr,AF_amr,AF_asj,AF_eas,AF_fin,AF_nfe,AF_sas) for a batch of variants, indexed per alternate allele on multi-allelic records. Missing or.-valued subfields are omitted from the returned map rather than read as zero. The existing global-onlylookup_batchand itsFrequencyDatabasecontract are unchanged.
Fixed¶
- SV report version field: the structural variant JSON report stamps the installed package version, read from
vartriage.__version__at write time, so the recorded version tracks the release that produced the report.
[0.18.1] - 2026-09-14¶
Performance¶
- Logarithmic gene-model overlap and splice-site checks: the pure-Python interval index answers each overlap query through a max-end segment tree over the start-sorted intervals, visiting only the candidate prefix and pruning subtrees whose maximum end falls at or below the query start (O(log n + k) per query). Splice-site membership is resolved by binary search over precomputed 4 bp donor/acceptor windows rather than a per-call scan of every exon boundary. Annotation now scales with the variant count instead of variants times features; GIAB HG002 chr22 (50,284 variants) drops from tens of minutes to ~20-27 s on the in-memory backend with byte-identical classification output. The segment tree is rebuilt transparently when a chromosome index is restored from an older on-disk cache.
Added¶
- gnomAD genomes remote preset (
gnomad-genomes-v4-grch38): named remote-tabix preset for the public gnomAD genomes v4.1.1 per-chromosome VCFs, alongside the existing exomes preset. Query the genomes release by name via--gnomad-remote gnomad-genomes-v4-grch38orRemoteTabixConfig(gnomad_remote_url=...)over HTTP byte-range. Returns the global allele frequency; per-population subfields are not yet parsed by the remote backend.
[0.18.0] - 2026-08-27¶
Added¶
- VCF quality control (
vartriage/qc/): pre-flight sample QC computed in a single streaming pass before annotation. Metrics: Ti/Tv ratio, het/hom ratio, total variant count, ins/del ratio, and per-chromosome counts. Each metric is validated against assay-specific ranges (wgs,wes,panel) and reported as PASS, WARN, or FAIL with an overall verdict. vartriage qcsubcommand: standalone QC-only mode with no annotation. Supports--sample,--assay-type,--output-json,--strict,--expected-titv, and--expected-het-hom.- QC pipeline gate:
--strict-qchalts the pipeline with exit code 3 when any metric reaches FAIL, before annotation runs.--skip-qcbypasses QC.--assay-typeselects the threshold preset.--expected-titvand--expected-het-homoverride warn-level ranges. - QCConfig: new config dataclass with startup validation and a
[qc]TOML section (~/.vartriage/config.toml) for threshold overrides. CLI values take precedence over TOML. - PipelineConfig.qc: new field wiring QC into
Pipeline.run(). The QC report is exposed via thePipeline.qc_reportproperty. - QC computation uses no additional dependencies: counts are basic arithmetic over variant records; no scipy or numpy required for the QC pass.
- Sample Quality Control section in clinical reports:
clinical-html,clinical-pdf, andclinical-docxnow render the QC metric table under the executive summary. A WARN or FAIL verdict also adds a note to the limitations section.
Fixed¶
- Clinical PDF dependencies: the
clinicalandallextras requireweasyprint>=69.0,<71.0withpydyf>=0.11,<0.13. This pairing carries the WeasyPrint SSRF hardening for the URL fetcher (CVE-2025-68616, addressed in 69.0) and keepswrite_pdfcompatible with the matching pydyf API. The lockfile resolves WeasyPrint to 70.0.
Notes¶
- QC edge cases are handled without false failures: a VCF with no indels or no genotype calls reports the affected ratio but does not count it against the verdict. Single-sample VCFs auto-detect the sample for het/hom.
[0.17.5] - 2026-08-25¶
Added¶
- SpliceAI SQLite backend (
--spliceai-db): query precomputed delta scores directly from the OpenCRAVAT SQLite database (1.7 GB, all chromosomes). Eliminates the need to pre-filter scores into per-analysis TSV files. Returns max(ds_ag, ds_al, ds_dg, ds_dl) per variant. Mutually exclusive with--spliceai-scores(TSV). - PrioritizationConfig.spliceai_db_path: new field with mutual exclusion validation against
spliceai_scores_path. - SpliceAISQLiteLoader class (
vartriage/prioritization/spliceai_db.py): read-only connection, chromosome normalization, batch lookups grouped by chromosome, context manager protocol.
Fixed¶
- Remote tabix timeout enforcement:
pysam.TabixFile.fetch()previously blocked indefinitely on stalled S3 connections (gnomAD). Now wrapped in aThreadPoolExecutorwithfuture.result(timeout=read_timeout). TimeoutError triggers retry with backoff and records circuit breaker failures. Prevents the 6+ hour hang observed during eRepo validation.
Validation Results¶
eRepo (21,506 classified variants, SpliceAI SQLite + remote gnomAD):
| Metric | v0.17.3 (no SpliceAI) | v0.17.5 (SpliceAI SQLite) |
|---|---|---|
| Pathogenic sensitivity | 65.6% | 70.5% |
| Splice-site sensitivity | 9.8% | 55.9% |
| Pathogenic PPV | 99.2% | 99.2% |
| Missense sensitivity | 46.9% | 46.9% |
| Frameshift sensitivity | 99.7% | 99.7% |
[0.17.4] - 2026-08-24¶
Changed¶
- README: updated criteria count from 10 to 13+ (PM1, PM4, BS2, PVS1_Strong added). Updated benchmark table to v0.17.2 GIAB numbers (50,284 variants, 26s, 670 MB). Added validation summary table with eRepo results. Added known limitations section documenting splice sensitivity gap and missing benign criteria.
- docs/pipeline-stages.md: pipeline flow diagram now includes GeneKnowledgeAnnotator and PhenotypeBoost stages. ACMG section updated with PM1, PM4, BS2, PVS1_Strong criteria. Combining rules updated to reflect Bayesian-adapted LP rules (2M=LP, 1M+4Sup=LP). BA1 standalone override documented.
- docs/architecture.md: package tree updated with mito/ (v0.15.0), remote/ (v0.16.0), data/knowledge/, and data/mito/ directories.
- docs/tutorial.md: pipeline flow shows optional GeneKnowledgeAnnotator and PhenotypeBoost stages. Example JSON output uses
prioritization_scoreas primary ranking field;composite_rankmarked deprecated. - docs/configuration.md: PipelineConfig table includes
mitoandremotefields. Added full MitoConfig and RemoteTabixConfig documentation sections.- docs/validation.md: renamed from "GIAB Validation" to "Validation". Added summary table. Added complete ClinVar eRepo validation section with stratified metrics, interpretation, tool comparison, and provenance. - docs/examples/README.md: added sections 13 (mitochondrial, v0.15.0) and 14 (remote tabix, v0.16.0). Updated sample output file descriptions to document
prioritization_score. - docs/examples/sample_pipeline_output.json: added
prioritization_scorefield. Updated evidence tags (PM1, BS1, BP4, BA1, BP7) and classifications (Likely_Benign, Benign) to reflect current criteria. - docs/examples/sample_pipeline_output.csv: added
prioritization_scorecolumn. Same tag and classification updates.
[0.17.3] - 2026-08-23¶
Fixed¶
- PM2 fires on absent-from-gnomAD variants: previously, variants absent from gnomAD (AF=None) were treated as "data unavailable" and PM2 was withheld. Per ACMG/AMP 2015, "absent from controls" satisfies PM2. With gnomAD v4.1.1 covering 730K+ exomes, absence is strong evidence of rarity. This fix improved eRepo pathogenic sensitivity from 17.6% to 65.6%.
- BA1 standalone override: BA1 (AF > 5%) now returns Benign unconditionally before conflict detection, per ACMG Table 5. Previously, BA1 + any pathogenic tag produced VUS (conflicting evidence).
Added¶
- Bayesian-adapted LP rules (Tavtigian et al. 2018): 2 Moderate = Likely Pathogenic, 1 Moderate + 4 Supporting = Likely Pathogenic. These are standard ClinGen-accepted extensions.
- eRepo validation script (
scripts/validate_erepo.py): end-to-end ClinGen Expert Panel concordance validation. Extracts expert-panel variants from ClinVar, filters REVEL, runs the pipeline, reports sensitivity/PPV/per-consequence metrics. - GTF annotation pickle caching: first GTF load parses the file (~70s); subsequent loads deserialize the cache (~2s). Cache validated by source file mtime and size.
- REVEL cache fingerprint validation: filtered REVEL file is only reused when a SHA-256 sidecar matches the current variant set and REVEL source mtime.
Validation Results¶
eRepo (21,928 input variants, 21,614 classified after filtering, gnomAD v4.1.1 remote):
| Metric | Value |
|---|---|
| Pathogenic sensitivity | 65.6% |
| Pathogenic PPV | 99.2% |
| Frameshift sensitivity | 99.7% |
| In-frame deletion | 99.1% |
| In-frame insertion | 100% |
| Missense sensitivity | 46.9% |
| Splice-site sensitivity | 9.8% |
| Benign sensitivity | 6.0% |
[0.17.2] - 2026-08-21¶
Changed¶
- 70x performance improvement: replaced polars per-batch DataFrame joins (O(n*m) per batch) with pre-materialized Python dict lookups (O(1) per variant) for gnomAD and ClinVar annotation backends. chr22 (50,284 variants) now completes in ~26s, down from 30+ minutes.
- Pickle caching for reference data: gnomAD and ClinVar DataFrames are serialized to disk on first load; subsequent runs skip TSV parsing entirely (~0.3s load from cache vs ~30s parse).
Fixed¶
- PyPI publishing: versions 0.10.0-0.17.1 were never published to PyPI due to version mismatch between pyproject.toml and git tags. Fixed version management and added CI gate (tests must pass before publish).
[0.17.1] - 2026-08-15¶
Fixed¶
- BP4 no longer fires for in-frame indels: IN_FRAME_INSERTION and IN_FRAME_DELETION are now excluded from BP4 computational benign evidence. These variants fire PM4 (moderate pathogenic); awarding BP4 simultaneously created logically contradictory evidence and an incorrect conflicting-evidence VUS.
[0.17.0] - 2026-08-15¶
Added¶
- PVS1 LoF constraint gating: PVS1 strength now depends on whether loss-of-function is the established disease mechanism. Genes with pLI > 0.9 get PVS1 at Very Strong; genes with pLI < 0.9 get PVS1_Strong (Strong). An optional
lof_gene_listparameter allows labs to curate an explicit gene list. - PM1 evaluator: fires for missense variants in genes with gnomAD missense constraint (mis_z > 3.09), indicating functional domain intolerance.
- PM4 evaluator: fires for in-frame insertions, in-frame deletions, and stop-loss variants.
STOP_LOSSconsequence type added toFunctionalConsequenceenum.PVS1_STRONGevidence tag with Strong strength for downgraded PVS1 in LoF-tolerant genes.has_conflicting_evidencefield onClassifiedVariant: True when both pathogenic and benign tags are present. Distinguishes "VUS due to insufficient evidence" from "VUS due to conflicting evidence."2 VS = Pathogeniccombining rule per ACMG 2015 Table 5.- Limitations documentation:
docs/acmg-criteria.mdnow includes an explicit table of evaluated vs. placeholder criteria, documented limitations (PVS1 scope, multiallelic handling, genome build, haplogroup context, PP5 model).
Changed¶
- CADD normalization unified:
compute_prioritization_scorenow divides by 99.0 (matchingnormalize_cadd_scores) instead of the previous inconsistent 60.0. Variants with CADD Phred 35 now get a prioritization score of 0.354 (was 0.583). This may change relative rankings for variants where CADD is the only available score. - Mitochondrial Rule 4: now requires MITOMAP
confirmedstatus for Likely Pathogenic classification. Previously, any MITOMAP entry (including unconfirmed reports) triggered LP. Unconfirmed entries now classify as VUS with a reason noting "requires validation." ACMGClassifierconstructor accepts optionallof_gene_list: frozenset[str]parameter._assign_tagsnow evaluates PM1 and PM4 in addition to existing criteria.
Migration Notes¶
- Breaking for CADD-dependent rankings: variants scored primarily by CADD will shift in prioritization order. Re-run existing analyses to see the impact. Classification (ACMG tiers) is not affected since CADD drives PP3/BP4 via REVEL thresholds, not directly.
- PVS1 strength change: variants in genes with pLI < 0.9 that previously received PVS1 (Very Strong) now receive PVS1_Strong (Strong). This may downgrade some variants from Pathogenic to Likely Pathogenic when PVS1 was the only very-strong evidence. Users who don't provide gene-disease linkage data (no
--knowledge-diror bundles) see no change (benefit of the doubt applies). - Mito LP reclassification: variants previously classified LP based on unconfirmed MITOMAP entries are now VUS.
[0.16.0] - 2026-08-14¶
Added¶
- Remote tabix CADD backend: query CADD Phred scores from the 80+ GB pre-scored file via HTTP byte-range requests using
pysam.TabixFile. No local download required. Activate with--cadd-remote cadd-v1.7-grch38or pass a full URL. - Remote tabix gnomAD backend: query gnomAD allele frequencies from per-chromosome tabix-indexed VCFs hosted on public cloud storage. Activate with
--gnomad-remote gnomad-exomes-v4-grch38. - Named presets: 4 bundled presets (
cadd-v1.7-grch38,cadd-v1.7-grch37,cadd-v1.7-indels-grch38,gnomad-exomes-v4-grch38) so users don't need to memorize download URLs. List withvartriage remote list-presets. - Score cache (
~/.vartriage/remote_cache.db): SQLite cache stores fetched scores locally with configurable TTL (default 30 days). Pinned mode (--remote-cache-ttl -1) for clinical reproducibility. - Batch query optimization: variants within 10 kb on the same chromosome are grouped into a single range query, reducing HTTP round-trips for gene panels and WGS.
- Circuit breaker: after 5 consecutive network failures within 60 seconds, remote queries are disabled for the remainder of the run. The pipeline continues with scores marked as unavailable rather than crashing.
- CLI flags:
--cadd-remote,--gnomad-remote,--remote-cache-ttl. vartriage remote list-presetssubcommand: displays all named presets with source, genome build, and description.- Priority rules: local file (
--cadd-scores) always takes precedence over remote tabix (--cadd-remote), which takes precedence over REST API (--mode api). RemoteTabixConfigdataclass: configures URLs, cache path/TTL, connect/read timeouts, retry count, and batch window size.- User guide:
docs/remote-tabix.mdwith CLI examples, Python API, caching details, and configuration reference.
Changed¶
AnnotationConfig.gnomad_pathis now optional (Path | None). WhenNoneand remote gnomAD is configured, the pipeline injects the remote backend automatically.AnnotationEnginegainedset_frequency_db()andhas_frequency_dbfor runtime backend injection.PrioritizationEngineaccepts an optionalremote_configparameter. When no local CADD file is available, it usesRemoteTabixCADDfor batch lookups.PipelineConfiggained aremote: RemoteTabixConfig | Nonefield.
Performance Targets¶
| Workload | Variants | Target (cold cache) | Target (warm cache) |
|---|---|---|---|
| Gene panel | 500 | < 30 seconds | < 5 seconds |
| chr22 WGS | 42,000 | < 10 minutes | < 10 seconds |
Migration Notes¶
- No existing CLI flags or config fields changed behavior. Users who don't specify
--cadd-remoteor--gnomad-remotesee zero difference. - No new pip dependencies.
pysamalready supports remote tabix URLs via HTSlib's libcurl integration. Cache uses stdlibsqlite3.
[0.15.0] - 2026-08-07¶
Added¶
- Mitochondrial variant support: automatic detection and classification of chrM/MT variants using mtDNA-specific criteria separate from the nuclear ACMG/AMP 2015 framework.
- Mitochondrial genetic code: vertebrate mitochondrial codon table (4 differences from standard). CodonResolver auto-selects the correct table based on chromosome.
- Heteroplasmy extraction: computes alternate allele fraction from AD or AF fields with 5-level classification (homoplasmic/high/moderate/low/sub-threshold).
- MITOMAP database: bundled pathogenic mutations (~60 entries) with disease association and confirmation status lookups.
- HelixMTdb frequency: bundled population allele frequencies (~125 entries) for distinguishing haplogroup markers from rare variants.
- MT gene map: all 37 mitochondrial genes with coordinate-based annotation (protein_coding/tRNA/rRNA/control_region/intergenic).
- MitochondrialClassifier: rule-based classification (Pathogenic/Likely Pathogenic/VUS/Likely Benign/Benign) using MITOMAP confirmation, heteroplasmy level, and population frequency.
- Maternal inheritance check: when trio data is available, verifies maternal transmission and flags potential de novo mtDNA mutations.
- CLI flags:
--skip-mitoto bypass mitochondrial analysis,--mt-min-heteroplasmyto set the reporting threshold (default 1.0%). - JSON/CSV output: mitochondrial findings section with heteroplasmy_level, mitomap_disease, mt_classification, and gene context fields.
- Clinical report: "Mitochondrial Findings" section with variant table and methodology note about mtDNA-specific criteria.
- Data update scripts:
scripts/download_mitomap.pyandscripts/download_helixmtdb.pyfor refreshing bundled reference data.
Changed¶
Pipeline.run()now splits chrM/MT variants to a dedicated MitochondrialPipeline and merges results at the report stage.PipelineConfiggained amitofield (optionalMitoConfig).- VCFParser sample extraction is now enabled when mitochondrial analysis is active (needed for heteroplasmy AD field access).
[0.14.0] - 2026-08-02¶
Added¶
- PS1 criterion (Strong pathogenic): fires when a different nucleotide change at the same codon produces the same amino acid substitution as a known ClinVar Pathogenic variant. Requires a ClinVar protein index file and reference FASTA for codon resolution.
- PM5 criterion (Moderate pathogenic): fires when a novel missense change occurs at an amino acid position where a different pathogenic missense is established in ClinVar. Shares the same protein index as PS1.
- PP3_Moderate and BP4_Moderate evidence tags: ClinGen SVI strength-modulated computational evidence (Pejaver et al. 2022). PP3 now fires at two strength levels; BP4 fires at two strength levels.
- gnomAD GraphQL API client (
vartriage.api.gnomad_client.GnomADClient): queries gnomAD v4 directly for per-population allele frequencies. ReturnsPopulationFrequencieswith all 7 ancestry groups. Cached in the same SQLite database as other API responses. - ClinVar protein index (
vartriage.annotation.clinvar_protein_index): loads a TSV of ClinVar pathogenic missense variants keyed by (gene, amino acid position) for O(1) PS1/PM5 lookups. ProteinChangedataclass onAnnotatedVariant: stores resolved amino acid substitution (gene, position, ref_aa, alt_aa) from codon resolution. Enables PS1/PM5 without re-running codon analysis at classification time.
Changed¶
- PP3 threshold updated to ClinGen-calibrated values: REVEL > 0.644 (supporting), REVEL > 0.773 (moderate). Previously used REVEL > 0.7 (supporting only).
- BP4 threshold updated to ClinGen-calibrated values: REVEL < 0.290 (supporting), REVEL < 0.183 (moderate). Previously used REVEL < 0.15 (supporting only).
- Benign combining rules now handle moderate-level benign evidence: 1 BS + 1 BM = Likely Benign, 2 BM = Likely Benign, 1 BM + 2 BP = Likely Benign.
ACMGClassifierconstructor accepts an optionalprotein_indexparameter for PS1/PM5 evaluation.EvidenceTagenum expanded withPS1,PM5,PP3_MODERATE,BP4_MODERATEmembers.EVIDENCE_STRENGTH_MAPupdated with strength assignments for all new tags.
Migration Notes¶
- Breaking for threshold-sensitive analyses: variants with REVEL between 0.644-0.7 now receive PP3 (previously they didn't). Variants with REVEL between 0.15-0.290 now receive BP4 (previously they didn't). This means some variants will shift classification. Re-run existing analyses to see the impact.
- PS1/PM5 are additive: they fire only when a protein index is provided. Without it, behavior is identical to 0.13.0 (the criteria are simply omitted).
- The gnomAD API client is optional: it's available under
pip install vartriage[api]and activates only when explicitly used.
[0.13.0] - 2026-07-31¶
Added¶
- Structural variant triage pipeline (
vartriage/structural/): complete SV analysis from VCF parsing through ClinGen-based classification. SVParser: streams SV records from VCF via pysam. Handles SVTYPE, END/SVLEN, CIPOS/CIEND, BND bracket notation, copy number. Supports Manta, DELLY, GATK-SV, GRIDSS, and LUMPY INFO field conventions.SVAnnotator: determines gene overlap (whole-gene vs partial), attaches ClinGen dosage sensitivity (HI/TS scores), and matches gnomAD-SV population frequencies via reciprocal overlap.SVScorer: composite pathogenicity score from gene impact (35%), dosage sensitivity (30%), population frequency rarity (20%), and SV size (15%).SVClassifier: implements ACMG/ClinGen Technical Standards for CNV interpretation (Riggs et al. 2020). Evaluates Sections 1-4, maps accumulated evidence to 5-tier classification.SVTriagePipeline: orchestrates all stages, validates config at construction, writes JSON/CSV reports.SVReportBuilder: clinical report section with findings table, summary stats, and per-SV evidence narratives.merge_findings(): combines SNV ClassifiedVariants + SV ClassifiedSVs into unified ranked output.- CLI
vartriage svsubcommand with--sv-vcf,--min-sv-size,--max-sv-size,--sv-types,--dosage-sensitivity,--gnomad-sv,--pathogenic-regions,--benign-regions,--reciprocal-overlap,--whole-gene-threshold,--include-benignflags. - Integrated SV mode:
--sv-vcfflag on the mainvartriagecommand runs both SNV and SV pipelines, producing separate output files. - Syndrome name resolution: pathogenic region matches return the associated syndrome name (e.g., "22q11.2 deletion syndrome") in classification output.
- Bundled reference data:
clingen_dosage.tsv(54 genes),clingen_pathogenic_regions.bed(16 syndromes),clingen_benign_regions.bed(18 regions). - gnomAD-SV bundle:
gnomad-sventry in the bundle registry for automated download. - Download script:
scripts/download_clingen_dosage.pyfetches latest ClinGen dosage data from FTP. SVTriageConfigfrozen dataclass with startup validation for all parameters.- Data models:
StructuralVariant,AnnotatedSV,ScoredSV,ClassifiedSV,GeneOverlap,Breakpoint. - Enums:
SVType(6 types),SVConsequence(8 levels),SVClassification(5 tiers),SVEvidenceCategory(ClinGen sections 1-5). - 60 new unit/integration tests covering parser, annotator, frequency matching, classifier, and end-to-end 22q11.2 deletion scenario.
docs/structural-variants.mduser guide with CLI reference, Python API, and output format documentation.
[0.12.0] - 2026-07-29¶
Added¶
- Gene-disease linkage knowledge base (
vartriage/knowledge/): connects variants to their clinical context via gene-level annotations from OMIM, ClinGen, HPO, and gnomAD constraint data. OMIMDatabase: gene-disease associations with MIM numbers and inheritance modes (AD, AR, XL, XLD, XLR, MT).HPODatabase: gene-to-HPO phenotype term mappings for overlap scoring.ClinGenValidityDB: gene-disease validity levels (Definitive, Strong, Moderate, Limited, Disputed, Refuted).ConstraintDB: gnomAD constraint metrics (pLI, LOEUF, mis_z) per gene.ActionabilityDB: ClinGen actionability curations with intervention types.GeneKnowledgeRegistry: flyweight-cached composite lookup (O(1) per gene). Owns phenotype overlap scoring internally.GeneKnowledgeAnnotator: pipeline stage enriching variants withGeneContext. Filters by inheritance mode and actionability. Providesboost_scores()for prioritization.apply_phenotype_boost(): stateless utility applyingscore * (1 + overlap)bounded at 2.0.- Phenotype-driven prioritization (
--hpo-terms HP:0001250,HP:0001249): computes overlap between patient HPO terms and each gene's phenotype annotations. Variants in phenotype-matching genes receive a multiplicative score boost (1.0-2.0) applied toprioritization_scorebetween scoring and ACMG classification. Tier isolation preserved: boost operates within classification tiers only. - Inheritance mode filtering (
--inheritance-mode AD|AR|XL|XLD|XLR|MT): filters variants to genes matching the specified OMIM inheritance pattern. Intergenic variants and genes without OMIM data pass through unfiltered. - Actionability filtering (
--flag-actionable): when set, only variants in ClinGen-curated medically actionable genes (plus intergenic) pass through. Actionability annotation is always populated when knowledge base is active. - Gene context in output (JSON and CSV):
disease_associations(with MIM numbers and inheritance mode),clingen_validity,gene_constraint(pLI/LOEUF/mis_z),is_actionable, andphenotype_match_scorefields added to classified variant output. - Custom knowledge directory (
--knowledge-dir): override the bundled data with custom pre-processed TSV files. - Bundled knowledge data: pre-processed TSV files for 22 clinically relevant genes (BRCA1, BRCA2, TP53, CFTR, SCN1A, MECP2, NF2, etc.) shipped as package data.
KnowledgeBaseConfigdataclass with HPO term format validation and inheritance mode validation.GeneContextdataclass attached toAnnotatedVariantfor variant-facing gene knowledge.gene_contextfield onAnnotatedVariant.- Spike-in validation script (
scripts/validate_affected_patient.py): injects pathogenic NF2 variants into GIAB HG002 and validates end-to-end gene-disease linkage. - 73 new tests: 56 unit tests for all knowledge modules + 17 clinical scenario integration tests (Dravet syndrome patient simulation).
docs/gene-disease-linkage.mduser guide.
Changed¶
PipelineConfiggainsknowledge: KnowledgeBaseConfig | Nonefield.GeneKnowledgeAnnotatoris constructed once atPipeline.__init__(not per-run) to avoid repeated TSV reloads.- CSV
disease_associationscolumn now includes MIM numbers:"Disease name [MIM:123456] (AD)". - Pipeline stage order:
AnnotationEngine → [GeneFilter] → [GeneKnowledgeAnnotator] → PrioritizationEngine → [PhenotypeBoost] → ACMGClassifier. - Package version bumped to 0.12.0.
[0.11.1] - 2026-07-28¶
Documentation¶
- Document CADD score normalization: composite uses Phred/99.0, prioritization_score uses Phred/60.0, both capped at 1.0.
- Add ACMG evidence strength column (Very Strong, Moderate, Supporting, Standalone, Strong) to classification table.
- Clarify cohort parallel mode uses ThreadPoolExecutor with GIL-releasing pysam I/O.
- Add API mode Python snippet showing programmatic
APIConfig.load()usage.
[0.11.0] - 2026-07-19¶
Added¶
- Multi-sample cohort analysis (
vartriage cohort): process multiple VCF files together to find shared variants, compute cross-sample recurrence frequencies, and produce per-gene burden tables. CohortConfigdataclass with validation (min 2 samples, AF threshold, recurrence threshold, parallel options).CohortAggregator: merges classified variants by coordinate(chrom, pos, ref, alt)across samples. Filters by population AF before aggregation. Resolves most-severe classification and consequence when a variant appears in multiple samples.CohortStatistics: per-gene burden (variant count, pathogenic count, penetrance), recurrence distribution, per-sample counts, classification and consequence breakdowns.CohortReportGenerator: writes three output files per run (variants, gene burden, summary) in JSON or CSV format.CohortPipeline: orchestrates per-sample Pipeline execution then aggregates. Supports sequential and parallel (ThreadPoolExecutor) processing modes.- CLI subcommand
vartriage cohortwith--manifest(tab-separated, optional labels column) or--vcf(multiple paths). Options:--min-recurrence,--max-af,--no-singletons,--parallel,--max-workers,--cohort-name,--output-format. Accepts all standard reference file flags. CohortVariantdataclass withcohort_frequency,is_singleton,is_universalcomputed properties.GeneBurdendataclass withpenetranceproperty (fraction of cohort samples affected).CohortSummarydataclass aggregating top-level cohort metrics.- Public API exports:
CohortPipeline,CohortAggregator,CohortReportGenerator,CohortStatistics,CohortConfig,CohortVariant,CohortSummary,GeneBurden.
Changed¶
- Package version bumped to 0.11.0.
__init__.pyexports expanded with all cohort public classes.- README updated with cohort CLI usage, Python API examples, and
CohortConfigreference table.
[0.10.0] - 2026-07-18¶
Added¶
- Zygosity model:
Zygosityenum (HETEROZYGOUS, HOMOZYGOUS_ALT, HEMIZYGOUS, UNKNOWN) andzygosityfield onAnnotatedVariant. Ready for population from VCF FORMAT GT field. - Variant quality metrics:
VariantQualityMetricsdataclass (depth, genotype_quality, allele_balance, is_low_confidence) andquality_metricsfield onAnnotatedVariant. - ACMG Secondary Findings (SF v3.2): shipped 71-gene list as package data (
vartriage/data/acmg_sf_v3.2.txt).SecondaryFindingsFilterclass withis_secondary_finding()andsplit_stream(). CLI flag--secondary-findings. - Computational-only disclaimer: clinical reports now display a prominent banner after the header citing ACMG/AMP 2015 (Richards et al.) and stating that all findings require review by a clinical geneticist.
Notes¶
- Zygosity extraction from VCF FORMAT fields and quality metrics population planned for v0.10.1 (requires VCFParser changes for FORMAT field access).
- HGVS nomenclature generation planned for v0.10.1 (requires CodonContext pipeline integration).
[0.9.0] - 2026-07-17¶
Added¶
- Benign evidence criteria: the pipeline can now classify variants as Benign or Likely Benign for the first time.
- BA1: any gnomAD population AF > 5% = standalone Benign
- BS1: any population AF > 1% = strong benign evidence
- BP4: low computational predictor scores (REVEL < 0.15 for missense, CADD Phred < 10 for non-missense)
- BP7: synonymous variant with SpliceAI < 0.1 (no splice impact)
- Population-specific frequency model (
PopulationFrequenciesdataclass): stores per-population gnomAD AFs (AFR, AMR, ASJ, EAS, FIN, NFE, SAS) withmax_population_af,any_exceeds(), andall_below()helpers. - Updated combining rules (ACMG Table 5): BA1 standalone = Benign, 2 BS = Benign, 1 BS + 1 BP = Likely Benign, conflicting pathogenic + benign = VUS.
STANDALONEstrength tier inEvidenceStrengthenum for BA1.- 9 new
EvidenceTagenum members: PS1, PM1, PM4, PM5 (pathogenic, evaluators planned for v0.9.1), BA1, BS1, BS2, BP4, BP7 (benign, evaluators active). has_conflicting_evidence()helper in combining module.population_frequenciesfield onAnnotatedVariant.- 24 unit tests for benign criteria covering positive/negative/edge cases.
Changed¶
- PM2 now checks ALL population-specific frequencies below threshold (not just global AF). Falls back to global AF when per-population data is absent.
EVIDENCE_STRENGTH_MAPexpanded from 4 to 13 entries covering all new tags.combine_evidence()rewritten: separates pathogenic and benign evidence, detects conflicts, applies benign combining rules.- README links converted to absolute GitHub URLs (fixes dead links on PyPI).
Notes¶
- PS1, PM1, PM4, PM5 evaluators are not yet implemented (tags exist in the enum, evaluators planned for v0.9.1 when ClinVar amino acid index and functional domain data are available).
- BS2 tag exists but evaluator requires gnomAD homozygote count data not currently parsed (planned for v0.9.1).
[0.8.0] - 2026-07-16¶
Added¶
- Codon-level consequence calling (
--reference-fasta): SNVs in CDS regions now use actual amino acid comparison instead of the positional heuristic. Requires an indexed reference FASTA. Correctly distinguishes synonymous, missense, and nonsense changes by extracting the reference codon, substituting the variant base, and translating both codons. - Variant normalization: left-align and trim indels before gnomAD/ClinVar/score lookups using the reference FASTA (Tan et al. 2015 algorithm). Reduces silent lookup failures caused by representation differences between VCF callers and reference databases.
- Prioritization score (
prioritization_scoreoutput field): literature-backed scoring using REVEL directly for missense (validated 0.7 threshold), SpliceAI for splice-adjacent, CADD Phred/60 for non-missense. Replaces the unvalidated 0.4/0.6 weighted average as the recommended ranking method. TranscriptCDSIndex: per-transcript CDS exon map built from GTF, enabling genomic-to-CDS position mapping for forward and reverse strand genes.CodonResolver: FASTA-backed codon extraction with split-codon-at-exon-junction handling and negative-strand reverse complement.VariantNormalizer: three-step normalization (right-trim, left-trim, left-align) with 1000-iteration safety cap.compute_prioritization_score()function in scoring module.- Standard genetic code table (
_internal/genetic_code.py) withtranslate_codon()andreverse_complement(). - 36 unit tests covering genetic code, transcript index, normalizer, and prioritization score.
- 3 live VEP concordance integration tests (marked
@pytest.mark.slow).
Changed¶
ScoredVariantnow carries bothcomposite_rank(legacy) andprioritization_score(new). Both synced via__post_init__when only one is set.- JSON output includes
prioritization_scorefield alongsidecomposite_rank. _determine_consequence()accepts optionalcodon_resolverparameter; uses it for SNVs in CDS when available, falls back to positional heuristic without FASTA.AnnotationConfiggainsreference_fasta_path: Optional[Path]field.- GTF parsing now populates a
TranscriptCDSIndexfrom CDS features (frame column extracted).
Notes¶
- Without
--reference-fasta, all behavior is unchanged from v0.7.0 (backward compatible). - The positional heuristic ("SNV in CDS = Missense") remains as a fallback. To get correct consequence calling, provide the reference FASTA.
composite_rankis deprecated in favor ofprioritization_score. Both are present in output for backward compatibility.composite_rankwill be removed in v1.0.0.
[0.7.0] - 2026-07-14¶
Added¶
- API annotation mode (
--mode api|hybrid): query Ensembl VEP, ClinVar E-utilities, CADD, and SpliceAI via HTTP instead of local reference files. Zero-config variant triage for gene panels and exploratory analysis. - Ensembl VEP client: batch POST annotation (200 variants/request) with consequence, gene symbol, gnomAD frequency, and CADD Phred extraction. VEP Sequence Ontology terms mapped to FunctionalConsequence enum across 14 severity ranks. GRCh37 and GRCh38 support.
- ClinVar E-utilities client: esearch + esummary lookups with clinical significance mapping, review status ranking for conflicting interpretations, and NCBI API key support for higher rate limits.
- CADD REST API client: position-based Phred score lookups with ref/alt allele filtering. Used as fallback when VEP's CADD plugin has no pre-computed score.
- SpliceAI Lookup client: max delta score extraction with consequence-based smart filtering (only queries splice-relevant variants to conserve rate limit).
- API annotation engine: composes VEP + ClinVar into the same
annotate()interface as the local engine. Concurrent ClinVar lookups via ThreadPoolExecutor. - API score provider: CADD score hierarchy (VEP plugin first, standalone API fallback) and SpliceAI lookups. Documents REVEL limitation (no API exists).
- Response caching: SQLite-backed at
~/.vartriage/api_cache.dbwith configurable TTL (7 days default, 30 days for ClinVar). Pinned mode (cache_ttl_days = -1) for clinical reproducibility. - Resilience stack: token bucket rate limiter (per-service with daily caps), circuit breaker (CLOSED/OPEN/HALF_OPEN with 60s recovery), exponential backoff retry (1s/2s/4s, max 3), Retry-After header parsing.
- VCF-to-VEP notation converter: handles SNVs, deletions, insertions, MNVs, complex indels, and chr prefix stripping.
- CLI flags:
--mode(local/api/hybrid),--api-key(NCBI),--no-confirm(skip large-run prompt). - HTTP proxy support via
HTTPS_PROXY/HTTP_PROXYenv vars and[api.proxy]TOML config. - User-Agent header on all requests identifying vartriage version and project URL.
docs/api-mode.mduser guide.pip install vartriage[api]optional dependency group (httpx).
Changed¶
PipelineConfigaccepts an optionalapifield for API backend configuration.Pipeline.run()routes toAPIAnnotationEnginewhen API mode is active.pyproject.toml: addedapiand updatedalloptional dependency groups.
[0.6.0] - 2026-07-13¶
Added¶
- Score bundle downloader (
vartriage bundle): automated downloading, transformation, and management of reference files. Supported bundles: ClinVar, gnomAD exomes (chr22), REVEL, GENCODE, SpliceAI. Subcommands:download,list,verify,status,update-registry. - Bundle-aware pipeline (
--use-bundles,--genome-build): auto-resolve reference file paths from installed bundles. Explicit CLI paths take precedence. No configuration needed aftervartriage bundle download. - HTTP download engine with resume support (Range header), atomic
.partialfiles, exponential backoff retry (1s, 2s, 4s) for 429/5xx errors, and streaming progress bar on stderr. - Post-download transformers: VCF-to-TSV (bcftools with pysam fallback), ClinVar normalizer (adds
chrprefix + CLNSIG mapping), CSV-to-TSV (REVEL), SpliceAI max-delta extractor, and passthrough (gz decompression). - Per-bundle manifests tracking version, checksums, download timestamps, and source URLs.
- Configurable storage layout at
~/.vartriage/bundles/with env var override (VARTRIAGE_BUNDLE_STORAGE). - TOML configuration file support (
~/.vartriage/config.toml) for default build, concurrency, and proxy settings. - Disk space pre-flight validation before downloads.
- SHA-256 checksum verification for downloaded and transformed files.
use_bundlesandgenome_buildfields onPipelineConfig.docs/bundles.mduser guide with command reference and storage layout documentation.
Changed¶
pyproject.toml: addedtomliandtomllibto mypy ignore list for Python 3.10 compatibility.
[0.5.0] - 2026-07-13¶
Added¶
- Clinical report generation (
--output-format clinical-html|clinical-pdf|clinical-docx): produces structured, sign-off-ready clinical variant reports from ClassifiedVariant data. Per-variant evidence narratives with plain-language ACMG criteria explanations. Structured sections: Header, Executive Summary, Findings Table, Evidence Cards, Limitations, Methodology, Sign-off. JSON audit trail sidecar (.audit.json) with run manifest and per-variant decision log. - CLI flags
--patient-idand--panel-namefor clinical report metadata. Required when using anyclinical-*output format. ClinicalReportConfigdataclass:patient_id,panel_name,output_format,report_template.- Self-contained HTML output (all CSS inlined, no JavaScript, no external dependencies).
- PDF rendering via WeasyPrint (optional dependency). DOCX rendering via python-docx (optional dependency).
- Atomic file writing for clinical reports (temp file + rename).
- Findings table sorted by classification tier (Pathogenic > Likely_Pathogenic > VUS), then by composite rank descending within each tier.
- Empty variant set produces a valid report with zero counts and a negative-finding statement.
Dependencies¶
weasyprint(optional, forclinical-pdf)python-docx(optional, forclinical-docx)
[0.4.0] - 2026-07-12¶
Added¶
- Trio-based inheritance pattern classification (
--proband,--mother,--father): classifies variants into Mendelian inheritance patterns (de novo, dominant, recessive, compound heterozygous, X-linked) based on family genotypes. Multi-label: a single variant can carry multiple patterns when criteria overlap. Compound het uses gene-aware buffering to detect trans inheritance. Replaces SampleExtractor when trio mode is active. Mutually exclusive with--sample. - SpliceAI score integration (
--spliceai-scores): third pathogenicity predictor alongside CADD and REVEL. Dynamic weight redistribution in the composite formula (0.5/0.3/0.2 when all three present, proportional fallback for any two, single-score identity). PP3 now also fires when SpliceAI > 0.5 on splice-site or missense variants. PVS1 fires for SPLICE_SITE + SpliceAI > 0.8. Fully backward-compatible: existing two-score behavior unchanged when SpliceAI is not configured. - VCF output format (
--output-format vcf): produces an annotated bgzipped VCF (.vcf.gz) with a tabix index (.tbi). Re-reads the source VCF, injects VARTRIAGE_CONSEQUENCE, VARTRIAGE_AF, VARTRIAGE_RANK, VARTRIAGE_ACMG, and VARTRIAGE_TAGS INFO fields for classified variants, and writes all records (matched or not) to output. Directly loadable in IGV and queryable with bcftools. - Gene list filtering (
--gene-list): restrict analysis to variants in a user-supplied gene panel file. One gene symbol per line, case-insensitive matching. Genes in the list with zero matching variants produce a logged WARNING so you can catch typos or outdated nomenclature. gene_namefield onAnnotatedVariantpopulated during consequence annotation.
[0.3.0] - 2026-07-12¶
Added¶
- BED-based region filtering (
--regions): target analysis to specific genomic intervals from a BED file. - Multi-sample VCF support (
--sample,--min-gq): extract a single sample from multi-sample VCFs with optional genotype quality filtering. RegionFilterConfig,SampleConfigconfiguration dataclasses.GeneFilterConfig,RegionFilterConfig,SampleConfigconfiguration dataclasses.- VCF output format (
--output-format vcf): produces an annotated bgzipped VCF (.vcf.gz) with a tabix index (.tbi). Re-reads the source VCF, injects VARTRIAGE_CONSEQUENCE, VARTRIAGE_AF, VARTRIAGE_RANK, VARTRIAGE_ACMG, and VARTRIAGE_TAGS INFO fields for classified variants, and writes all records (matched or not) to output. Directly loadable in IGV and queryable with bcftools. - SpliceAI score integration (
--spliceai-scores): third pathogenicity predictor alongside CADD and REVEL. Dynamic weight redistribution in the composite formula (0.5/0.3/0.2 when all three present, proportional fallback for any two, single-score identity). PP3 now also fires when SpliceAI > 0.5 on splice-site or missense variants. PVS1 fires for SPLICE_SITE + SpliceAI > 0.8. Fully backward-compatible: existing two-score behavior unchanged when SpliceAI is not configured.
[0.2.0] - 2026-07-10¶
Added¶
- Vectorized batch annotation in
PyRangesConsequenceAnnotator:assign_batchnow performs a single PyRanges join instead of per-variant loops. 680x speedup (7 variants/sec to 4,749 variants/sec). - Reference file caching (
_internal/cache.py): GTF interval trees, CADD score dicts, and REVEL score dicts are serialized to pickle after first parse. Subsequent runs skip parsing entirely. Cache invalidation uses source file mtime and a version stamp. Writes are atomic (write-to-temp then rename). Cache files sit next to their source with a.vartriage.cachesuffix. - TabixFrequencyDatabase (
annotation/frequency_tabix.py): queries bgzipped+tabix-indexed gnomAD VCFs on the fly via pysam. Zero memory footprint for the reference file. Auto-selected whengnomad_pathends with.vcf.bgzor.vcf.gz. - Extension-based backend routing in
AnnotationEngine:.vcf.bgz/.vcf.gzpaths route to the tabix backend;.tsv/.tsv.gzpaths use the existing polars/dict backends.
Performance¶
- chr22 full annotation benchmark (130K variants + GENCODE + 4.8M gnomAD entries): 36.3s wall time, ~2 GB peak RSS
- With 100K gnomAD subset: 19.5s wall time, 453 MB peak RSS
[0.1.1] - 2026-07-10¶
Fixed¶
- Frequency loader now treats
'.'in the af column as null instead of raising a parse error (gnomAD files use'.'for missing values)
Added¶
- CLI entry point:
vartriage --vcf ... --output ...with full argument parsing - Streaming report generation: JSON and CSV write directly from iterators, no full-variant buffering
ScoreLoaderclass for loading CADD/REVEL TSV files into coordinate-keyed dictsVarTriageWarningbase class for the warning hierarchy;ScoreValidationWarningandMissingDataSummaryWarninginherit from itpy.typedmarker for PEP 561 compliance- Typed Protocol return values in annotation engine (removed all
Anyfrom return annotations) assign_batchmethod on theIntervalIndexProtocol- Upper bounds on all dependency version specs (pysam <1.0, numpy <3.0, polars <2.0, etc.)
- GitHub Actions CI workflow (Python 3.10, 3.11, 3.12)
- Automated PyPI publishing via trusted publisher workflow
CONTRIBUTING.mdLICENSEfile (MIT)
[0.1.0] - 2026-07-09¶
Initial release on PyPI.
Added¶
- VCF parsing via pysam with streaming iteration
- Quality filtering on FILTER field and QUAL score threshold
- Annotation engine with auto-detected backends:
- Functional consequence via GTF/GFF gene models
- Population frequency via gnomAD
- Clinical significance via ClinVar
- Prioritization: allele frequency gate + composite CADD/REVEL scoring
- ACMG/AMP 2015 evidence classification (PVS1, PM2, PP3, PP5)
- Report generation in JSON, CSV, and PDF
- Protocol-based backend system with pure-Python fallbacks
- Memory-bounded processing for whole-genome scale data (4M+ variants under 2 GB RSS)
- Optional extras:
[accelerated](polars, pyranges),[pdf](reportlab),[all]