Configuration¶
All configuration classes are frozen dataclasses. They validate parameters at construction and raise ValueError for out-of-range values.
QualityFilterConfig¶
Controls which variants pass the quality gate.
| Field | Type | Default | Valid range |
|---|---|---|---|
min_qual |
float |
20.0 |
0 to 1,000,000 |
from vartriage import QualityFilterConfig
# Default: QUAL >= 20
config = QualityFilterConfig()
# Stringent clinical threshold
config = QualityFilterConfig(min_qual=50.0)
AnnotationConfig¶
Paths to reference data and batch size for the annotation engine.
| Field | Type | Default | Valid range / notes |
| ------- | ------ | --------- | --------- | ------------------- |
| gene_annotation_path | Path | required | GTF or GFF file |
| gnomad_path | Path | required | gnomAD TSV or tabix VCF (.vcf.bgz/.vcf.gz) |
| clinvar_path | Optional[Path] | None | ClinVar TSV, or None to skip |
| reference_fasta_path | Optional[Path] | None | Indexed FASTA (.fa + .fai) for codon-level consequence calling |
| batch_size | int | 10_000 | 1,000 to 100,000 |
from pathlib import Path
from vartriage import AnnotationConfig
config = AnnotationConfig(
gene_annotation_path=Path("gencode.v44.gtf"),
gnomad_path=Path("gnomad.v4.sites.tsv"),
clinvar_path=Path("clinvar_20240101.tsv"),
reference_fasta_path=Path("GRCh38.fa"), # enables codon-level consequence calling
batch_size=50_000, # larger batches for faster processing
)
PrioritizationConfig¶
Controls frequency filtering and pathogenicity scoring.
| Field | Type | Default | Valid range |
|---|---|---|---|
max_allele_frequency |
float |
0.01 |
0.0 to 1.0 |
cadd_scores_path |
Optional[Path] |
None |
CADD Phred TSV |
revel_scores_path |
Optional[Path] |
None |
REVEL scores TSV |
spliceai_scores_path |
Optional[Path] |
None |
SpliceAI scores TSV |
batch_size |
int |
10_000 |
1,000 to 100,000 |
from vartriage import PrioritizationConfig
# Rare disease: very stringent frequency cutoff
config = PrioritizationConfig(
max_allele_frequency=0.0001,
cadd_scores_path=Path("cadd_v1.7.tsv"),
revel_scores_path=Path("revel_v1.3.tsv"),
spliceai_scores_path=Path("spliceai_scores.tsv"),
)
# Research: relaxed frequency, no score files
config = PrioritizationConfig(
max_allele_frequency=0.05,
)
ReportConfig¶
Output format selection.
| Field | Type | Default | Options |
|---|---|---|---|
output_format |
Literal[...] |
"json" |
"json", "csv", "pdf", "vcf", "clinical-html", "clinical-pdf", "clinical-docx" |
When the format is clinical-html, clinical-pdf, or clinical-docx, the pipeline constructs a ClinicalReportConfig instead and delegates to the clinical report generator. See the ClinicalReportConfig section below.
InheritanceConfig¶
Trio-based inheritance pattern classification settings.
| Field | Type | Default | Notes |
| ------- | ------ | --------- | ---- | ------- |
| proband | str | required | Proband sample name |
| mother | str | required | Mother sample name |
| father | str | required | Father sample name |
| patterns | list[str] | all five | Patterns to evaluate |
Supported patterns: de_novo, dominant, recessive, compound_het, x_linked.
from vartriage import InheritanceConfig
# All patterns (default)
config = InheritanceConfig(proband="CHILD", mother="MOM", father="DAD")
# Only de novo and recessive
config = InheritanceConfig(
proband="CHILD",
mother="MOM",
father="DAD",
patterns=["de_novo", "recessive"],
)
Raises ValueError if sample names are empty, patterns list is empty, or any pattern is not in the supported set.
ClinicalReportConfig¶
Configuration for structured clinical report generation. Required when --output-format is clinical-html, clinical-pdf, or clinical-docx.
| Field | Type | Default | Notes |
|---|---|---|---|
patient_id |
str |
required | Patient identifier (non-empty) |
panel_name |
str |
required | Gene panel name (non-empty) |
output_format |
Literal |
required | "clinical-pdf", "clinical-html", or "clinical-docx" |
report_template |
str |
"standard" |
Report template name |
from vartriage.models.config import ClinicalReportConfig
config = ClinicalReportConfig(
patient_id="PAT-2026-001",
panel_name="Cardiac Panel v3",
output_format="clinical-html",
)
# With custom template
config = ClinicalReportConfig(
patient_id="PAT-2026-002",
panel_name="Hereditary Cancer Panel",
output_format="clinical-pdf",
report_template="standard",
)
Raises ValueError at construction if patient_id or panel_name is empty or whitespace-only.
The clinical report produces:
- Self-contained HTML (no external resources, no JavaScript)
- PDF via WeasyPrint (install with
pip install weasyprint) - DOCX via python-docx (install with
pip install python-docx)
A JSON audit trail sidecar (.audit.json) is written alongside every clinical report. It contains the run manifest (config, reference checksums, timestamps) and a per-variant decision log.
GeneFilterConfig¶
Restricts analysis to variants in a user-supplied gene list.
| Field | Type | Default | Notes |
|---|---|---|---|
gene_list_path |
Path |
required | Plain text file, one gene symbol per line |
from pathlib import Path
from vartriage import GeneFilterConfig
config = GeneFilterConfig(gene_list_path=Path("cardiac_panel.txt"))
The gene list file format: one symbol per line, blank lines and lines starting with # are skipped, matching is case-insensitive.
MissingDataConfig¶
Controls the missing data warning threshold.
| Field | Type | Default | Notes |
|---|---|---|---|
warning_threshold |
int |
1000 |
Summary warning fires when exceeded |
PipelineConfig¶
Top-level configuration aggregating all sub-configs.
| Field | Type | Default | Notes |
|---|---|---|---|
vcf_path |
Path |
required | .vcf or .vcf.gz |
output_path |
Path |
required | Output report path |
quality_filter |
QualityFilterConfig |
default instance | |
annotation |
Optional[AnnotationConfig] |
None |
None skips annotation |
prioritization |
PrioritizationConfig |
default instance | |
report |
ReportConfig |
default instance | |
missing_data |
MissingDataConfig |
default instance | |
inheritance |
Optional[InheritanceConfig] |
None |
None skips trio analysis |
gene_filter |
Optional[GeneFilterConfig] |
None |
None skips gene filtering |
region_filter |
Optional[RegionFilterConfig] |
None |
None skips region filtering |
sample |
Optional[SampleConfig] |
None |
None skips sample extraction |
clinical_report |
Optional[ClinicalReportConfig] |
None |
Required for clinical formats |
use_bundles |
bool |
False |
Auto-resolve reference paths from bundles |
genome_build |
str |
"grch38" |
Build for bundle resolution |
api |
Optional[object] |
None |
APIConfig for API/hybrid mode |
knowledge |
Optional[KnowledgeBaseConfig] |
None |
Gene-disease linkage knowledge base config |
mito |
Optional[MitoConfig] |
None |
Mitochondrial analysis config (auto-enabled when chrM detected) |
remote |
Optional[RemoteTabixConfig] |
None |
Remote tabix scoring config (CADD/gnomAD via HTTP) |
sv_vcf_path |
Optional[Path] |
None |
Structural variant VCF for integrated SV triage |
qc |
Optional[QCConfig] |
None |
Pre-flight quality control config (None skips QC) |
RegionFilterConfig¶
Restricts analysis to variants within BED file intervals.
| Field | Type | Default | Notes |
|---|---|---|---|
bed_path |
Path |
required | BED file with target genomic intervals |
from pathlib import Path
from vartriage.models.config import RegionFilterConfig
config = RegionFilterConfig(bed_path=Path("target_regions.bed"))
SampleConfig¶
Extracts a single sample from multi-sample VCFs.
| Field | Type | Default | Notes |
|---|---|---|---|
sample_name |
str |
required | Sample name from VCF header |
min_gq |
int \| None |
None |
Genotype quality threshold (0-99) |
from vartriage.models.config import SampleConfig
config = SampleConfig(sample_name="PROBAND_01", min_gq=20)
Raises ValueError if min_gq is not None and outside range [0, 99].
CohortConfig¶
Multi-sample cohort analysis configuration. See Cohort Analysis Guide for full usage.
| Field | Type | Default | Notes |
|---|---|---|---|
sample_vcfs |
list[Path] |
required | At least 2 VCF file paths |
output_path |
Path |
required | Output directory for reports |
cohort_name |
str |
"cohort" |
Identifier for output filenames |
min_recurrence |
int |
2 |
Minimum samples for recurrence (>= 1) |
output_format |
Literal |
"json" |
"json" or "csv" |
max_af_threshold |
float |
0.05 |
Max population AF for inclusion (0.0-1.0) |
include_singletons |
bool |
True |
Include single-sample variants in output |
sample_labels |
dict \| None |
None |
Map file stems to display labels |
parallel |
bool |
False |
Process samples concurrently |
max_workers |
int |
4 |
Thread pool size (>= 1) |
from pathlib import Path
from vartriage import CohortConfig
config = CohortConfig(
sample_vcfs=[
Path("sample1.vcf.gz"),
Path("sample2.vcf.gz"),
Path("sample3.vcf.gz"),
],
output_path=Path("cohort_results/"),
cohort_name="cardiac_study",
min_recurrence=2,
max_af_threshold=0.01,
parallel=True,
max_workers=8,
)
Raises ValueError if fewer than 2 samples, min_recurrence < 1, max_af_threshold outside [0.0, 1.0], or max_workers < 1.
KnowledgeBaseConfig¶
Gene-disease linkage knowledge base configuration. Enables OMIM disease associations, ClinGen validity, HPO phenotype matching, gnomAD constraint, and actionability annotations.
| Field | Type | Default | Notes |
|---|---|---|---|
data_dir |
Path \| None |
None |
Directory with knowledge TSV files. None uses bundled data. |
hpo_terms |
frozenset[str] |
frozenset() |
Patient HPO terms for phenotype boosting (HP:NNNNNNN format) |
inheritance_mode |
str \| None |
None |
Filter to genes matching mode: AD, AR, XL, XLD, XLR, MT |
flag_actionable |
bool |
False |
Filter to ClinGen medically actionable genes |
from vartriage.knowledge.config import KnowledgeBaseConfig
# Phenotype-driven prioritization for epilepsy patient
config = KnowledgeBaseConfig(
hpo_terms=frozenset({"HP:0001250", "HP:0001249", "HP:0002069"}),
inheritance_mode="AD",
)
# Custom knowledge directory
config = KnowledgeBaseConfig(
data_dir=Path("/data/custom_knowledge/"),
flag_actionable=True,
)
Raises ValueError if HPO terms don't match HP:NNNNNNN format or inheritance_mode is unrecognized.
MitoConfig¶
Mitochondrial variant analysis configuration. Controls detection and classification of chrM/MT variants using mtDNA-specific criteria.
| Field | Type | Default | Notes |
|---|---|---|---|
enabled |
bool |
True |
Enable/disable mitochondrial analysis |
min_heteroplasmy |
float |
1.0 |
Minimum heteroplasmy % for reporting (0.0-100.0) |
gene_map_path |
Path \| None |
None |
Custom mt_gene_map.tsv (defaults to bundled) |
mitomap_path |
Path \| None |
None |
Custom mitomap_pathogenic.tsv (defaults to bundled) |
helixmtdb_path |
Path \| None |
None |
Custom helixmtdb_frequency.tsv (defaults to bundled) |
from vartriage.mito.config import MitoConfig
# Default: auto-detect chrM, report heteroplasmy >= 1%
config = MitoConfig()
# Custom threshold for high-confidence calls
config = MitoConfig(min_heteroplasmy=5.0)
# Disable mitochondrial analysis (targeted panels without mtDNA capture)
config = MitoConfig(enabled=False)
Mitochondrial analysis is auto-enabled when chrM/MT variants are present in the VCF. Use --skip-mito (CLI) or MitoConfig(enabled=False) to disable. Raises ValueError if min_heteroplasmy is outside [0.0, 100.0].
RemoteTabixConfig¶
Remote tabix scoring configuration. Queries CADD and gnomAD scores from public HTTP servers via byte-range requests without downloading multi-GB reference files locally.
| Field | Type | Default | Notes |
|---|---|---|---|
cadd_remote_url |
str \| None |
None |
URL or named preset (e.g., cadd-v1.7-grch38) |
gnomad_remote_url |
str \| None |
None |
URL or named preset. Supports {chrom} placeholder. |
cache_ttl_days |
int |
30 |
Days until cache expiry. -1 = pinned |
cache_path |
Path |
~/.vartriage/remote_cache.db |
SQLite score cache location |
connect_timeout |
float |
10.0 |
TCP connect timeout (seconds) |
read_timeout |
float |
30.0 |
Per-query read timeout (seconds) |
max_retries |
int |
3 |
Retry attempts for transient failures (5xx, timeout) |
batch_window_bp |
int |
10_000 |
Variants within this window (bp) are grouped into one range query |
from vartriage.remote.config import RemoteTabixConfig
# CADD scores via remote tabix (using named preset)
config = RemoteTabixConfig(cadd_remote_url="cadd-v1.7-grch38")
# Both CADD and gnomAD remote, pinned cache for reproducibility
config = RemoteTabixConfig(
cadd_remote_url="cadd-v1.7-grch38",
gnomad_remote_url="gnomad-exomes-v4-grch38",
cache_ttl_days=-1,
)
Local files always take priority over remote. The pipeline falls back gracefully on network failures (retries with backoff, then omits the score). See Remote Tabix Guide for presets, caching, and Python API details.
QCConfig¶
Pre-flight quality control configuration. QC runs a single streaming pass over the VCF before annotation, computing Ti/Tv, het/hom, variant count, and ins/del ratios, then validating them against assay-specific ranges.
| Field | Type | Default | Notes |
|---|---|---|---|
assay_type |
str |
"wes" |
wgs, wes, or panel (selects the threshold preset) |
strict |
bool |
False |
FAIL halts the pipeline with exit code 3 before annotation |
skip |
bool |
False |
Bypass QC entirely |
expected_ti_tv |
tuple \| None |
None |
Override Ti/Tv warn range (min, max) |
expected_het_hom |
tuple \| None |
None |
Override het/hom warn range (min, max) |
sample_id |
str \| None |
None |
Sample for het/hom; single-sample VCFs auto-detect |
from vartriage.qc.config import QCConfig
# Default WES thresholds, warn-only
config = QCConfig(assay_type="wes")
# WGS with a strict gate: FAIL halts the pipeline
config = QCConfig(assay_type="wgs", strict=True)
# Override the Ti/Tv warn range
config = QCConfig(assay_type="wes", expected_ti_tv=(1.9, 2.3))
Set PipelineConfig(qc=QCConfig(...)) to run QC as part of the pipeline, or use --assay-type, --strict-qc, and --skip-qc on the CLI. Warn ranges also read from a [qc] section in ~/.vartriage/config.toml, with CLI values taking precedence. Raises ValueError when assay_type is unrecognized or an override range has min >= max. See Quality Control Guide for metric definitions and thresholds.
ClinVarProteinIndex (PS1/PM5)¶
The ClinVarProteinIndex is passed directly to ACMGClassifier for PS1 and PM5 evidence evaluation. It loads a pre-processed TSV of ClinVar pathogenic missense variants keyed by gene and amino acid position.
from pathlib import Path
from vartriage.annotation.clinvar_protein_index import ClinVarProteinIndex
from vartriage.classification.acmg import ACMGClassifier
# Load the protein index
protein_index = ClinVarProteinIndex()
protein_index.load(Path("clinvar_protein_index.tsv"))
# Pass to the classifier
classifier = ACMGClassifier(protein_index=protein_index)
The TSV columns: gene, position, ref_aa, alt_aa, chrom, pos, ref, alt, significance. Generated using scripts/prepare_clinvar_protein_index.py from the ClinVar VCF.
Without a protein index, PS1 and PM5 are omitted (recorded as missing data sources). Behavior is identical to v0.13.0.
GnomADClient (gnomAD API)¶
The gnomAD GraphQL API client queries gnomAD v4 for per-population allele frequencies. Used as a fallback when local gnomAD files are unavailable or in API mode.
| Parameter | Type | Default | Notes |
|---|---|---|---|
dataset |
str |
"gnomad_r4" |
Dataset version: gnomad_r4, gnomad_r3, gnomad_r2_1 |
prefer_source |
str |
"combined" |
Data source: exome, genome, or combined |
max_retries |
int |
3 |
Retry attempts for transient failures |
timeout |
tuple[float, float] |
(10.0, 30.0) |
(connect, read) timeouts in seconds |
from vartriage.api.gnomad_client import GnomADClient
from vartriage.api._rate_limiter import RateLimiter
from vartriage.api._cache import ResponseCache
from vartriage.api._circuit_breaker import CircuitBreaker
client = GnomADClient(
rate_limiter=RateLimiter(requests_per_second=5, daily_limit=10000),
cache=ResponseCache(db_path=Path("~/.vartriage/api_cache.db")),
circuit_breaker=CircuitBreaker(),
dataset="gnomad_r4",
prefer_source="combined",
)
# Lookup a single variant
freq = client.lookup_frequency("chr17", 43091429, "T", "G")
if freq:
print(f"NFE AF: {freq.nfe}, Global AF: {freq.global_af}")
The client caches responses in the shared SQLite database. Requires pip install vartriage[api] (httpx).
Example configurations¶
Stringent clinical filtering¶
For rare Mendelian disease panels:
config = PipelineConfig(
vcf_path=Path("patient.vcf.gz"),
output_path=Path("clinical_report.json"),
quality_filter=QualityFilterConfig(min_qual=50.0),
annotation=AnnotationConfig(
gene_annotation_path=Path("gencode.v44.gtf"),
gnomad_path=Path("gnomad.v4.sites.tsv"),
clinvar_path=Path("clinvar.tsv"),
),
prioritization=PrioritizationConfig(
max_allele_frequency=0.0001,
cadd_scores_path=Path("cadd.tsv"),
revel_scores_path=Path("revel.tsv"),
),
report=ReportConfig(output_format="pdf"),
)
Relaxed research filtering¶
For exploratory variant discovery:
config = PipelineConfig(
vcf_path=Path("cohort_merged.vcf.gz"),
output_path=Path("research_candidates.csv"),
quality_filter=QualityFilterConfig(min_qual=10.0),
annotation=AnnotationConfig(
gene_annotation_path=Path("gencode.v44.gtf"),
gnomad_path=Path("gnomad.v4.sites.tsv"),
),
prioritization=PrioritizationConfig(
max_allele_frequency=0.05,
),
report=ReportConfig(output_format="csv"),
)
Annotation-only mode¶
Run annotation without full pipeline scoring. Use individual stages:
from vartriage import VCFParser, QualityFilter, AnnotationEngine
with VCFParser(Path("input.vcf.gz")) as parser:
qf = QualityFilter(QualityFilterConfig(min_qual=20.0))
engine = AnnotationEngine(
AnnotationConfig(
gene_annotation_path=Path("gencode.v44.gtf"),
gnomad_path=Path("gnomad.v4.sites.tsv"),
)
)
for annotated in engine.annotate(qf.apply(iter(parser))):
print(annotated.consequence, annotated.allele_frequency)
Clinical report output¶
For sign-off-ready clinical reporting with audit trail:
from vartriage.models.config import ClinicalReportConfig
config = PipelineConfig(
vcf_path=Path("patient.vcf.gz"),
output_path=Path("clinical_report.html"),
quality_filter=QualityFilterConfig(min_qual=50.0),
annotation=AnnotationConfig(
gene_annotation_path=Path("gencode.v44.gtf"),
gnomad_path=Path("gnomad.v4.sites.tsv"),
clinvar_path=Path("clinvar.tsv"),
),
prioritization=PrioritizationConfig(
max_allele_frequency=0.0001,
cadd_scores_path=Path("cadd.tsv"),
revel_scores_path=Path("revel.tsv"),
spliceai_scores_path=Path("spliceai.tsv"),
),
report=ReportConfig(output_format="clinical-html"),
clinical_report=ClinicalReportConfig(
patient_id="PAT-2026-001",
panel_name="Cardiac Panel v3",
output_format="clinical-html",
),
)
This writes clinical_report.html (self-contained, viewable offline) and clinical_report.html.audit.json (machine-parseable decision log).
Bundle Configuration (v0.6.0+)¶
The bundle system uses a TOML configuration file at ~/.vartriage/config.toml:
[bundle]
default_build = "grch38"
download_concurrency = 2
storage_path = "~/.vartriage/bundles"
auto_verify = true
[bundle.proxy]
http_proxy = ""
https_proxy = ""
Environment variable overrides:
VARTRIAGE_BUNDLE_STORAGE- override storage pathVARTRIAGE_DEFAULT_BUILD- override default genome build
Pipeline integration¶
Add use_bundles=True to auto-resolve missing reference paths:
config = PipelineConfig(
vcf_path=Path("patient.vcf.gz"),
output_path=Path("results.json"),
use_bundles=True, # resolve from ~/.vartriage/bundles/
genome_build="grch38", # which build to look up
)
Or via CLI:
Explicitly provided paths always take precedence over bundle resolution. See docs/bundles.md for the full bundle user guide.
API Configuration (v0.7.0+)¶
Configure the API annotation backend via ~/.vartriage/config.toml:
[api]
mode = "local" # "local" | "api" | "hybrid"
genome_build = "grch38" # "grch37" | "grch38"
ncbi_api_key = "" # or set NCBI_API_KEY env var
cache_ttl_days = 7 # -1 to pin indefinitely
cache_path = "~/.vartriage/api_cache.db"
vep_batch_size = 200 # max 200 (Ensembl limit)
max_retries = 3
preferred_frequency_source = "gnomad_exome" # or "gnomad_genome"
[api.rate_limits]
vep_requests_per_second = 15
clinvar_requests_per_second = 10 # 3 without API key
cadd_requests_per_second = 2
spliceai_requests_per_minute = 5
[api.timeouts]
connect_seconds = 10
read_seconds = 30
[api.proxy]
url = "" # e.g., "http://proxy:8080"
APIConfig Fields¶
| Field | Type | Default | Notes |
| ------- | ------ | --------- | --------- | --------- | ------- |
| mode | Literal | "local" | "local", "api", "hybrid" |
| genome_build | Literal | "grch38" | "grch37", "grch38" |
| ncbi_api_key | str \| None | None | NCBI API key for ClinVar |
| cache_path | Path | ~/.vartriage/api_cache.db | SQLite cache location |
| cache_ttl_days | int | 7 | Days until cache expiry. -1 = pinned |
| vep_batch_size | int | 200 | 1 to 200 |
| max_retries | int | 3 | 0 to 10 |
| connect_timeout | float | 10.0 | Seconds |
| read_timeout | float | 30.0 | Seconds |
| vep_rate_limit | float | 15.0 | Requests/second |
| clinvar_rate_limit | float | 10.0 | Requests/second (with key) |
| cadd_rate_limit | float | 2.0 | Requests/second |
| spliceai_rate_limit | float | 0.08 | Requests/second (5/min) |
| vep_daily_limit | int \| None | 55000 | VEP daily request cap |
| proxy_url | str \| None | None | HTTP proxy URL |
| preferred_frequency_source | Literal | "gnomad_exome" | "gnomad_exome", "gnomad_genome" |
Environment Variables¶
| Variable | Overrides |
| ------- | --------- | ----------- |
| NCBI_API_KEY | ncbi_api_key |
| HTTPS_PROXY | proxy_url |
| HTTP_PROXY | proxy_url (lower priority than HTTPS_PROXY) |
Priority Order¶
Settings merge with this precedence (highest wins):
- CLI flags (
--mode,--api-key) - Environment variables
- TOML config file (
~/.vartriage/config.toml[api]section) - Built-in defaults
See docs/api-mode.md for the full API mode user guide.