Skip to content

Configuration

All configuration classes are frozen dataclasses. They validate parameters at construction and raise ValueError for out-of-range values.

QualityFilterConfig

Controls which variants pass the quality gate.

Field Type Default Valid range
min_qual float 20.0 0 to 1,000,000
from vartriage import QualityFilterConfig

# Default: QUAL >= 20
config = QualityFilterConfig()

# Stringent clinical threshold
config = QualityFilterConfig(min_qual=50.0)

AnnotationConfig

Paths to reference data and batch size for the annotation engine.

| Field | Type | Default | Valid range / notes | | ------- | ------ | --------- | --------- | ------------------- | | gene_annotation_path | Path | required | GTF or GFF file | | gnomad_path | Path | required | gnomAD TSV or tabix VCF (.vcf.bgz/.vcf.gz) | | clinvar_path | Optional[Path] | None | ClinVar TSV, or None to skip | | reference_fasta_path | Optional[Path] | None | Indexed FASTA (.fa + .fai) for codon-level consequence calling | | batch_size | int | 10_000 | 1,000 to 100,000 |

from pathlib import Path
from vartriage import AnnotationConfig

config = AnnotationConfig(
    gene_annotation_path=Path("gencode.v44.gtf"),
    gnomad_path=Path("gnomad.v4.sites.tsv"),
    clinvar_path=Path("clinvar_20240101.tsv"),
    reference_fasta_path=Path("GRCh38.fa"),  # enables codon-level consequence calling
    batch_size=50_000,  # larger batches for faster processing
)

PrioritizationConfig

Controls frequency filtering and pathogenicity scoring.

Field Type Default Valid range
max_allele_frequency float 0.01 0.0 to 1.0
cadd_scores_path Optional[Path] None CADD Phred TSV
revel_scores_path Optional[Path] None REVEL scores TSV
spliceai_scores_path Optional[Path] None SpliceAI scores TSV
batch_size int 10_000 1,000 to 100,000
from vartriage import PrioritizationConfig

# Rare disease: very stringent frequency cutoff
config = PrioritizationConfig(
    max_allele_frequency=0.0001,
    cadd_scores_path=Path("cadd_v1.7.tsv"),
    revel_scores_path=Path("revel_v1.3.tsv"),
    spliceai_scores_path=Path("spliceai_scores.tsv"),
)

# Research: relaxed frequency, no score files
config = PrioritizationConfig(
    max_allele_frequency=0.05,
)

ReportConfig

Output format selection.

Field Type Default Options
output_format Literal[...] "json" "json", "csv", "pdf", "vcf", "clinical-html", "clinical-pdf", "clinical-docx"
from vartriage import ReportConfig

config = ReportConfig(output_format="csv")

When the format is clinical-html, clinical-pdf, or clinical-docx, the pipeline constructs a ClinicalReportConfig instead and delegates to the clinical report generator. See the ClinicalReportConfig section below.

InheritanceConfig

Trio-based inheritance pattern classification settings.

| Field | Type | Default | Notes | | ------- | ------ | --------- | ---- | ------- | | proband | str | required | Proband sample name | | mother | str | required | Mother sample name | | father | str | required | Father sample name | | patterns | list[str] | all five | Patterns to evaluate |

Supported patterns: de_novo, dominant, recessive, compound_het, x_linked.

from vartriage import InheritanceConfig

# All patterns (default)
config = InheritanceConfig(proband="CHILD", mother="MOM", father="DAD")

# Only de novo and recessive
config = InheritanceConfig(
    proband="CHILD",
    mother="MOM",
    father="DAD",
    patterns=["de_novo", "recessive"],
)

Raises ValueError if sample names are empty, patterns list is empty, or any pattern is not in the supported set.

ClinicalReportConfig

Configuration for structured clinical report generation. Required when --output-format is clinical-html, clinical-pdf, or clinical-docx.

Field Type Default Notes
patient_id str required Patient identifier (non-empty)
panel_name str required Gene panel name (non-empty)
output_format Literal required "clinical-pdf", "clinical-html", or "clinical-docx"
report_template str "standard" Report template name
from vartriage.models.config import ClinicalReportConfig

config = ClinicalReportConfig(
    patient_id="PAT-2026-001",
    panel_name="Cardiac Panel v3",
    output_format="clinical-html",
)

# With custom template
config = ClinicalReportConfig(
    patient_id="PAT-2026-002",
    panel_name="Hereditary Cancer Panel",
    output_format="clinical-pdf",
    report_template="standard",
)

Raises ValueError at construction if patient_id or panel_name is empty or whitespace-only.

The clinical report produces:

  • Self-contained HTML (no external resources, no JavaScript)
  • PDF via WeasyPrint (install with pip install weasyprint)
  • DOCX via python-docx (install with pip install python-docx)

A JSON audit trail sidecar (.audit.json) is written alongside every clinical report. It contains the run manifest (config, reference checksums, timestamps) and a per-variant decision log.

GeneFilterConfig

Restricts analysis to variants in a user-supplied gene list.

Field Type Default Notes
gene_list_path Path required Plain text file, one gene symbol per line
from pathlib import Path
from vartriage import GeneFilterConfig

config = GeneFilterConfig(gene_list_path=Path("cardiac_panel.txt"))

The gene list file format: one symbol per line, blank lines and lines starting with # are skipped, matching is case-insensitive.

MissingDataConfig

Controls the missing data warning threshold.

Field Type Default Notes
warning_threshold int 1000 Summary warning fires when exceeded
from vartriage import MissingDataConfig

config = MissingDataConfig(warning_threshold=500)

PipelineConfig

Top-level configuration aggregating all sub-configs.

Field Type Default Notes
vcf_path Path required .vcf or .vcf.gz
output_path Path required Output report path
quality_filter QualityFilterConfig default instance
annotation Optional[AnnotationConfig] None None skips annotation
prioritization PrioritizationConfig default instance
report ReportConfig default instance
missing_data MissingDataConfig default instance
inheritance Optional[InheritanceConfig] None None skips trio analysis
gene_filter Optional[GeneFilterConfig] None None skips gene filtering
region_filter Optional[RegionFilterConfig] None None skips region filtering
sample Optional[SampleConfig] None None skips sample extraction
clinical_report Optional[ClinicalReportConfig] None Required for clinical formats
use_bundles bool False Auto-resolve reference paths from bundles
genome_build str "grch38" Build for bundle resolution
api Optional[object] None APIConfig for API/hybrid mode
knowledge Optional[KnowledgeBaseConfig] None Gene-disease linkage knowledge base config
mito Optional[MitoConfig] None Mitochondrial analysis config (auto-enabled when chrM detected)
remote Optional[RemoteTabixConfig] None Remote tabix scoring config (CADD/gnomAD via HTTP)
sv_vcf_path Optional[Path] None Structural variant VCF for integrated SV triage
qc Optional[QCConfig] None Pre-flight quality control config (None skips QC)

RegionFilterConfig

Restricts analysis to variants within BED file intervals.

Field Type Default Notes
bed_path Path required BED file with target genomic intervals
from pathlib import Path
from vartriage.models.config import RegionFilterConfig

config = RegionFilterConfig(bed_path=Path("target_regions.bed"))

SampleConfig

Extracts a single sample from multi-sample VCFs.

Field Type Default Notes
sample_name str required Sample name from VCF header
min_gq int \| None None Genotype quality threshold (0-99)
from vartriage.models.config import SampleConfig

config = SampleConfig(sample_name="PROBAND_01", min_gq=20)

Raises ValueError if min_gq is not None and outside range [0, 99].

CohortConfig

Multi-sample cohort analysis configuration. See Cohort Analysis Guide for full usage.

Field Type Default Notes
sample_vcfs list[Path] required At least 2 VCF file paths
output_path Path required Output directory for reports
cohort_name str "cohort" Identifier for output filenames
min_recurrence int 2 Minimum samples for recurrence (>= 1)
output_format Literal "json" "json" or "csv"
max_af_threshold float 0.05 Max population AF for inclusion (0.0-1.0)
include_singletons bool True Include single-sample variants in output
sample_labels dict \| None None Map file stems to display labels
parallel bool False Process samples concurrently
max_workers int 4 Thread pool size (>= 1)
from pathlib import Path
from vartriage import CohortConfig

config = CohortConfig(
    sample_vcfs=[
        Path("sample1.vcf.gz"),
        Path("sample2.vcf.gz"),
        Path("sample3.vcf.gz"),
    ],
    output_path=Path("cohort_results/"),
    cohort_name="cardiac_study",
    min_recurrence=2,
    max_af_threshold=0.01,
    parallel=True,
    max_workers=8,
)

Raises ValueError if fewer than 2 samples, min_recurrence < 1, max_af_threshold outside [0.0, 1.0], or max_workers < 1.

KnowledgeBaseConfig

Gene-disease linkage knowledge base configuration. Enables OMIM disease associations, ClinGen validity, HPO phenotype matching, gnomAD constraint, and actionability annotations.

Field Type Default Notes
data_dir Path \| None None Directory with knowledge TSV files. None uses bundled data.
hpo_terms frozenset[str] frozenset() Patient HPO terms for phenotype boosting (HP:NNNNNNN format)
inheritance_mode str \| None None Filter to genes matching mode: AD, AR, XL, XLD, XLR, MT
flag_actionable bool False Filter to ClinGen medically actionable genes
from vartriage.knowledge.config import KnowledgeBaseConfig

# Phenotype-driven prioritization for epilepsy patient
config = KnowledgeBaseConfig(
    hpo_terms=frozenset({"HP:0001250", "HP:0001249", "HP:0002069"}),
    inheritance_mode="AD",
)

# Custom knowledge directory
config = KnowledgeBaseConfig(
    data_dir=Path("/data/custom_knowledge/"),
    flag_actionable=True,
)

Raises ValueError if HPO terms don't match HP:NNNNNNN format or inheritance_mode is unrecognized.

MitoConfig

Mitochondrial variant analysis configuration. Controls detection and classification of chrM/MT variants using mtDNA-specific criteria.

Field Type Default Notes
enabled bool True Enable/disable mitochondrial analysis
min_heteroplasmy float 1.0 Minimum heteroplasmy % for reporting (0.0-100.0)
gene_map_path Path \| None None Custom mt_gene_map.tsv (defaults to bundled)
mitomap_path Path \| None None Custom mitomap_pathogenic.tsv (defaults to bundled)
helixmtdb_path Path \| None None Custom helixmtdb_frequency.tsv (defaults to bundled)
from vartriage.mito.config import MitoConfig

# Default: auto-detect chrM, report heteroplasmy >= 1%
config = MitoConfig()

# Custom threshold for high-confidence calls
config = MitoConfig(min_heteroplasmy=5.0)

# Disable mitochondrial analysis (targeted panels without mtDNA capture)
config = MitoConfig(enabled=False)

Mitochondrial analysis is auto-enabled when chrM/MT variants are present in the VCF. Use --skip-mito (CLI) or MitoConfig(enabled=False) to disable. Raises ValueError if min_heteroplasmy is outside [0.0, 100.0].

RemoteTabixConfig

Remote tabix scoring configuration. Queries CADD and gnomAD scores from public HTTP servers via byte-range requests without downloading multi-GB reference files locally.

Field Type Default Notes
cadd_remote_url str \| None None URL or named preset (e.g., cadd-v1.7-grch38)
gnomad_remote_url str \| None None URL or named preset. Supports {chrom} placeholder.
cache_ttl_days int 30 Days until cache expiry. -1 = pinned
cache_path Path ~/.vartriage/remote_cache.db SQLite score cache location
connect_timeout float 10.0 TCP connect timeout (seconds)
read_timeout float 30.0 Per-query read timeout (seconds)
max_retries int 3 Retry attempts for transient failures (5xx, timeout)
batch_window_bp int 10_000 Variants within this window (bp) are grouped into one range query
from vartriage.remote.config import RemoteTabixConfig

# CADD scores via remote tabix (using named preset)
config = RemoteTabixConfig(cadd_remote_url="cadd-v1.7-grch38")

# Both CADD and gnomAD remote, pinned cache for reproducibility
config = RemoteTabixConfig(
    cadd_remote_url="cadd-v1.7-grch38",
    gnomad_remote_url="gnomad-exomes-v4-grch38",
    cache_ttl_days=-1,
)

Local files always take priority over remote. The pipeline falls back gracefully on network failures (retries with backoff, then omits the score). See Remote Tabix Guide for presets, caching, and Python API details.

QCConfig

Pre-flight quality control configuration. QC runs a single streaming pass over the VCF before annotation, computing Ti/Tv, het/hom, variant count, and ins/del ratios, then validating them against assay-specific ranges.

Field Type Default Notes
assay_type str "wes" wgs, wes, or panel (selects the threshold preset)
strict bool False FAIL halts the pipeline with exit code 3 before annotation
skip bool False Bypass QC entirely
expected_ti_tv tuple \| None None Override Ti/Tv warn range (min, max)
expected_het_hom tuple \| None None Override het/hom warn range (min, max)
sample_id str \| None None Sample for het/hom; single-sample VCFs auto-detect
from vartriage.qc.config import QCConfig

# Default WES thresholds, warn-only
config = QCConfig(assay_type="wes")

# WGS with a strict gate: FAIL halts the pipeline
config = QCConfig(assay_type="wgs", strict=True)

# Override the Ti/Tv warn range
config = QCConfig(assay_type="wes", expected_ti_tv=(1.9, 2.3))

Set PipelineConfig(qc=QCConfig(...)) to run QC as part of the pipeline, or use --assay-type, --strict-qc, and --skip-qc on the CLI. Warn ranges also read from a [qc] section in ~/.vartriage/config.toml, with CLI values taking precedence. Raises ValueError when assay_type is unrecognized or an override range has min >= max. See Quality Control Guide for metric definitions and thresholds.

ClinVarProteinIndex (PS1/PM5)

The ClinVarProteinIndex is passed directly to ACMGClassifier for PS1 and PM5 evidence evaluation. It loads a pre-processed TSV of ClinVar pathogenic missense variants keyed by gene and amino acid position.

from pathlib import Path
from vartriage.annotation.clinvar_protein_index import ClinVarProteinIndex
from vartriage.classification.acmg import ACMGClassifier

# Load the protein index
protein_index = ClinVarProteinIndex()
protein_index.load(Path("clinvar_protein_index.tsv"))

# Pass to the classifier
classifier = ACMGClassifier(protein_index=protein_index)

The TSV columns: gene, position, ref_aa, alt_aa, chrom, pos, ref, alt, significance. Generated using scripts/prepare_clinvar_protein_index.py from the ClinVar VCF.

Without a protein index, PS1 and PM5 are omitted (recorded as missing data sources). Behavior is identical to v0.13.0.

GnomADClient (gnomAD API)

The gnomAD GraphQL API client queries gnomAD v4 for per-population allele frequencies. Used as a fallback when local gnomAD files are unavailable or in API mode.

Parameter Type Default Notes
dataset str "gnomad_r4" Dataset version: gnomad_r4, gnomad_r3, gnomad_r2_1
prefer_source str "combined" Data source: exome, genome, or combined
max_retries int 3 Retry attempts for transient failures
timeout tuple[float, float] (10.0, 30.0) (connect, read) timeouts in seconds
from vartriage.api.gnomad_client import GnomADClient
from vartriage.api._rate_limiter import RateLimiter
from vartriage.api._cache import ResponseCache
from vartriage.api._circuit_breaker import CircuitBreaker

client = GnomADClient(
    rate_limiter=RateLimiter(requests_per_second=5, daily_limit=10000),
    cache=ResponseCache(db_path=Path("~/.vartriage/api_cache.db")),
    circuit_breaker=CircuitBreaker(),
    dataset="gnomad_r4",
    prefer_source="combined",
)

# Lookup a single variant
freq = client.lookup_frequency("chr17", 43091429, "T", "G")
if freq:
    print(f"NFE AF: {freq.nfe}, Global AF: {freq.global_af}")

The client caches responses in the shared SQLite database. Requires pip install vartriage[api] (httpx).

Example configurations

Stringent clinical filtering

For rare Mendelian disease panels:

config = PipelineConfig(
    vcf_path=Path("patient.vcf.gz"),
    output_path=Path("clinical_report.json"),
    quality_filter=QualityFilterConfig(min_qual=50.0),
    annotation=AnnotationConfig(
        gene_annotation_path=Path("gencode.v44.gtf"),
        gnomad_path=Path("gnomad.v4.sites.tsv"),
        clinvar_path=Path("clinvar.tsv"),
    ),
    prioritization=PrioritizationConfig(
        max_allele_frequency=0.0001,
        cadd_scores_path=Path("cadd.tsv"),
        revel_scores_path=Path("revel.tsv"),
    ),
    report=ReportConfig(output_format="pdf"),
)

Relaxed research filtering

For exploratory variant discovery:

config = PipelineConfig(
    vcf_path=Path("cohort_merged.vcf.gz"),
    output_path=Path("research_candidates.csv"),
    quality_filter=QualityFilterConfig(min_qual=10.0),
    annotation=AnnotationConfig(
        gene_annotation_path=Path("gencode.v44.gtf"),
        gnomad_path=Path("gnomad.v4.sites.tsv"),
    ),
    prioritization=PrioritizationConfig(
        max_allele_frequency=0.05,
    ),
    report=ReportConfig(output_format="csv"),
)

Annotation-only mode

Run annotation without full pipeline scoring. Use individual stages:

from vartriage import VCFParser, QualityFilter, AnnotationEngine

with VCFParser(Path("input.vcf.gz")) as parser:
    qf = QualityFilter(QualityFilterConfig(min_qual=20.0))
    engine = AnnotationEngine(
        AnnotationConfig(
            gene_annotation_path=Path("gencode.v44.gtf"),
            gnomad_path=Path("gnomad.v4.sites.tsv"),
        )
    )
    for annotated in engine.annotate(qf.apply(iter(parser))):
        print(annotated.consequence, annotated.allele_frequency)

Clinical report output

For sign-off-ready clinical reporting with audit trail:

from vartriage.models.config import ClinicalReportConfig

config = PipelineConfig(
    vcf_path=Path("patient.vcf.gz"),
    output_path=Path("clinical_report.html"),
    quality_filter=QualityFilterConfig(min_qual=50.0),
    annotation=AnnotationConfig(
        gene_annotation_path=Path("gencode.v44.gtf"),
        gnomad_path=Path("gnomad.v4.sites.tsv"),
        clinvar_path=Path("clinvar.tsv"),
    ),
    prioritization=PrioritizationConfig(
        max_allele_frequency=0.0001,
        cadd_scores_path=Path("cadd.tsv"),
        revel_scores_path=Path("revel.tsv"),
        spliceai_scores_path=Path("spliceai.tsv"),
    ),
    report=ReportConfig(output_format="clinical-html"),
    clinical_report=ClinicalReportConfig(
        patient_id="PAT-2026-001",
        panel_name="Cardiac Panel v3",
        output_format="clinical-html",
    ),
)

This writes clinical_report.html (self-contained, viewable offline) and clinical_report.html.audit.json (machine-parseable decision log).

Bundle Configuration (v0.6.0+)

The bundle system uses a TOML configuration file at ~/.vartriage/config.toml:

[bundle]
default_build = "grch38"
download_concurrency = 2
storage_path = "~/.vartriage/bundles"
auto_verify = true

[bundle.proxy]
http_proxy = ""
https_proxy = ""

Environment variable overrides:

  • VARTRIAGE_BUNDLE_STORAGE - override storage path
  • VARTRIAGE_DEFAULT_BUILD - override default genome build

Pipeline integration

Add use_bundles=True to auto-resolve missing reference paths:

config = PipelineConfig(
    vcf_path=Path("patient.vcf.gz"),
    output_path=Path("results.json"),
    use_bundles=True,  # resolve from ~/.vartriage/bundles/
    genome_build="grch38",  # which build to look up
)

Or via CLI:

vartriage --vcf patient.vcf.gz --output results.json --use-bundles --genome-build grch38

Explicitly provided paths always take precedence over bundle resolution. See docs/bundles.md for the full bundle user guide.

API Configuration (v0.7.0+)

Configure the API annotation backend via ~/.vartriage/config.toml:

[api]
mode = "local"                        # "local" | "api" | "hybrid"
genome_build = "grch38"               # "grch37" | "grch38"
ncbi_api_key = ""                     # or set NCBI_API_KEY env var
cache_ttl_days = 7                    # -1 to pin indefinitely
cache_path = "~/.vartriage/api_cache.db"
vep_batch_size = 200                  # max 200 (Ensembl limit)
max_retries = 3
preferred_frequency_source = "gnomad_exome"  # or "gnomad_genome"

[api.rate_limits]
vep_requests_per_second = 15
clinvar_requests_per_second = 10      # 3 without API key
cadd_requests_per_second = 2
spliceai_requests_per_minute = 5

[api.timeouts]
connect_seconds = 10
read_seconds = 30

[api.proxy]
url = ""                              # e.g., "http://proxy:8080"

APIConfig Fields

| Field | Type | Default | Notes | | ------- | ------ | --------- | --------- | --------- | ------- | | mode | Literal | "local" | "local", "api", "hybrid" | | genome_build | Literal | "grch38" | "grch37", "grch38" | | ncbi_api_key | str \| None | None | NCBI API key for ClinVar | | cache_path | Path | ~/.vartriage/api_cache.db | SQLite cache location | | cache_ttl_days | int | 7 | Days until cache expiry. -1 = pinned | | vep_batch_size | int | 200 | 1 to 200 | | max_retries | int | 3 | 0 to 10 | | connect_timeout | float | 10.0 | Seconds | | read_timeout | float | 30.0 | Seconds | | vep_rate_limit | float | 15.0 | Requests/second | | clinvar_rate_limit | float | 10.0 | Requests/second (with key) | | cadd_rate_limit | float | 2.0 | Requests/second | | spliceai_rate_limit | float | 0.08 | Requests/second (5/min) | | vep_daily_limit | int \| None | 55000 | VEP daily request cap | | proxy_url | str \| None | None | HTTP proxy URL | | preferred_frequency_source | Literal | "gnomad_exome" | "gnomad_exome", "gnomad_genome" |

Environment Variables

| Variable | Overrides | | ------- | --------- | ----------- | | NCBI_API_KEY | ncbi_api_key | | HTTPS_PROXY | proxy_url | | HTTP_PROXY | proxy_url (lower priority than HTTPS_PROXY) |

Priority Order

Settings merge with this precedence (highest wins):

  1. CLI flags (--mode, --api-key)
  2. Environment variables
  3. TOML config file (~/.vartriage/config.toml [api] section)
  4. Built-in defaults

See docs/api-mode.md for the full API mode user guide.