Skip to content

Configuration Classes

vartriage.PipelineConfig dataclass

Top-level pipeline configuration aggregating all sub-configs.

Parameters

vcf_path : Path Path to the input VCF file (.vcf or .vcf.gz). output_path : Path Path where the output report file will be written. quality_filter : QualityFilterConfig Quality filtering settings. Defaults to standard thresholds. annotation : AnnotationConfig Annotation engine settings including reference file paths. prioritization : PrioritizationConfig Prioritization engine settings for frequency filtering and scoring. report : ReportConfig Report generation format settings. missing_data : MissingDataConfig Missing data handling and warning threshold settings. gene_filter : GeneFilterConfig | None Gene list filtering settings. When None, gene filtering is disabled and the annotated stream passes directly to prioritization.

Source code in vartriage/models/config.py
@dataclass(frozen=True)
class PipelineConfig:
    """Top-level pipeline configuration aggregating all sub-configs.

    Parameters
    ----------
    vcf_path : Path
        Path to the input VCF file (``.vcf`` or ``.vcf.gz``).
    output_path : Path
        Path where the output report file will be written.
    quality_filter : QualityFilterConfig
        Quality filtering settings. Defaults to standard thresholds.
    annotation : AnnotationConfig
        Annotation engine settings including reference file paths.
    prioritization : PrioritizationConfig
        Prioritization engine settings for frequency filtering and scoring.
    report : ReportConfig
        Report generation format settings.
    missing_data : MissingDataConfig
        Missing data handling and warning threshold settings.
    gene_filter : GeneFilterConfig | None
        Gene list filtering settings. When None, gene filtering is disabled
        and the annotated stream passes directly to prioritization.
    """

    vcf_path: Path
    output_path: Path
    quality_filter: QualityFilterConfig = field(default_factory=QualityFilterConfig)
    annotation: AnnotationConfig | None = None
    prioritization: PrioritizationConfig = field(default_factory=PrioritizationConfig)
    report: ReportConfig = field(default_factory=ReportConfig)
    missing_data: MissingDataConfig = field(default_factory=MissingDataConfig)
    gene_filter: GeneFilterConfig | None = field(default=None)
    region_filter: RegionFilterConfig | None = field(default=None)
    sample: SampleConfig | None = field(default=None)
    inheritance: InheritanceConfig | None = field(default=None)
    clinical_report: ClinicalReportConfig | None = field(default=None)
    use_bundles: bool = False
    use_disease_thresholds: bool = False
    genome_build: str = "grch38"
    api: object | None = field(default=None)
    knowledge: KnowledgeBaseConfig | None = field(default=None)
    sv_vcf_path: Path | None = None
    mito: MitoConfig | None = field(default=None)
    remote: RemoteTabixConfig | None = field(default=None)
    qc: QCConfig | None = field(default=None)

    def __post_init__(self) -> None:
        fmt = self.report.output_format
        if fmt.startswith("clinical-") and self.clinical_report is None:
            raise ValueError(
                f"clinical_report config is required when "
                f"report.output_format is '{fmt}'"
            )

vartriage.QualityFilterConfig dataclass

Configuration for quality-based variant filtering.

Parameters

min_qual : float Minimum QUAL score threshold. Variants with a QUAL score below this value are excluded from downstream analysis. Must be in the range [0, 1_000_000]. Default is 20.0.

Raises

ValueError If min_qual is outside the range [0, 1_000_000].

Source code in vartriage/models/config.py
@dataclass(frozen=True)
class QualityFilterConfig:
    """Configuration for quality-based variant filtering.

    Parameters
    ----------
    min_qual : float
        Minimum QUAL score threshold. Variants with a QUAL score below this
        value are excluded from downstream analysis. Must be in the range
        [0, 1_000_000]. Default is 20.0.

    Raises
    ------
    ValueError
        If ``min_qual`` is outside the range [0, 1_000_000].
    """

    min_qual: float = 20.0

    def __post_init__(self) -> None:
        if not (0 <= self.min_qual <= 1_000_000):
            raise ValueError(
                f"min_qual must be between 0 and 1000000, got {self.min_qual}"
            )

vartriage.AnnotationConfig dataclass

Configuration for the annotation engine.

Parameters

gene_annotation_path : Path Path to a GTF/GFF gene annotation reference file used for functional consequence assignment via coordinate overlap. gnomad_path : Path Path to a local gnomAD reference file for population allele frequency lookups. clinvar_path : Optional[Path] Path to a ClinVar reference file for clinical significance lookups. When None, ClinVar annotation is skipped and variants receive a null clinical significance value. batch_size : int Number of variants processed per batch during vectorized annotation operations. Must be in the range [1_000, 100_000]. Default is 10_000.

Raises

ValueError If batch_size is outside the range [1_000, 100_000].

Source code in vartriage/models/config.py
@dataclass(frozen=True)
class AnnotationConfig:
    """Configuration for the annotation engine.

    Parameters
    ----------
    gene_annotation_path : Path
        Path to a GTF/GFF gene annotation reference file used for functional
        consequence assignment via coordinate overlap.
    gnomad_path : Path
        Path to a local gnomAD reference file for population allele frequency
        lookups.
    clinvar_path : Optional[Path]
        Path to a ClinVar reference file for clinical significance lookups.
        When None, ClinVar annotation is skipped and variants receive a null
        clinical significance value.
    batch_size : int
        Number of variants processed per batch during vectorized annotation
        operations. Must be in the range [1_000, 100_000]. Default is 10_000.

    Raises
    ------
    ValueError
        If ``batch_size`` is outside the range [1_000, 100_000].
    """

    gene_annotation_path: Path
    gnomad_path: Path | None = None
    clinvar_path: Path | None = None
    reference_fasta_path: Path | None = None
    batch_size: int = 10_000

    def __post_init__(self) -> None:
        if not (1_000 <= self.batch_size <= 100_000):
            raise ValueError(
                f"batch_size must be between 1000 and 100000, got {self.batch_size}"
            )

vartriage.PrioritizationConfig dataclass

Configuration for the prioritization engine.

Parameters

max_allele_frequency : float .. deprecated:: 0.14.0 The prioritization engine no longer applies a frequency gate. All variants now pass through to ACMG classification where BA1/BS1 benign evidence tags handle frequency-based filtering. This field is retained for backward compatibility and will be removed in v1.0.0. Maximum allele frequency threshold. Must be in the range [0.0, 1.0]. Default is 0.01. cadd_scores_path : Optional[Path] Path to a CADD Phred score reference file. When None, CADD scores are not incorporated into composite ranking. revel_scores_path : Optional[Path] Path to a REVEL score reference file. When None, REVEL scores are not incorporated into composite ranking. spliceai_scores_path : Optional[Path] Path to a SpliceAI score TSV reference file. When None, SpliceAI scores are not incorporated into composite ranking. Mutually exclusive with spliceai_db_path. spliceai_db_path : Optional[Path] Path to a SpliceAI SQLite database (OpenCRAVAT format). When set, the pipeline queries precomputed delta scores directly from the database. Mutually exclusive with spliceai_scores_path. batch_size : int Number of variants processed per batch during vectorized score normalization. Must be in the range [1_000, 100_000]. Default is 10_000.

Raises

ValueError If max_allele_frequency is outside the range [0.0, 1.0]. ValueError If batch_size is outside the range [1_000, 100_000].

Source code in vartriage/models/config.py
@dataclass(frozen=True)
class PrioritizationConfig:
    """Configuration for the prioritization engine.

    Parameters
    ----------
    max_allele_frequency : float
        .. deprecated:: 0.14.0
            The prioritization engine no longer applies a frequency gate.
            All variants now pass through to ACMG classification where BA1/BS1
            benign evidence tags handle frequency-based filtering. This field
            is retained for backward compatibility and will be removed in v1.0.0.
        Maximum allele frequency threshold. Must be in the range [0.0, 1.0].
        Default is 0.01.
    cadd_scores_path : Optional[Path]
        Path to a CADD Phred score reference file. When None, CADD scores are
        not incorporated into composite ranking.
    revel_scores_path : Optional[Path]
        Path to a REVEL score reference file. When None, REVEL scores are not
        incorporated into composite ranking.
    spliceai_scores_path : Optional[Path]
        Path to a SpliceAI score TSV reference file. When None, SpliceAI
        scores are not incorporated into composite ranking. Mutually
        exclusive with ``spliceai_db_path``.
    spliceai_db_path : Optional[Path]
        Path to a SpliceAI SQLite database (OpenCRAVAT format). When set,
        the pipeline queries precomputed delta scores directly from the
        database. Mutually exclusive with ``spliceai_scores_path``.
    batch_size : int
        Number of variants processed per batch during vectorized score
        normalization. Must be in the range [1_000, 100_000]. Default is
        10_000.

    Raises
    ------
    ValueError
        If ``max_allele_frequency`` is outside the range [0.0, 1.0].
    ValueError
        If ``batch_size`` is outside the range [1_000, 100_000].
    """

    max_allele_frequency: float = 0.01
    cadd_scores_path: Path | None = None
    revel_scores_path: Path | None = None
    spliceai_scores_path: Path | None = None
    spliceai_db_path: Path | None = None
    batch_size: int = 10_000

    def __post_init__(self) -> None:
        if not (0.0 <= self.max_allele_frequency <= 1.0):
            raise ValueError(
                f"max_allele_frequency must be between 0.0 and 1.0, "
                f"got {self.max_allele_frequency}"
            )
        if not (1_000 <= self.batch_size <= 100_000):
            raise ValueError(
                f"batch_size must be between 1000 and 100000, got {self.batch_size}"
            )
        if self.spliceai_scores_path and self.spliceai_db_path:
            raise ValueError(
                "Cannot configure both spliceai_scores_path (TSV) and "
                "spliceai_db_path (SQLite). Choose one SpliceAI backend."
            )

vartriage.ReportConfig dataclass

Configuration for report generation.

Parameters

output_format : str Desired output format for the final report. Accepts "json", "csv", "pdf", "vcf", "clinical-pdf", "clinical-html", or "clinical-docx". Default is "json".

Source code in vartriage/models/config.py
@dataclass(frozen=True)
class ReportConfig:
    """Configuration for report generation.

    Parameters
    ----------
    output_format : str
        Desired output format for the final report. Accepts "json",
        "csv", "pdf", "vcf", "clinical-pdf", "clinical-html", or
        "clinical-docx". Default is ``"json"``.
    """

    output_format: Literal[
        "json",
        "csv",
        "pdf",
        "vcf",
        "clinical-pdf",
        "clinical-html",
        "clinical-docx",
    ] = "json"

vartriage.MissingDataConfig dataclass

Configuration for missing data handling behavior.

Parameters

warning_threshold : int Maximum number of MissingDataWarning events allowed before the pipeline emits a summary warning. The summary includes the total count of missing-data events and the reference sources that contributed. Default is 1000.

Source code in vartriage/models/config.py
@dataclass(frozen=True)
class MissingDataConfig:
    """Configuration for missing data handling behavior.

    Parameters
    ----------
    warning_threshold : int
        Maximum number of ``MissingDataWarning`` events allowed before the
        pipeline emits a summary warning. The summary includes the total count
        of missing-data events and the reference sources that contributed.
        Default is 1000.
    """

    warning_threshold: int = 1000