Cohort Analysis¶
Multi-sample cohort analysis classes for cross-sample variant aggregation and reporting.
CohortPipeline¶
vartriage.cohort.pipeline.CohortPipeline
¶
Orchestrate multi-sample cohort analysis.
Processes each sample VCF through the standard vartriage pipeline, collects classified variants, then merges them via CohortAggregator for cross-sample analysis.
Parameters¶
cohort_config : CohortConfig Cohort-level settings (sample list, thresholds, output). pipeline_config : PipelineConfig | None Base pipeline configuration applied to each sample. When None, a minimal config is constructed per-sample using only the cohort_config's sample VCF paths with default quality/prioritization settings. annotation_config : AnnotationConfig | None Shared annotation config for all samples. Overrides the pipeline_config's annotation setting when provided. prioritization_config : PrioritizationConfig | None Shared prioritization config. Overrides pipeline_config when provided.
Source code in vartriage/cohort/pipeline.py
36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 | |
gene_burdens
property
¶
Per-gene burden records (populated after run()).
summary
property
¶
Cohort summary statistics (populated after run()).
variants
property
¶
Aggregated cohort variants (populated after run()).
run()
¶
Execute the full cohort analysis pipeline.
Sequence: 1. Process each sample VCF through the standard pipeline 2. Aggregate variants across samples 3. Compute cohort statistics 4. Generate reports
Returns¶
list[Path] Paths to generated report files.
Raises¶
FileNotFoundError If any sample VCF file does not exist.
Source code in vartriage/cohort/pipeline.py
Orchestrates per-sample Pipeline execution, aggregation, statistics computation, and report generation.
from vartriage import CohortPipeline, CohortConfig
config = CohortConfig(
sample_vcfs=[Path("a.vcf.gz"), Path("b.vcf.gz")],
output_path=Path("output/"),
)
pipeline = CohortPipeline(cohort_config=config)
report_paths = pipeline.run()
# Access results after run()
pipeline.variants # list[CohortVariant]
pipeline.gene_burdens # list[GeneBurden]
pipeline.summary # CohortSummary
Parameters:
| Name | Type | Description |
|---|---|---|
| cohort_config | CohortConfig | Cohort-level settings |
| pipeline_config | PipelineConfig, optional | Base config applied to each sample |
| annotation_config | AnnotationConfig, optional | Shared annotation config (overrides pipeline_config) |
| prioritization_config | PrioritizationConfig, optional | Shared scoring config (overrides pipeline_config) |
Methods:
run() -> list[Path]- Execute the full cohort analysis, returns paths to generated reports.
Properties (populated after run):
variants: list[CohortVariant]gene_burdens: list[GeneBurden]summary: CohortSummary | None
CohortAggregator¶
vartriage.cohort.aggregator.CohortAggregator
¶
Merges per-sample classified variants into cohort-level records.
Groups variants by genomic coordinate across all samples, then produces CohortVariant records with cross-sample frequency and merged evidence. Respects the config's min_recurrence threshold and max_af_threshold for filtering.
Parameters¶
config : CohortConfig Cohort analysis configuration.
Source code in vartriage/cohort/aggregator.py
89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 | |
samples_added
property
¶
Number of samples that have been ingested so far.
total_distinct_variants
property
¶
Number of distinct variant coordinates seen across all samples.
add_sample(sample_id, vcf_path, variants)
¶
Ingest classified variants from a single sample.
Parameters¶
sample_id : str Human-readable sample identifier. vcf_path : Path Source VCF path for traceability. variants : list[ClassifiedVariant] All classified variants from this sample's pipeline run.
Returns¶
int Number of variants added from this sample (after AF filtering).
Source code in vartriage/cohort/aggregator.py
aggregate()
¶
Produce merged cohort variants from all ingested samples.
Filtering logic: - Variants with sample_count >= min_recurrence are included. - Singletons (sample_count == 1) are included only when include_singletons is True. - Variants with 1 < sample_count < min_recurrence are excluded.
This means min_recurrence acts as a hard inclusion threshold, not just a highlight marker. Set min_recurrence=1 with include_singletons=True to get all variants regardless of recurrence.
Results are sorted by sample_count descending, then by genomic coordinate.
Returns¶
list[CohortVariant] Cohort-level variant records meeting inclusion criteria.
Source code in vartriage/cohort/aggregator.py
get_recurrent_variants(min_count=2)
¶
Return only variants appearing in >= min_count samples.
Convenience method for extracting shared variants without changing the config's min_recurrence permanently.
Parameters¶
min_count : int Minimum sample count threshold. Default is 2.
Returns¶
list[CohortVariant] Filtered and sorted cohort variants.
Source code in vartriage/cohort/aggregator.py
Merges classified variants from multiple samples by genomic coordinate.
Methods:
add_sample(sample_id, vcf_path, variants) -> int- Ingest one sample's results. Returns count after AF filtering.aggregate() -> list[CohortVariant]- Produce merged cohort variants respecting config thresholds.get_recurrent_variants(min_count=2) -> list[CohortVariant]- Convenience filter for shared variants.reset()- Clear all data for reuse.
Properties:
samples_added: inttotal_distinct_variants: int
CohortStatistics¶
vartriage.cohort.statistics.CohortStatistics
¶
Compute summary statistics from aggregated cohort variants.
Takes a list of CohortVariant records (output of CohortAggregator) and produces per-gene burden tables, recurrence distributions, and the overall CohortSummary dataclass.
Parameters¶
config : CohortConfig Cohort configuration (used for cohort_name and sample metadata). variants : list[CohortVariant] Aggregated cohort variants to analyze.
Source code in vartriage/cohort/statistics.py
33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 | |
variant_count
property
¶
Total distinct variants in the cohort.
classification_distribution()
¶
Count variants by their most severe ACMG classification.
Returns¶
dict[str, int] Mapping of classification name -> count.
Source code in vartriage/cohort/statistics.py
compute_gene_burden()
¶
Compute per-gene variant burden across the cohort.
Groups variants by gene, counts pathogenic/likely_pathogenic hits, and tracks how many samples are affected per gene. Results are sorted by pathogenic_count descending, then by samples_affected descending.
Returns¶
list[GeneBurden] Per-gene burden records, sorted by severity.
Source code in vartriage/cohort/statistics.py
compute_summary(samples_processed)
¶
Produce the top-level cohort summary.
Parameters¶
samples_processed : list[str] Ordered list of sample identifiers that were analyzed.
Returns¶
CohortSummary Aggregate statistics for the cohort run.
Source code in vartriage/cohort/statistics.py
consequence_distribution()
¶
Count variants by functional consequence type.
Returns¶
dict[str, int] Mapping of consequence name -> count.
Source code in vartriage/cohort/statistics.py
per_sample_counts()
¶
Count total variants per sample across the cohort.
Returns¶
dict[str, int] Mapping of sample_id -> number of cohort variants that include that sample. Sorted by count descending.
Source code in vartriage/cohort/statistics.py
recurrence_distribution()
¶
Count how many variants appear in exactly N samples.
Returns¶
dict[int, int] Mapping of sample_count -> number of variants with that count. Sorted by key ascending.
Source code in vartriage/cohort/statistics.py
Computes summary metrics from aggregated cohort variants.
Methods:
compute_summary(samples_processed) -> CohortSummary- Top-level metrics.compute_gene_burden() -> list[GeneBurden]- Per-gene mutation burden sorted by severity.recurrence_distribution() -> dict[int, int]- Sample count histogram.per_sample_counts() -> dict[str, int]- Variants per sample.classification_distribution() -> dict[str, int]- Count by ACMG class.consequence_distribution() -> dict[str, int]- Count by consequence type.
CohortReportGenerator¶
vartriage.cohort.report.CohortReportGenerator
¶
Generate cohort analysis reports in JSON or CSV format.
Writes three output files per cohort run: - variants report: all cohort variants with recurrence data - gene burden report: per-gene statistics - summary report: top-level cohort metrics
Parameters¶
config : CohortConfig Cohort configuration with output_path and output_format.
Source code in vartriage/cohort/report.py
27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 | |
generate(variants, gene_burdens, summary)
¶
Write all cohort report files.
Parameters¶
variants : list[CohortVariant] Aggregated cohort variants. gene_burdens : list[GeneBurden] Per-gene burden statistics. summary : CohortSummary Top-level cohort metrics.
Returns¶
list[Path] Paths to all generated report files.
Source code in vartriage/cohort/report.py
Writes cohort analysis results to disk in JSON or CSV format.
Methods:
generate(variants, gene_burdens, summary) -> list[Path]- Write all report files, returns paths.
Data Models¶
CohortConfig¶
Frozen dataclass. Configuration for multi-sample cohort analysis. See Cohort Analysis Guide for the full field reference.
CohortVariant¶
Frozen dataclass. A variant aggregated across multiple samples.
Key fields: chrom, pos, ref, alt, gene_name, consequence, sample_count, total_samples, occurrences, max_classification, all_evidence_tags, allele_frequency.
Properties: key, cohort_frequency, is_singleton, is_universal, sample_ids.
SampleOccurrence¶
Frozen dataclass. Record of a variant's appearance in one sample.
Fields: sample_id, vcf_path, classified (ClassifiedVariant).
GeneBurden¶
Frozen dataclass. Per-gene variant burden across the cohort.
Fields: gene_name, total_variants, pathogenic_count, samples_affected, total_samples, most_severe.
Properties: penetrance (float, fraction of cohort affected).
CohortSummary¶
Frozen dataclass. Aggregate statistics for a completed cohort run.
Fields: cohort_name, total_samples, total_variants, shared_variants, singleton_variants, universal_variants, pathogenic_variants, likely_pathogenic_variants, genes_affected, top_recurrent_genes, samples_processed.