Benchmarking somatic wes pipelines: A systematic evaluation of computational parameters and their clinical implications

Yükleniyor...
Küçük Resim

Akademik Birim

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Graduate School

Özet

Next-generation sequencing (NGS) technologies have transformed biomedical research and clinical diagnostics by enabling rapid, high-throughput profiling of genomic material at a fraction of the cost of earlier sequencing methods. Among the various applications of NGS, whole-exome sequencing (WES) has become particularly important for cancer genomics. WES captures the protein-coding regions of the genome---where the majority of disease-causing mutations reside---and is substantially more cost-effective than whole-genome sequencing (WGS). In the context of oncology, somatic variant calling---the computational process of identifying mutations that arise specifically in tumor tissue and are absent from the matched normal tissue of the same patient---serves as the foundation for precision medicine workflows that guide treatment selection, prognostication, and biomarker discovery. The bioinformatics pipelines used to perform somatic variant calling are composed of multiple sequential steps, and at each step researchers must select from a range of competing algorithms and configure their associated parameters. These steps typically include quality-control trimming of raw sequencing reads, alignment (mapping) of reads to a reference genome, handling of PCR duplicate reads, base quality score recalibration (BQSR), and the variant calling step itself. Although a substantial body of literature has examined how the choice of variant calling algorithm affects detection accuracy, far less attention has been paid to the upstream pre-processing steps or to the interactions between different stages of the pipeline. This thesis presents a comprehensive benchmarking study designed to address these gaps. A total of 480 pipeline configurations were constructed by combining six key factors: two mapping algorithms (BWA-MEM and Bowtie2), three variant callers (Mutect2, Strelka2, and SomaticSniper), two trimming options, two base recalibration options, two duplicate handling strategies, and two high-performance computing environments. These configurations were applied to paired tumor--normal WES datasets from five independent sequencing centers provided by the Sequencing Quality Control Phase~2 (SEQC2) consortium. Pipeline outputs were evaluated against the SEQC2 high-confidence variant set containing 1,161 validated single-nucleotide polymorphisms (SNPs). The results reveal substantial heterogeneity in variant calling outcomes driven primarily by the choice of variant caller, which accounted for 57.27\% of the total variance in F1-scores according to a Type~III ANOVA analysis. The sequencing center---which captures inter-center differences primarily driven by variations in read depth and library preparation---was the second largest contributor at 20.59\%. Adapter trimming was found to significantly improve recall, raising it from 0.717 to 0.800, without a statistically significant reduction in precision (FDR-corrected $p = 0.3318$). Base recalibration improved precision from 0.636 to 0.699 while reducing recall by less than 1\%. When both pre-processing steps were applied together, the overall F1-score reached 0.733, compared to 0.651 for the baseline. Importantly, trimming did not introduce meaningful computational overhead, whereas base recalibration increased execution times substantially. Among individual pipeline configurations, Bowtie combined with Mutect achieved the highest F1-score (0.922). A consensus approach combining 34 pipelines achieved the highest F1-score of 0.94. An ensemble voting strategy using as few as two optimally selected pipelines attained an F1-score of 0.926, falling within 1.48\% of the consensus peak while dramatically reducing computational requirements. The clinical implications of pipeline choice were examined through tumor mutational burden (TMB) analysis and the detection of cancer- and drug-associated gene variants, demonstrating that computational decisions directly influence therapeutic recommendations. Specifically, different mapper--caller combinations produced markedly different TMB estimates, and the application of trimming and base recalibration reduced TMB variability. In conclusion, the principal outcome of this thesis is that different somatic WES pipeline configurations yield heterogeneous variant lists from identical input data, raising serious consistency concerns for NGS-based clinical genomics. The choice of variant calling algorithm is the dominant driver of this heterogeneity, while pre-processing steps---particularly adapter trimming and base recalibration---introduce measurable trade-offs that propagate to clinically relevant endpoints such as TMB and actionable mutation detection. Based on these findings, when computational resources are limited, BWA-MEM combined with Mutect2 with both trimming and base recalibration enabled provides the best balance of accuracy and speed; when maximum accuracy is targeted, an ensemble approach consisting of at least two complementary pipelines should be used. Continuous benchmarking and transparent reporting of pipeline configurations are essential to mitigate this heterogeneity and ensure the reliability of NGS-based clinical decision-making.

Tanım

Thesis (M.Sc.) -- Istanbul Technical University, Graduate School, 2026

Dergi veya Seri

ISSN

ISBN

Haklar

Anahtar Kelimeler

Bioinformatics, Biyoinformatik, Next-Generation Sequencing, Yeni Nesil Dizileme, Cancer Genomics, Kanser Genomiği, Bioinformatics Pipelines, Biyoenformatik Boru Hatları

Alıntı

Onay

Gözden geçir

Tamamlayıcı Bilgiler

Referans Gösteren

0

Views

0

Downloads