Перейти к содержимому

WGGC Webinars | Dr. Julian Uszkoreit | Harmonizing peptide identification across search engines

West German Genome Center

0:00 / 0:00

WGGC Webinars | Dr. Julian Uszkoreit | Harmonizing peptide identification across search engines

75 просмотров · 1 месяц назад
West German Genome Center
13 подписчиков
75 просмотров · 1 месяц назад
Title: How post-processing harmonizes peptide identification across proteomics search engines Peptide search engines match fragmentation mass spectra against a protein sequence database to identify peptides from complex biological samples. Different tools produce different results from the same data, with limited consensus on which to trust - a fundamental challenge for reproducibility and comparability in proteomics workflows. In this talk, I present a systematic benchmarking study comparing seven widely used data-dependent peptide search engines across four datasets, multiple mass spectrometry platforms, and protein databases of varying size and composition. We evaluated the impact of three rescoring strategies - from a semi-supervised machine learning approach (Percolator) to prediction-based methods leveraging deep learning fragment ion intensity models (MS2Rescore and Oktoberfest) - alongside standard target-decoy FDR estimation. Our findings show that prediction-based rescoring dramatically reduces inter-tool variability, effectively harmonizing identification results that would otherwise differ substantially. We further examine how database composition shapes outcomes - particularly in metaproteomic settings, where the challenge of an incomplete reference mirrors the difficulties of working with poorly assembled or highly diverse genomes. FDR reliability is assessed through entrapment-based analyses, and practical considerations around computational runtime and resource consumption are discussed. Modern rescoring has fundamentally changed how we should think about search engine selection in proteomics: once prediction-based rescoring is applied, tool choice becomes far less critical than commonly assumed. What remains indispensable, regardless of the tool, is rigorous FDR validation at every step of the analysis.