Open AccessProceedings Article
Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra
John T. Halloran,David M. Rocke +1 more
- 01 Jan 2018
- Vol. 31, pp 5420-5430
TL;DR: This work greatly expands the parameter learning capabilities of a dynamic Bayesian network (DBN) peptide-scoring algorithm, Didea, by deriving emission distributions for which its conditional log-likelihood scoring function remains concave, and significantly improves DideA's runtime performance through successive optimizations to its message passing schedule.
read more
Abstract: The most widely used technology to identify the proteins present in a complex biological sample is tandem mass spectrometry, which quickly produces a large collection of spectra representative of the peptides (i.e., protein subsequences) present in the original sample. In this work, we greatly expand the parameter learning capabilities of a dynamic Bayesian network (DBN) peptide-scoring algorithm, Didea [25], by deriving emission distributions for which its conditional log-likelihood scoring function remains concave. We show that this class of emission distributions, called Convex Virtual Emissions (CVEs), naturally generalizes the log-sum-exp function while rendering both maximum likelihood estimation and conditional maximum likelihood estimation concave for a wide range of Bayesian networks. Utilizing CVEs in Didea allows efficient learning of a large number of parameters while ensuring global convergence, in stark contrast to Didea's previous parameter learning framework (which could only learn a single parameter using a costly grid search) and other trainable models [12, 13, 14] (which only ensure convergence to local optima). The newly trained scoring function substantially outperforms the state-of-the-art in both scoring function accuracy and downstream Fisher kernel analysis. Furthermore, we significantly improve Didea's runtime performance through successive optimizations to its message passing schedule and derive explicit connections between Didea's new concave score and related MS/MS scoring functions.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Proceedings Article
Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra
John T. Halloran,David M. Rocke +1 more
- 01 Jan 2018
TL;DR: This work greatly expands the parameter learning capabilities of a dynamic Bayesian network (DBN) peptide-scoring algorithm, Didea, by deriving emission distributions for which its conditional log-likelihood scoring function remains concave, and significantly improves DideA's runtime performance through successive optimizations to its message passing schedule.
Deep Semi-Supervised Learning Improves Universal Peptide Identification of Shotgun Proteomics Data
TL;DR: It is shown that deep models significantly improve the recalibration of PSMs compared to the most accurate and widely-used post-processors, such as Percolator and PeptideProphet, and that deep learning is able to adaptively analyze complex datasets and features for more accurate universal post-processing.
4
•Proceedings Article
GPU-accelerated primal learning for extremely fast large-scale classification
John T. Halloran,David M. Rocke +1 more
- 01 Jan 2020
TL;DR: In this paper, the authors show that using GPUs to train logistic regression classifiers in LIBLINEAR is up to an order-of-magnitude faster than solely using multithreading.
•Posted Content
GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification
John T. Halloran,David M. Rocke +1 more
TL;DR: This work shows how GPU speedups may be mixed with multithreading to enable such speedups when the dataset is too large for GPU memory requirements; on a massive dense proteomics dataset, these mixed-architecture speedups reduce SVM analysis time from over half a week to less than a single day while using limited GPU memory.
1
References
Semi-supervised learning for peptide identification from shotgun proteomics datasets
TL;DR: An algorithm, called Percolator, is described, for improving the rate of confident peptide identifications from a collection of tandem mass spectra, using semi-supervised machine learning to discriminate between correct and decoy spectrum identifications.
2.3K
Comet: An open-source MS/MS sequence database search tool
TL;DR: The Comet search engine is introduced, open source, freely available, and based on one of the original sequence database search tools that has been widely used for many years.
1.4K
•Posted Content
Open Mass Spectrometry Search Algorithm
Lewis Y. Geer,Sanford P. Markey,Jeffrey A. Kowalak,Lukas Wagner,Ming Xu,Dawn M. Maynard,Xiaoyu Yang,Wenyao Shi,Stephen H. Bryant +8 more
TL;DR: The Open Mass Spectrometry Search Algorithm (OMSSA) as mentioned in this paper was designed to be faster than published algorithms in searching large MS/MS datasets for peptide identification.
1.3K
MS-GF+ makes progress towards a universal database search tool for proteomics.
Sangtae Kim,Pavel A. Pevzner +1 more
TL;DR: A database search tool MS-GF+ is presented that is sensitive (it identifies more peptides than most other database search tools) and universal (it works well for diverse types of spectra, different configurations of MS instruments and different experimental protocols), and improves upon the performance of tools specifically designed for these applications.
Improvements to the Percolator Algorithm for Peptide Identification from Shotgun Proteomics Data Sets
TL;DR: Improvements to Percolator are described, including a method, Q-ranker, for directly optimizing the number of identified spectra at a specified q value, which achieves further gains.