Optimal simultaneous superpositioning of multiple structures with missing data
TL;DR: This work uses the expectation–maximization algorithm, a classic statistical technique for dealing with incomplete data, to find both maximum-likelihood solutions and the optimal least-squares solution as a special case for superposition when some of the data are missing.
read more
Abstract: Motivation: Superpositioning is an essential technique in structural biology that facilitates the comparison and analysis of conformational differences among topologically similar structures Performing a superposition requires a one-to-one correspondence, or alignment, of the point sets in the different structures However, in practice, some points are usually ‘missing’ from several structures, for example, when the alignment contains gaps Current superposition methods deal with missing data simply by superpositioning a subset of points that are shared among all the structures This practice is inefficient, as it ignores important data, and it fails to satisfy the common least-squares criterion In the extreme, disregarding missing positions prohibits the calculation of a superposition altogether
Results: Here, we present a general solution for determining an optimal superposition when some of the data are missing We use the expectation–maximization algorithm, a classic statistical technique for dealing with incomplete data, to find both maximum-likelihood solutions and the optimal least-squares solution as a special case
Availability and implementation: The methods presented here are implemented in THESEUS 20, a program for superpositioning macromolecular structures ANSI C source code and selected compiled binaries for various computing platforms are freely available under the GNU open source license from http://wwwtheseus3dorg
Contact: dtheobald@brandeisedu
Supplementary information:Supplementary data are available at Bioinformatics online
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Determinants of Base Editing Outcomes from Target Library Analysis and Machine Learning
Mandana Arbab,Mandana Arbab,Mandana Arbab,Max W. Shen,Beverly Mok,Beverly Mok,Beverly Mok,Christine D. Wilson,Christine D. Wilson,Christine D. Wilson,Żaneta Matuszek,Christopher A. Cassa,Christopher A. Cassa,David R. Liu,David R. Liu,David R. Liu +15 more
TL;DR: This work characterized sequence-activity relationships of cytosine and adenine base editors and used the resulting outcomes to train BE-Hive, a machine learning model that accurately predicts base editing genotypic outcomes and engineer novel CBE variants that modulate editing outcomes.
230
Conformational analysis of the DFG-out kinase motif and biochemical profiling of structurally validated type II inhibitors.
R. S. K. Vijayan,Peng He,Vivek Modi,Krisna C. Duong-Ly,Haiching Ma,Jeffrey R. Peterson,Roland L. Dunbrack,Ronald M. Levy +7 more
TL;DR: It is discovered that the number of structurally validated type II inhibitors that can be found in the PDB and that are also represented in publicly available biochemical profiling studies of kinase inhibitors is very small.
201
Evolutionary drivers of thermoadaptation in enzyme catalysis
Vy Nguyen,Christine D. Wilson,Marc Hoemberger,John B. Stiller,Roman V. Agafonov,Steffen Kutter,Justin English,Douglas L. Theobald,Dorothee Kern +8 more
TL;DR: It is shown that evolution solved the enzyme’s key kinetic obstacle—how to maintain catalytic speed on a cooler Earth—by exploiting transition-state heat capacity and suggests that the catalyticspeed of adenylate kinase is an evolutionary driver for organismal fitness.
The energy landscape of adenylate kinase during catalysis.
S. Jordan Kerns,Roman V. Agafonov,Young-Jin Cho,Francesco Pontiggia,Renee Otten,Dimitar V. Pachov,Steffen Kutter,Lien A. Phung,Padraig Niall Murphy,Vu Hong Thai,Tom Alber,Michael F. Hagan,Dorothee Kern +12 more
TL;DR: The results highlight the importance of the entire energy landscape in catalysis and suggest that adenylate kinases have evolved to activate key processes simultaneously by precise placement of a single, charged and very abundant cofactor in a preorganized active site.
179
References
•Book
The EM algorithm and extensions
Geoffrey J. McLachlan,Thriyambakam Krishnan +1 more
- 15 Nov 1996
TL;DR: The EM Algorithm and Extensions describes the formulation of the EM algorithm, details its methodology, discusses its implementation, and illustrates applications in many statistical contexts, opening the door to the tremendous potential of this remarkably versatile statistical tool.
•Book
Methods of Biochemical Analysis
David Glick
- 15 Jan 1968
TL;DR: The Radiation Inactivation Method as a Tool to Study Structure-Function Relationships in Proteins Immunoassay with Electrochemical Detection and Theory and Newer Technology Assays for Superoxide Dismutase are studied.
4.6K
Generalized procrustes analysis
TL;DR: In this article, the authors investigated the problem of translating, rotating, reflecting and scaling configurations to minimize the goodness-of-fit criterion, where Gi is the centroid of the points in p-dimensional space.
3.3K
•Book
Statistical Shape Analysis: With Applications in R
Ian L. Dryden,Kanti V. Mardia +1 more
- 06 Sep 2016
TL;DR: In this article, the authors proposed a planar procrustes analysis for two-dimensional data and showed that it is possible to estimate the size and shape of a shape in images.