Almut Heinken1, T. Hulshof, Bram Nap2, Filippo Martinelli1, Arianna Basile, Amy O’Brolchain, Neil Francis O’Sullivan, Celine Gallagher, Eimer Magee, Francesca McDonagh, Ian Lalor, Maeve Bergin, Phoebe Evans, Rachel Daly, Ronan Farrell3, Rose Mary Delaney, Saoirse Hill, Saoirse Roisin McAuliffe, Trevor Kilgannon, Ronan M. T. Fleming4, Cyrille C Thinnes, Ines Thiele•
Abstract: Genome-scale modeling of microbiome metabolism enables the simulation of diet-host-microbiome-disease interactions. However, current genome-scale reconstruction resources are limited in scope by computational challenges. We developed an optimized and highly parallelized reconstruction and analysis pipeline to build a resource of 247,092 microbial genome-scale metabolic reconstructions, deemed APOLLO. APOLLO spans 19 phyla, contains >60% of uncharacterized strains, and accounts for strains from 34 countries, all age groups, and multiple body sites. Using machine learning, we predicted with high accuracy the taxonomic assignment of strains based on the computed metabolic features. We then built 14,451 metagenomic sample-specific microbiome community models to systematically interrogate their community-level metabolic capabilities. We show that sample-specific metabolic pathways accurately stratify microbiomes by body site, age, and disease state. APOLLO is freely available, enables the systematic interrogation of the metabolic capabilities of largely still uncultured and unclassified species, and provides unprecedented opportunities for systems-level modeling of personalized host-microbiome co-metabolism.
Ryan Z. Friedman, Avinash Ramu, Sara Lichtarge, Yawei Wu, Lloyd D. Tripp, Delina Y. Lyon, Connie A. Myers, David M. Granas, Maria Gause, Joseph C. Corbo, Barak A. Cohen, Michael A. White
TL;DR: Researchers develop an active learning approach to train deep learning models that distinguish between enhancers and silencers in the developing neural retina, leveraging synthetic biology and uncertainty sampling to iteratively generate informative training data.
Abstract: Highlights•Transcription factor binding sites activate or repress depending on context•Genomic examples are insufficient to learn how context affects binding sites•Active learning iteratively generates informative new training data•A CNN trained with active learning distinguishes activating and repressing sitesSummaryDeep learning is a promising strategy for modeling cis-regulatory elements. However, models trained on genomic sequences often fail to explain why the same transcription factor can activate or repress transcription in different contexts. To address this limitation, we developed an active learning approach to train models that distinguish between enhancers and silencers composed of binding sites for the photoreceptor transcription factor cone-rod homeobox (CRX). After training the model on nearly all bound CRX sites from the genome, we coupled synthetic biology with uncertainty sampling to generate additional rounds of informative training data. This allowed us to iteratively train models on data from multiple rounds of massively parallel reporter assays. The ability of the resulting models to discriminate between CRX sites with identical sequence but opposite functions establishes active learning as an effective strategy to train models of regulatory DNA. A record of this paper's transparent peer review process is included in the supplemental information.Graphical abstract
Abstract: The three-dimensional (3D) morphology of cells emerges from complex cellular and environmental interactions, serving as an indicator of cell state and function. In this study, we used deep learning to discover morphology representations and understand cell states. This study introduced MorphoMIL, a computational pipeline combining geometric deep learning and attention-based multiple-instance learning to profile 3D cell and nuclear shapes. We used 3D point-cloud input and captured morphological signatures at single-cell and population levels, accounting for phenotypic heterogeneity. We applied these methods to over 95,000 melanoma cells treated with clinically relevant and cytoskeleton-modulating chemical and genetic perturbations. The pipeline accurately predicted drug perturbations and cell states. Our framework revealed subtle morphological changes associated with perturbations, key shapes correlating with signaling activity, and interpretable insights into cell-state heterogeneity. MorphoMIL demonstrated superior performance and generalized across diverse datasets, paving the way for scalable, high-throughput morphological profiling in drug discovery. A record of this paper's transparent peer review process is included in the supplemental information.
Helen O. Masson, Jasmine Tat, Pablo Di Giusto, Athanasios Antonakoudis, Isaac Shamie, Hratch Baghdassarian, Mojtaba Samoudi, Caressa M. Robinson, Chih-Chung Kuo, Natália Massaco Koga, Sonia Singh, Angel Gezalyan, Zerong Li, Alexia Movsessian, Anne Richelle, Nathan E. Lewis
TL;DR: Researchers developed an automated pipeline to reconstruct 495 metabolic and gene expression models, integrating them with multi-omics data to identify microbial associations with inflammatory bowel disease, providing testable hypotheses on gut microbiota activity.
Abstract: The gut microbiome plays a critical role in human health, spurring extensive research using multi-omic technologies. Although these tools offer valuable insights, they often fall short in capturing the complexity of microbial interactions that associate with disease onset, progression, and treatment. Thus, integration of multi-omics datasets with metabolic models is needed to predict associations between microbial activity and disease. Here, we automated the reconstruction of 495 metabolic and gene expression models (ME-models), overcoming the main limitation preventing the wide use of this approach. We integrated them with multi-omics data from patients with inflammatory bowel disease (IBD), identifying taxa associated with variations in amino acids, short-chain fatty acids, and pH in the gut of IBD patients. In general, this approach provides testable hypotheses of the metabolic activity of the gut microbiota, and the automated pipeline opens the opportunity to study microbial interactions in other biologically relevant settings using ME-models.
Elliot L. Chaikof, Jichao Chen, Martha U. Gillette, Laurie A. Boyer, Tara L. Deans, Pulin Li, Isaac B. Hilton, Kyle Daniels, Yogesh Goyal, Ying Mei, Chang-di Linghu, Theresa B. Loveless, David M. Truong, Michael R. Blatchley1, Mingxia Gu, Caleb J. Bashor2, Jason H. Yang, Ritu Raman, Akhilesh B. Reddy, Krishanu Saha, Jennifer Davis, Kalpna Gupta, Xiaojing J. Gao3, Kate E. Galloway•
TL;DR: Researchers reveal a core passive mechanism and facultative mTOR-mediated mechanisms that coordinate protein synthesis and decay in mammalian cells, maintaining cellular homeostasis and proteome integrity despite variations in protein synthesis rates.
Abstract: The maintenance of cellular homeostasis requires tight regulation of proteome concentration and composition. To achieve this, protein production and elimination must be robustly coordinated. However, the mechanistic basis of this coordination remains unclear. Here, we address this question using quantitative live-cell imaging, computational modeling, transcriptomics, and proteomics approaches. We found that protein decay rates systematically adapt to global alterations of protein synthesis rates. This adaptation is driven by a core passive mechanism supplemented by facultative changes in mechanistic/mammalian target of rapamycin (mTOR) signaling. Passive adaptation hinges on changes in the production rate of the machinery governing protein decay and allows for partial maintenance of the cellular proteome. Sustained changes in mTOR signaling provide an additional layer of adaptation unique to naive pluripotent stem cells, allowing for near-perfect maintenance of proteome composition. Our work unravels the mechanisms protecting the integrity of mammalian proteomes upon variations in protein synthesis rates. A record of this paper's transparent peer review process is included in the supplemental information.
Abstract: The circulating antibody (Ab) repertoire is crucial for immune protection, holding significant immunological and biotechnological value. While bottom-up mass spectrometry (MS) is widely used for profiling the sequence diversity of circulating Abs (Ab repertoire sequencing [Ab-seq]), it has not been thoroughly benchmarked. We quantified the replicability and robustness of Ab-seq using six monoclonal Ab spike-ins in 70 combinations of concentration and oligoclonality, with and without polyclonal serum immunoglobulin G (IgG) background. Each combination underwent four protease treatments and was analyzed across four experimental and three technical replicates, totaling 3,360 liquid chromatography-tandem MS (LC-MS/MS) runs. We quantified the dependence of Ab-seq identification on Ab sequence, concentration, protease, presence of background IgGs, and bioinformatics methods. Integrating the data from experimental replicates, proteases, and bioinformatics tools enhanced Ab identification. De novo sequencing performed similarly to database-dependent methods at higher Ab concentrations, but de novo Ab reconstruction remains challenging. Our work provides a foundational resource for the field of MS-based Ab profiling. A record of this paper's transparent peer review process is included in the supplemental information.