ClassyFire: automated chemical classification with a comprehensive, computable taxonomy
Yannick Djoumbou Feunang,Roman Eisner,Craig Knox,Leonid L. Chepelev,Janna Hastings,Gareth Owen,Eoin Fahy,Christoph Steinbeck,Shankar Subramanian,Evan E Bolton,Russell Greiner,David S. Wishart +11 more
TL;DR: A comprehensive, flexible, and computable, purely structure-based chemical taxonomy (ChemOnt) is developed along with a computer program (ClassyFire) that uses only chemical structures and structural features to automatically assign all known chemical compounds to a taxonomy consisting of >4800 different categories.
read more
Abstract: Scientists have long been driven by the desire to describe, organize, classify, and compare objects using taxonomies and/or ontologies. In contrast to biology, geology, and many other scientific disciplines, the world of chemistry still lacks a standardized chemical ontology or taxonomy. Several attempts at chemical classification have been made; but they have mostly been limited to either manual, or semi-automated proof-of-principle applications. This is regrettable as comprehensive chemical classification and description tools could not only improve our understanding of chemistry but also improve the linkage between chemistry and many other fields. For instance, the chemical classification of a compound could help predict its metabolic fate in humans, its druggability or potential hazards associated with it, among others. However, the sheer number (tens of millions of compounds) and complexity of chemical structures is such that any manual classification effort would prove to be near impossible. We have developed a comprehensive, flexible, and computable, purely structure-based chemical taxonomy (ChemOnt), along with a computer program (ClassyFire) that uses only chemical structures and structural features to automatically assign all known chemical compounds to a taxonomy consisting of >4800 different categories. This new chemical taxonomy consists of up to 11 different levels (Kingdom, SuperClass, Class, SubClass, etc.) with each of the categories defined by unambiguous, computable structural rules. Furthermore each category is named using a consensus-based nomenclature and described (in English) based on the characteristic common structural properties of the compounds it contains. The ClassyFire webserver is freely accessible at http://classyfire.wishartlab.com/
. Moreover, a Ruby API version is available at https://bitbucket.org/wishartlab/classyfire_api
, which provides programmatic access to the ClassyFire server and database. ClassyFire has been used to annotate over 77 million compounds and has already been integrated into other software packages to automatically generate textual descriptions for, and/or infer biological properties of over 100,000 compounds. Additional examples and applications are provided in this paper. ClassyFire, in combination with ChemOnt (ClassyFire’s comprehensive chemical taxonomy), now allows chemists and cheminformaticians to perform large-scale, rapid and automated chemical classification. Moreover, a freely accessible API allows easy access to more than 77 million “ClassyFire” classified compounds. The results can be used to help annotate well studied, as well as lesser-known compounds. In addition, these chemical classifications can be used as input for data integration, and many other cheminformatics-related tasks.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Impacts of long-term organic fertilization on metabolomic and metagenomic characteristics of soils in a greenhouse vegetable production system
Yun‐Cheng Hsieh,Chun‐Han Su,Tzung-Han Lee,Lean‐Teik Ng +3 more
Ex vivo study of molecular changes of stained teeth following hydrogen peroxide and peroxymonosulfate treatments
Paulo Wender P. Gomes,Simone Zuffa,Anelize Bauermeister,Andrés Mauricio Caraballo-Rodríguez,Haoqi Nina Zhao,Helena Mannochio-Russo,Cajetan Dogo-isonagie,Om Patel,Paloma Pimenta,Jennifer Gronlund,Stacey Lavender,Shira Pilch,V. Maloney,Michael North,Pieter C. Dorrestein +14 more
TL;DR: Show how two different tooth-whitening peroxides could affect the molecular profiles of human teeth by comparing the molecular profiles of teeth bleached with hydrogen peroxide and peroxymonosulfate.
Exploring the Exclusive Isolation of Pseudomonas syringae in Peltigera Lichens via metabolite analysis and growth assays
Natalia Ramírez,Diana Marcela Vinchira-Villarraga,Mojgan Rabiey,Margrét Auður Sigurbjörnsdóttir,Starri Heiđmarsson,Oddur Vilhelmsson,Robert W. Jackson +6 more
- 06 Sep 2024
TL;DR: This study investigates the exclusive association of Pseudomonas syringae with Peltigera lichens, revealing lower metabolite richness in Peltigera but higher investment in specific compounds, which may facilitate P. syringae growth and inhibit antagonists.
Integrated targeted and untargeted metabolomics profiling of Vanilla species from the Atlantic Forest: Unveiling the bioeconomic potential of Vanilla cribbiana
Joana Paula da Silva Oliveira,Rafael Garrett,Maria Gabriela Bello Koblitz,Andrea Furtado Macedo +3 more
Abstract: In this study, we employ both targeted and untargeted approaches to explore the metabolomic profiles of Vanilla spp., with a particular focus on V. cribbiana (VCR) and its comparison with V. planifolia (VP). We also examine V. bahiana and V. chamissonis using targeted approaches. Through advanced analytical techniques, our untargeted LC-HRMS approach led to the annotation of 60 metabolites, revealing a complex chemical composition with 34 novel compounds in the Vanilla genus in VCR and VP. These findings highlight significant flavoring compounds and lay the foundation for a subsequent quantitative estimation approach. Our targeted analysis, which measured key molecules, underscores VCR's potential in producing vanillin and acetovanillone at levels comparable to the commercially valuable VP and even higher levels of vanillic acid. This research enriches our understanding of flavor composition in vanilla species and emphasizes the importance of exploring wild relatives of vanilla crop for sustainable production and biodiversity conservation.
Metabolome responses of Enterococcus faecium to acid shock and nitrite stress
TL;DR: The approach uncovered the hidden interactions between intracellular metabolites and exogenous stress, and will improve the understanding of host‐microbe interactions.
References
Gene Ontology: tool for the unification of biology
M Ashburner,Catherine A. Ball,Judith A. Blake,David Botstein,Heather Butler,J. M. Cherry,Allan Peter Davis,Kara Dolinski,Selina S. Dwight,J.T. Eppig,Midori A. Harris,David P. Hill,Laurie Issel-Tarver,Andrew Kasarskis,Suzanna E. Lewis,John C. Matese,Joel E. Richardson,M. Ringwald,Gerald M. Rubin,Gavin Sherlock +19 more
TL;DR: The goal of the Gene Ontology Consortium is to produce a dynamic, controlled vocabulary that can be applied to all eukaryotes even as knowledge of gene and protein roles in cells is accumulating and changing.
Toward principles for the design of ontologies used for knowledge sharing
TL;DR: The role of ontology in supporting knowledge sharing activities is described, and a set of criteria to guide the development of ontologies for these purposes are presented, and it is shown how these criteria are applied in case studies from the design ofOntologies for engineering mathematics and bibliographic data.
KEGG as a reference resource for gene and protein annotation
TL;DR: The KEGG GENES database now includes viruses, plasmids, and the addendum category for functionally characterized proteins that are not represented in complete genomes, and new automatic annotation servers, BlastKOalA and GhostKOALA, are made available utilizing the non-redundant pangenome data set generated from theGENES database.
SMILES, a chemical language and information system. 1. introduction to methodology and encoding rules
TL;DR: This chapter discusses the construction of Benzenoid and Coronoid Hydrocarbons through the stages of enumeration, classification, and topological properties in a number of computers used for this purpose.
6.7K
PubChem Substance and Compound databases
Sunghwan Kim,Paul A. Thiessen,Evan E Bolton,Jie Chen,Gang Fu,Asta Gindulyte,Lianyi Han,Jane He,Siqian He,Benjamin A. Shoemaker,Jiyao Wang,Bo Yu,Jian-Jian Zhang,Stephen H. Bryant +13 more
TL;DR: An overview of the PubChem Substance and Compound databases is provided, including data sources and contents, data organization, data submission using PubChem Upload, chemical structure standardization, web-based interfaces for textual and non-textual searches, and programmatic access.
4.7K
Related Papers (5)
Mingxun Wang,Jeremy Carver,Vanessa V. Phelan,Laura M. Sanchez,Neha Garg,Yao Peng,Don D. Nguyen,Jeramie D. Watrous,Clifford A. Kapono,Tal Luzzatto-Knaan,Carla Porto,Amina Bouslimani,Alexey V. Melnik,Michael J. Meehan,Wei-Ting Liu,Max Crüsemann,Paul D. Boudreau,Eduardo Esquenazi,Mario Sandoval-Calderón,Roland D. Kersten,Laura A. Pace,Robert A. Quinn,Katherine R. Duncan,Cheng-Chih Hsu,Dimitrios J. Floros,Ronnie G. Gavilan,Karin Kleigrewe,Trent R. Northen,Rachel J. Dutton,Delphine Parrot,Erin E. Carlson,Bertrand Aigle,Charlotte Frydenlund Michelsen,Lars Jelsbak,Christian Sohlenkamp,Pavel A. Pevzner,Anna Edlund,Anna Edlund,Jeffrey S. McLean,Jeffrey S. McLean,Jörn Piel,Brian T. Murphy,Lena Gerwick,Chih-Chuang Liaw,Yu-Liang Yang,Hans-Ulrich Humpf,Maria Maansson,Robert A. Keyzers,Amy C. Sims,Andrew R. Johnson,Ashley M. Sidebottom,Brian E. Sedio,Andreas Klitgaard,Charles B. Larson,Charles B. Larson,Cristopher A. Boya P.,Daniel Torres-Mendoza,David Gonzalez,Denise Brentan Silva,Denise Brentan Silva,Lucas Miranda Marques,Daniel P. Demarque,Egle Pociute,Ellis C. O’Neill,Enora Briand,Enora Briand,Eric J. N. Helfrich,Eve A. Granatosky,Evgenia Glukhov,Florian Ryffel,Hailey Houson,Hosein Mohimani,Jenan J. Kharbush,Yi Zeng,Julia A. Vorholt,Kenji L. Kurita,Pep Charusanti,Kerry L. McPhail,Kristian Fog Nielsen,Lisa Vuong,Maryam Elfeki,Matthew F. Traxler,Niclas Engene,Nobuhiro Koyama,Oliver B. Vining,Ralph S. Baric,Ricardo Pianta Rodrigues da Silva,Samantha J. Mascuch,Sophie Tomasi,Stefan Jenkins,Venkat R. Macherla,Thomas Hoffman,Vinayak Agarwal,Philip G. Williams,Jingqui Dai,Ram P. Neupane,Joshua R. Gurr,Andrés M. C. Rodríguez,Anne Lamsa,Chen Zhang,Kathleen Dorrestein,Brendan M. Duggan,Jehad Almaliti,Pierre-Marie Allard,Prasad Phapale,Louis-Félix Nothias,Theodore Alexandrov,Marc Litaudon,Jean-Luc Wolfender,Jennifer E. Kyle,Thomas O. Metz,Tyler Peryea,Dac-Trung Nguyen,Danielle VanLeer,Paul Shinn,Ajit Jadhav,Rolf Müller,Katrina M. Waters,Wenyuan Shi,Xueting Liu,Lixin Zhang,Rob Knight,Paul R. Jensen,Bernhard O. Palsson,Kit Pogliano,Roger G. Linington,Marcelino Gutiérrez,Norberto Peporine Lopes,William H. Gerwick,William H. Gerwick,Bradley S. Moore,Bradley S. Moore,Pieter C. Dorrestein,Pieter C. Dorrestein,Nuno Bandeira,Nuno Bandeira +135 more
Lloyd W. Sumner,Alexander Amberg,Dave Barrett,Michael H. Beale,Richard D. Beger,Clare A. Daykin,Teresa W.-M. Fan,Oliver Fiehn,Royston Goodacre,Julian L. Griffin,Thomas Hankemeier,Nigel Hardy,James M. Harnly,Richard M. Higashi,Joachim Kopka,Andrew N. Lane,John C. Lindon,Philip J. Marriott,Andrew W. Nicholls,Michael D. Reily,John J. Thaden,Mark R. Viant +21 more