Open Access
SMART, a simple modular architecture research tool: Identification of signaling domains (computer analysisydiacylglycerol kinasesyDEATH domainydisease genesyautomatic sequence annotation)
F Rank Milpetz,P Eer Bork,Chris P. Ponting +2 more
- 01 Jan 1998
3.1K
TL;DR: SMART as discussed by the authors is a web-based tool that allows rapid identification and annotation of signaling domain sequences, which can be used to determine the modular architectures of single sequences or genomes.
read more
Abstract: Accurate multiple alignments of 86 domains that occur in signaling proteins have been constructed and used to provide a Web-based tool (SMART: simple modular architecture research tool) that allows rapid identification and annotation of signaling domain sequences. The majority of signaling proteins are multidomain in character with a considerable variety of domain combinations known. Com- parison with established databases showed that 25% of our domain set could not be deduced from SwissProt and 41% could not be annotated by Pfam. SMART is able to determine the modular architectures of single sequences or genomes; application to the entire yeast genome revealed that at least 6.7% of its genes contain one or more signaling domains, approximately 350 greater than previously annotated. The process of constructing SMART predicted (i) novel domain homologues in unexpected locations such as band 4.1- homologous domains in focal adhesion kinases; (ii) previously unknown domain families, including a citron-homology do- main; (iii) putative functions of domain families after identi- fication of additional family members, for example, a ubiq- uitin-binding role for ubiquitin-associated domains (UBA); (iv) cellular roles for proteins, such predicted DEATH do- mains in netrin receptors further implicating these molecules in axonal guidance; (v) signaling domains in known disease genes such as SPRY domains in both marenostrinypyrin and Midline 1; (vi) domains in unexpected phylogenetic contexts such as diacylglycerol kinase homologues in yeast and bacte- ria; and (vii) likely protein misclassifications exemplified by a predicted pleckstrin homology domain in a Candida albicans protein, previously described as an integrin.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
MUSCLE: multiple sequence alignment with high accuracy and high throughput
TL;DR: MUSCLE is a new computer program for creating multiple alignments of protein sequences that includes fast distance estimation using kmer counting, progressive alignment using a new profile function the authors call the log-expectation score, and refinement using tree-dependent restricted partitioning.
45.1K
MUSCLE: a multiple sequence alignment method with reduced time and space complexity
TL;DR: MUSCLE offers a range of options that provide improved speed and / or alignment accuracy compared with currently available programs, and a new option, MUSCLE-fast, designed for high-throughput applications.
The COG database: an updated version includes eukaryotes
Roman L. Tatusov,Natalie D. Fedorova,John D. Jackson,Aviva R. Jacobs,Boris Kiryutin,Eugene V. Koonin,Dmitri M. Krylov,Raja Mazumder,Sergei L. Mekhedov,Anastasia N. Nikolskaya,B Sridhar Rao,Sergei Smirnov,Alexander V. Sverdlov,Sona Vasudevan,Yuri I. Wolf,Jodie J. Yin,Darren A. Natale +16 more
TL;DR: A major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes is described and is expected to be a useful platform for functional annotation of newlysequenced genomes, including those of complex eukARYotes, and genome-wide evolutionary studies.
PANTHER: a library of protein families and subfamilies indexed by function.
Paul Thomas,Michael J. Campbell,Anish Kejariwal,Huaiyu Mi,Brian Karlak,Robin Daverman,Karen Diemer,Anushya Muruganujan,Apurva Narechania +8 more
TL;DR: The PANTHER/X ontology is used to give a high-level representation of gene function across the human and mouse genomes, and the family HMMs are used to rank missense single nucleotide polymorphisms (SNPs) according to their likelihood of affecting protein function.
The Calpain System
TL;DR: How calpain activity is regulated in cells is still unclear, but the calpains ostensibly participate in a variety of cellular processes including remodeling of cytoskeletal/membrane attachments, different signal transduction pathways, and apoptosis.
2.9K
References
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Stephen F. Altschul,Thomas L. Madden,Alejandro A. Schäffer,Jinghui Zhang,Zheng Zhang,Webb Miller,David J. Lipman +6 more
TL;DR: A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original.
Clustal w: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice
TL;DR: The sensitivity of the commonly used progressive multiple sequence alignment method has been greatly improved and modifications are incorporated into a new program, CLUSTAL W, which is freely available.
Predicting coiled coils from protein sequences
TL;DR: This method was used to delineate coiled-coil domains in otherwise globular proteins, such as the leucine zipper domains in transcriptional regulators, and to predict regions of discontinuity within coiled -coil structures,such as the hinge region in myosin.
4.2K
Proteins regulating Ras and its relatives.
Mark S. Boguski,Frank McCormick +1 more
TL;DR: Many of these proteins are much larger and more complex than their targets, containing multiple domains capable of interacting with an intricate network of cellular enzymes and structures.
2K
The PROSITE database, its status in 1997
TL;DR: The PROSITE database (http://www.expasy.ch/sprot/prosite.htm l) consists of biologically significant patterns and profiles formulated in such a way that with appropriate computational tools it can help to determine to which known family of protein a new sequence belongs, or which known domain(s) it contains.
Related Papers (5)
Etienne Formstecher,Sandra Aresta,Vincent Collura,Alexandre Hamburger,Alain Meil,Alexandra Trehin,Céline Reverdy,Virginie Betin,Sophie Maire,Christine Brun,Bernard Jacq,Monique Arpin,Yohanns Bellaïche,Saverio Bellusci,Philippe Benaroch,Michel Bornens,Roland Chanet,Philippe Chavrier,Olivier Delattre,Valérie Doye,Richard G. Fehon,Richard G. Fehon,Gérard Faye,Thierry Galli,Jean-Antoine Girault,Bruno Goud,Jean de Gunzburg,Ludger Johannes,Marie-Pierre Junier,Vincent Mirouse,Ashim Mukherjee,Dora Papadopoulo,Franck Perez,Anne Plessis,Carine Rossé,Simon Saule,Dominique Stoppa-Lyonnet,Alain Vincent,Michael A. White,Pierre Legrain,Jérôme Wojcik,Jacques Camonis,Jacques Camonis,Laurent Daviet +43 more