Journal Article10.1038/NRM2281
Predicting protein function from sequence and structure
TL;DR: There is an increasing number of noteworthy methods for predicting protein function from sequence and structural data alone, many of which are readily available to cell biologists who are aware of the strengths and pitfalls of each available technique.
read more
Abstract: While the number of sequenced genomes continues to grow, experimentally verified functional annotation of whole genomes remains patchy. Structural genomics projects are yielding many protein structures that have unknown function. Nevertheless, subsequent experimental investigation is costly and time-consuming, which makes computational methods for predicting protein function very attractive. There is an increasing number of noteworthy methods for predicting protein function from sequence and structural data alone, many of which are readily available to cell biologists who are aware of the strengths and pitfalls of each available technique.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Evolutionary conservation and emerging functional diversity of the cytosolic Hsp70:J protein chaperone network of Arabidopsis thaliana
Amit Kumar Verma,Danish Diwan,Sandeep Raut,Neha Dobriyal,Rebecca E. Brown,Vinita Gowda,Justin K. Hines,Chandan Sahi +7 more
TL;DR: It is proposed that higher plants have orchestrated their “chaperome,” especially their J protein complement, according to their specialized cellular and physiological stipulations, to contribute to the emerging functional diversity and complexity in the Hsp70:J protein network in higher plants.
•Dissertation
Automatic and manual functional annotation in a distributed web service environment
Anika Jöcker
- 01 Jan 2009
TL;DR: A system for manual functional annotation, called AFAWE, which makes use of annotated functional attributes to calculate the functional mutation rate between genes and groups of genes in a phylogenetic tree and to find out if the function of a gene can be transferred or not and how to improve the annotation.
•Dissertation
Characterising functional diversity in protein domain superfamilies and metagenomes
NL Dawson
- 28 Feb 2015
TL;DR: Functional site diversity was explored with a much larger dataset of superfamilies, by examining residues involved in protein interfaces and catalytic sites, and a complex relationship was found between metagenome sequences and sequence/structure positions.
Sequence Motifs: Highly Predictive Features of Protein Function
Asa Ben-Hur,Douglas L. Brutlag +1 more
- 01 Jan 2006
TL;DR: It is shown that despite the fact that motif composition is a very high dimensional representation of a sequence, that most classes of enzymes can be classified using a handful of motifs, yielding accurate and interpretable classifiers.
GPU-based Cloud computing for comparing the structure of protein binding sites
Matthias Leinweber,Lars Baumgärtner,Marco Mernberger,Thomas Fober,Eyke Hüllermeier,Gerhard Klebe,Bernd Freisleben +6 more
- 18 Jun 2012
TL;DR: This new implementation of SEGA has been tested on a subset of protein structure data contained in the CavBase, providing a structural comparison of protein binding sites on a much larger scale than in previous research efforts reported in the literature.
References
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Stephen F. Altschul,Thomas L. Madden,Alejandro A. Schäffer,Jinghui Zhang,Zheng Zhang,Webb Miller,David J. Lipman +6 more
TL;DR: A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original.
Clustal w: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice
TL;DR: The sensitivity of the commonly used progressive multiple sequence alignment method has been greatly improved and modifications are incorporated into a new program, CLUSTAL W, which is freely available.
Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles
Aravind Subramanian,Pablo Tamayo,Vamsi K. Mootha,Sayan Mukherjee,Benjamin L. Ebert,Michael A. Gillette,Amanda G. Paulovich,Scott L. Pomeroy,Todd R. Golub,Eric S. Lander,Jill P. Mesirov +10 more
TL;DR: The Gene Set Enrichment Analysis (GSEA) method as discussed by the authors focuses on gene sets, that is, groups of genes that share common biological function, chromosomal location, or regulation.
MUSCLE: multiple sequence alignment with high accuracy and high throughput
TL;DR: MUSCLE is a new computer program for creating multiple alignments of protein sequences that includes fast distance estimation using kmer counting, progressive alignment using a new profile function the authors call the log-expectation score, and refinement using tree-dependent restricted partitioning.
45.1K
Gene Ontology: tool for the unification of biology
M Ashburner,Catherine A. Ball,Judith A. Blake,David Botstein,Heather Butler,J. M. Cherry,Allan Peter Davis,Kara Dolinski,Selina S. Dwight,J.T. Eppig,Midori A. Harris,David P. Hill,Laurie Issel-Tarver,Andrew Kasarskis,Suzanna E. Lewis,John C. Matese,Joel E. Richardson,M. Ringwald,Gerald M. Rubin,Gavin Sherlock +19 more
TL;DR: The goal of the Gene Ontology Consortium is to produce a dynamic, controlled vocabulary that can be applied to all eukaryotes even as knowledge of gene and protein roles in cells is accumulating and changing.
Related Papers (5)
M Ashburner,Catherine A. Ball,Judith A. Blake,David Botstein,Heather Butler,J. M. Cherry,Allan Peter Davis,Kara Dolinski,Selina S. Dwight,J.T. Eppig,Midori A. Harris,David P. Hill,Laurie Issel-Tarver,Andrew Kasarskis,Suzanna E. Lewis,John C. Matese,Joel E. Richardson,M. Ringwald,Gerald M. Rubin,Gavin Sherlock +19 more