Scinovex
articleTop 1% cited

Combining evolutionary information and neural networks to predict protein secondary structure

Proteins Structure Function and Bioinformatics · 1994 · Vol. 19(1) · pp. 55–72
Burkhard RostChris Sander

Abstract

Using evolutionary information contained in multiple sequence alignments as input to neural networks, secondary structure can be predicted at significantly increased accuracy. Here, we extend our previous three-level system of neural networks by using additional input information derived from multiple alignments. Using a position-specific conservation weight as part of the input increases performance. Using the number of insertions and deletions reduces the tendency for overprediction and increases overall accuracy. Addition of the global amino acid content yields a further improvement, mainly in predicting structural class. The final network system has sustained overall accuracy of 71.6% in a multiple cross-validation test on 126 unique protein chains. A test on a new set of 124 recently solved protein structures that have no significant sequence similarity to the learning set confirms the high level of accuracy. The average cross-validated accuracy for all 250 sequence-unique chains is above 72%. Using various data sets, the method is compared to alternative prediction methods, some of which also use multiple alignments: the performance advantage of the network system is at least 6 percentage points in three-state accuracy. In addition, the network estimates secondary structure content from multiple sequence alignments about as well as circular dichroism spectroscopy on a single protein and classifies 75% of the 250 proteins correctly into one of four protein structural classes. Of particular practical importance is the definition of a position-specific reliability index. For 40% of all residues the method has a sustained three-state accuracy of 88%, as high as the overall average for homology modelling. A further strength of the method is greatly increased accuracy in predicting the placement of secondary structure segments.

Protein Structure and DynamicsMachine Learning in BioinformaticsRNA and protein synthesis mechanismsProtein secondary structureArtificial neural networkSequence (biology)Computer scienceSet (abstract data type)Similarity (geometry)Reliability (semiconductor)Artificial intelligenceTest setPosition (finance)

MeSH terms

Amino Acid SequenceBiological EvolutionSequence AlignmentNeural Networks, ComputerProtein Structure, Secondary
Citations
1,462
FWCI
41.57
field-weighted impact
References
136
Percentile
100%
vs. same field & year
Citations per year
Cited by
Machine Learning and Its Applications to Biology
PLoS Computational Biology · 2007 · 657 citations
The 26S Proteasome: A Molecular Machine Designed for Controlled Proteolysis
Annual Review of Biochemistry · 1999 · 1,887 citations
Length-dependent prediction of protein intrinsic disorder
BMC Bioinformatics · 2006 · 946 citations
Transmembrane helices predicted at 95% accuracy
Protein Science · 1995 · 669 citations
Conservation and prediction of solvent accessibility in protein families
Proteins Structure Function and Bioinformatics · 1994 · 674 citations
Genealogy of the α-crystallin—small heat-shock protein superfamily
International Journal of Biological Macromolecules · 1998 · 490 citations
Biochemical and genetic analysis of PHA synthases and other proteins required for PHA synthesis
International Journal of Biological Macromolecules · 1999 · 398 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.