Scinovex
article Open AccessTop 1% cited

Annotation Error in Public Databases: Misannotation of Molecular Function in Enzyme Superfamilies

PLoS Computational Biology · 2009 · Vol. 5(12) · pp. e1000605–e1000605
Alexandra M. SchnoesShoshana BrownIgor DodevskiPatricia C. Babbitt

Abstract

Due to the rapid release of new data from genome sequencing projects, the majority of protein sequences in public databases have not been experimentally characterized; rather, sequences are annotated using computational analysis. The level of misannotation and the types of misannotation in large public databases are currently unknown and have not been analyzed in depth. We have investigated the misannotation levels for molecular function in four public protein sequence databases (UniProtKB/Swiss-Prot, GenBank NR, UniProtKB/TrEMBL, and KEGG) for a model set of 37 enzyme families for which extensive experimental information is available. The manually curated database Swiss-Prot shows the lowest annotation error levels (close to 0% for most families); the two other protein sequence databases (GenBank NR and TrEMBL) and the protein sequences in the KEGG pathways database exhibit similar and surprisingly high levels of misannotation that average 5%-63% across the six superfamilies studied. For 10 of the 37 families examined, the level of misannotation in one or more of these databases is >80%. Examination of the NR database over time shows that misannotation has increased from 1993 to 2005. The types of misannotation that were found fall into several categories, most associated with "overprediction" of molecular function. These results suggest that misannotation in enzyme superfamilies containing multiple families that catalyze different reactions is a larger problem than has been recognized. Strategies are suggested for addressing some of the systematic problems contributing to these high levels of misannotation.

Machine Learning in BioinformaticsGenomics and Phylogenetic StudiesAdvanced Proteomics Techniques and ApplicationsUniProtGenBankKEGGAnnotationDatabaseSequence databaseBiologyFunction (biology)Computational biologyProtein sequencing

MeSH terms

Database Management SystemsDatabases, ProteinBiocatalysis

Funding

  • National Science Foundation
  • National Institutes of Health
Citations
704
FWCI
14.39
field-weighted impact
References
75
Percentile
99%
vs. same field & year
Citations per year
Cited by
NetSurfP‐2.0: Improved prediction of protein structural features by integrated deep learning
Proteins Structure Function and Bioinformatics · 2019 · 602 citations
Embracing the unknown: disentangling the complexities of the soil microbiome
Nature Reviews Microbiology · 2017 · 3,256 citations
Application of metagenomics in the human gut microbiome
World Journal of Gastroenterology · 2015 · 399 citations
References
Gene Ontology: tool for the unification of biology
Nature Genetics · 2000 · 43,975 citations
Conservation of gene order: a fingerprint of proteins that physically interact
Trends in Biochemical Sciences · 1998 · 1,055 citations
KEGG for linking genomes to life and the environment
Nucleic Acids Research · 2007 · 6,968 citations
MUSCLE: multiple sequence alignment with high accuracy and high throughput
Nucleic Acids Research · 2004 · 45,964 citations
The Pfam Protein Families Database
Nucleic Acids Research · 2002 · 14,220 citations
The COG database: an updated version includes eukaryotes
BMC Bioinformatics · 2003 · 4,484 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

Annotation Error in Public Databases: Misannotation of Molecular Function in Enzyme Superfamilies · Scinovex