Scinovex
article Open AccessTop 1% cited

NCBI prokaryotic genome annotation pipeline

Nucleic Acids Research · 2016 · Vol. 44(14) · pp. 6614–6624
Tatiana TatusovaMichael DiCuccioAzat BadretdinVyacheslav ChetverninEric P. NawrockiLeonid ZaslavskyAlexandre LomsadzeKim D. PruittMark BorodovskyJames Ostell

Abstract

Recent technological advances have opened unprecedented opportunities for large-scale sequencing and analysis of populations of pathogenic species in disease outbreaks, as well as for large-scale diversity studies aimed at expanding our knowledge across the whole domain of prokaryotes. To meet the challenge of timely interpretation of structure, function and meaning of this vast genetic information, a comprehensive approach to automatic genome annotation is critically needed. In collaboration with Georgia Tech, NCBI has developed a new approach to genome annotation that combines alignment based methods with methods of predicting protein-coding and RNA genes and other functional elements directly from sequence. A new gene finding tool, GeneMarkS+, uses the combined evidence of protein and RNA placement by homology as an initial map of annotation to generate and modify ab initio gene predictions across the whole genome. Thus, the new NCBI's Prokaryotic Genome Annotation Pipeline (PGAP) relies more on sequence similarity when confident comparative data are available, while it relies more on statistical predictions in the absence of external evidence. The pipeline provides a framework for generation and analysis of annotation on the full breadth of prokaryotic taxonomy. For additional information on PGAP see https://www.ncbi.nlm.nih.gov/genome/annotation_prok/ and the NCBI Handbook, https://www.ncbi.nlm.nih.gov/books/NBK174280/.

Genomics and Phylogenetic StudiesRNA and protein synthesis mechanismsBacteriophages and microbial interactionsAnnotationGenomeBiologyGenome projectComputational biologyGene AnnotationWhole genome sequencingGene predictionDECIPHERGene

MeSH terms

BacteriaBacterial ProteinsGenes, BacterialProkaryotic CellsGenome, BacterialDatabases, Nucleic AcidMolecular Sequence Annotation

Funding

  • Georgia Institute of Technology
Citations
6,790
FWCI
183.64
field-weighted impact
References
39
Percentile
100%
vs. same field & year
Citations per year
References
Prokka: rapid prokaryotic genome annotation
Bioinformatics · 2014 · 18,891 citations
Search and clustering orders of magnitude faster than BLAST
Bioinformatics · 2010 · 21,473 citations
Infernal 1.1: 100-fold faster RNA homology searches
Bioinformatics · 2013 · 3,809 citations
BLAST+: architecture and applications
BMC Bioinformatics · 2009 · 22,452 citations
Evolution and classification of the CRISPR–Cas systems
Nature Reviews Microbiology · 2011 · 2,472 citations
UniProt: a hub for protein information
Nucleic Acids Research · 2014 · 5,247 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.