Scinovex
article Open AccessTop 1% cited

Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation

Nucleic Acids Research · 2015 · Vol. 44(D1) · pp. D733–D745
Nuala A. O’LearyMatt W. WrightJ. Rodney BristerStacy CiufoDiana HaddadRich McVeighBhanu RajputBarbara RobbertseBrian Smith-WhiteDanso Ako-adjeiAlexander AstashynAzat BadretdinYīmíng BàoOlga BlinkovaVyacheslav BroverVyacheslav ChetverninJinna ChoiEric CoxOlga ErmolaevaCatherine M. FarrellTamara GoldfarbTripti GuptaDaniel H. HaftEneida HatcherWratko HlavinaVinita JoardarVamsi K. KodaliWenjun LiDonna MaglottPatrick MastersonKelly M. McGarveyMichael R. MurphyKathleen O’NeillShashikant PujarSanjida H RangwalaDaniel RauschLillian D. RiddickConrad L. SchochAndrei ShkedaSusan S. StorzHanzhen SunFrançoise Thibaud‐NissenIgor TolstoyRaymond E. TullyAnjana R. VatsanCraig WallinDavid WebbWendy WuMelissa LandrumAvi KimchiTatiana TatusovaMichael DiCuccioPaul KittsTerence D. MurphyKim D. Pruitt

Abstract

The RefSeq project at the National Center for Biotechnology Information (NCBI) maintains and curates a publicly available database of annotated genomic, transcript, and protein sequence records (http://www.ncbi.nlm.nih.gov/refseq/). The RefSeq project leverages the data submitted to the International Nucleotide Sequence Database Collaboration (INSDC) against a combination of computation, manual curation, and collaboration to produce a standard set of stable, non-redundant reference sequences. The RefSeq project augments these reference sequences with current knowledge including publications, functional features and informative nomenclature. The database currently represents sequences from more than 55,000 organisms (>4800 viruses, >40,000 prokaryotes and >10,000 eukaryotes; RefSeq release 71), ranging from a single record to complete genomes. This paper summarizes the current status of the viral, prokaryotic, and eukaryotic branches of the RefSeq project, reports on improvements to data access and details efforts to further expand the taxonomic representation of the collection. We also highlight diverse functional curation initiatives that support multiple uses of RefSeq data including taxonomic validation, genome annotation, comparative genomics, and clinical testing. We summarize our approach to utilizing available RNA-Seq and other data types in our manual curation process for vertebrate, plant, and other species, and describe a new direction for prokaryotic genomes and protein name management.

Genomics and Phylogenetic StudiesBacteriophages and microbial interactionsPlant Virus Research StudiesRefSeqAnnotationBiologyEnsemblComparative genomicsGenomeGenome projectGeneticistGenomicsComputational biology

MeSH terms

AnimalsCattleHumansInvertebratesNematodaPhylogenyReference StandardsVertebratesGenome, HumanGenome, ViralGenome, FungalSequence Analysis, RNAGenome, PlantSequence Analysis, ProteinGene Expression Profiling
Citations
6,951
FWCI
113.13
field-weighted impact
References
67
Percentile
100%
vs. same field & year
Citations per year
Cited by
KEGG: new perspectives on genomes, pathways, diseases and drugs
Nucleic Acids Research · 2016 · 9,241 citations
RESCRIPt: Reproducible sequence taxonomy reference database management
PLoS Computational Biology · 2021 · 895 citations
Best practices for analysing microbiomes
Nature Reviews Microbiology · 2018 · 1,823 citations
KEGG for taxonomy-based analysis of pathways and genomes
Nucleic Acids Research · 2022 · 6,257 citations
References
miRBase: annotating high confidence microRNAs using deep sequencing data
Nucleic Acids Research · 2013 · 5,177 citations
Nuclear ribosomal internal transcribed spacer (ITS) region as a universal DNA barcode marker for <i>Fungi</i>
Proceedings of the National Academy of Sciences · 2012 · 4,986 citations
Initial sequencing and analysis of the human genome
Nature · 2001 · 24,452 citations
UniProt: a hub for protein information
Nucleic Acids Research · 2014 · 5,247 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.