Scinovex
article Open AccessTop 1% cited

RESCRIPt: Reproducible sequence taxonomy reference database management

PLoS Computational Biology · 2021 · Vol. 17(11) · pp. e1009581–e1009581
Michael S. RobesonDevon O’RourkeBenjamin D. KaehlerMichał ZiemskiMatthew R. DillonJeffrey T. FosterNicholas A. Bokulich

Abstract

Nucleotide sequence and taxonomy reference databases are critical resources for widespread applications including marker-gene and metagenome sequencing for microbiome analysis, diet metabarcoding, and environmental DNA (eDNA) surveys. Reproducibly generating, managing, using, and evaluating nucleotide sequence and taxonomy reference databases creates a significant bottleneck for researchers aiming to generate custom sequence databases. Furthermore, database composition drastically influences results, and lack of standardization limits cross-study comparisons. To address these challenges, we developed RESCRIPt, a Python 3 software package and QIIME 2 plugin for reproducible generation and management of reference sequence taxonomy databases, including dedicated functions that streamline creating databases from popular sources, and functions for evaluating, comparing, and interactively exploring qualitative and quantitative characteristics across reference databases. To highlight the breadth and capabilities of RESCRIPt, we provide several examples for working with popular databases for microbiome profiling (SILVA, Greengenes, NCBI-RefSeq, GTDB), eDNA and diet metabarcoding surveys (BOLD, GenBank), as well as for genome comparison. We show that bigger is not always better, and reference databases with standardized taxonomies and those that focus on type strains have quantitative advantages, though may not be appropriate for all use cases. Most databases appear to benefit from some curation (quality filtering), though sequence clustering appears detrimental to database quality. Finally, we demonstrate the breadth and extensibility of RESCRIPt for reproducible workflows with a comparison of global hepatitis genomes. RESCRIPt provides tools to democratize the process of reference database acquisition and management, enabling researchers to reproducibly and transparently create reference materials for diverse research applications. RESCRIPt is released under a permissive BSD-3 license at https://github.com/bokulich-lab/RESCRIPt.

Environmental DNA in Biodiversity StudiesGenomics and Phylogenetic StudiesMicrobial Community Ecology and PhysiologyGenBankDatabaseReference genomeComputer scienceWorkflowMetagenomicsInformation retrievalReference databaseRefSeqDNA sequencing

MeSH terms

AnimalsClassificationDatabase Management SystemsHumansPhylogenyRNA, Ribosomal, 16SSoftwareSequence AnalysisComputational BiologyGenomicsDatabases, GeneticDatabases, Nucleic AcidMetagenomeMetagenomicsDNA Barcoding, Taxonomic
Citations
895
FWCI
92.20
field-weighted impact
References
127
Percentile
100%
vs. same field & year
Citations per year
References
Greengenes, a Chimera-Checked 16S rRNA Gene Database and Workbench Compatible with ARB
Applied and Environmental Microbiology · 2006 · 11,173 citations
Global patterns of 16S rRNA diversity at a depth of millions of sequences per sample
Proceedings of the National Academy of Sciences · 2010 · 9,826 citations
Biological identifications through DNA barcodes
Proceedings of the Royal Society B Biological Sciences · 2003 · 13,103 citations
Big Data: Astronomical or Genomical?
PLoS Biology · 2015 · 1,384 citations
Nuclear ribosomal internal transcribed spacer (ITS) region as a universal DNA barcode marker for <i>Fungi</i>
Proceedings of the National Academy of Sciences · 2012 · 4,986 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.