Scinovex
article Open AccessTop 1% cited

Statistical Methods for Detecting Differentially Abundant Features in Clinical Metagenomic Samples

PLoS Computational Biology · 2009 · Vol. 5(4) · pp. e1000352–e1000352
James R. WhiteNiranjan NagarajanMihai Pop

Abstract

Numerous studies are currently underway to characterize the microbial communities inhabiting our world. These studies aim to dramatically expand our understanding of the microbial biosphere and, more importantly, hope to reveal the secrets of the complex symbiotic relationship between us and our commensal bacterial microflora. An important prerequisite for such discoveries are computational tools that are able to rapidly and accurately compare large datasets generated from complex bacterial communities to identify features that distinguish them.We present a statistical method for comparing clinical metagenomic samples from two treatment populations on the basis of count data (e.g. as obtained through sequencing) to detect differentially abundant features. Our method, Metastats, employs the false discovery rate to improve specificity in high-complexity environments, and separately handles sparsely-sampled features using Fisher's exact test. Under a variety of simulations, we show that Metastats performs well compared to previously used methods, and significantly outperforms other methods for features with sparse counts. We demonstrate the utility of our method on several datasets including a 16S rRNA survey of obese and lean human gut microbiomes, COG functional profiles of infant and mature gut microbiomes, and bacterial and viral metabolic subsystem data inferred from random sequencing of 85 metagenomes. The application of our method to the obesity dataset reveals differences between obese and lean subjects not reported in the original study. For the COG and subsystem datasets, we provide the first statistically rigorous assessment of the differences between these populations. The methods described in this paper are the first to address clinical metagenomic datasets comprising samples from multiple subjects. Our methods are robust across datasets of varied complexity and sampling level. While designed for metagenomic applications, our software can also be applied to digital gene expression studies (e.g. SAGE). A web server implementation of our methods and freely available source code can be found at http://metastats.cbcb.umd.edu/.

Gut microbiota and healthGene expression and cancer classificationMetabolomics and Mass Spectrometry StudiesMetagenomicsMicrobiomeComputational biologyBiologyCogFalse discovery rateComputer scienceArtificial intelligenceBioinformaticsGenetics

MeSH terms

BacteriaChromosome MappingDNA, BacterialHumansIntestinesObesityGene Expression Profiling

Funding

  • Bill and Melinda Gates Foundation
Citations
1,639
FWCI
15.06
field-weighted impact
References
44
Percentile
99%
vs. same field & year
Citations per year
Cited by
Linkage of gut microbiome with cognition in hepatic encephalopathy
American Journal of Physiology-Gastrointestinal and Liver Physiology · 2011 · 546 citations
Colonic microbiome is altered in alcoholism
American Journal of Physiology-Gastrointestinal and Liver Physiology · 2012 · 749 citations
A Primer on Metagenomics
PLoS Computational Biology · 2010 · 695 citations
Waste Not, Want Not: Why Rarefying Microbiome Data Is Inadmissible
PLoS Computational Biology · 2014 · 3,022 citations
References
Statistical significance for genomewide studies
Proceedings of the National Academy of Sciences · 2003 · 10,009 citations
Greengenes, a Chimera-Checked 16S rRNA Gene Database and Workbench Compatible with ARB
Applied and Environmental Microbiology · 2006 · 11,173 citations
Development of the Human Infant Intestinal Microbiota
PLoS Biology · 2007 · 2,829 citations
Human gut microbes associated with obesity
Nature · 2006 · 8,677 citations
Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1995 · 106,483 citations
MEGAN analysis of metagenomic data
Genome Research · 2007 · 3,436 citations
Naive Bayesian Classifier for Rapid Assignment of rRNA Sequences into the New Bacterial Taxonomy
Applied and Environmental Microbiology · 2007 · 20,269 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.