Scinovex
articleTop 1% cited

Selecting the Best-Fit Model of Nucleotide Substitution

Systematic Biology · 2001 · Vol. 50(4) · pp. 580–601
David PosadaKeith A. Crandall

Abstract

Despite the relevant role of models of nucleotide substitution in phylogenetics, choosing among different models remains a problem. Several statistical methods for selecting the model that best fits the data at hand have been proposed, but their absolute and relative performance has not yet been characterized. In this study, we compare under various conditions the performance of different hierarchical and dynamic likelihood ratio tests, and of Akaike and Bayesian information methods, for selecting best-fit models of nucleotide substitution. We specifically examine the role of the topology used to estimate the likelihood of the different models and the importance of the order in which hypotheses are tested. We do this by simulating DNA sequences under a known model of nucleotide substitution and recording how often this true model is recovered by the different methods. Our results suggest that model selection is reasonably accurate and indicate that some likelihood ratio test methods perform overall better than the Akaike or Bayesian information criteria. The tree used to estimate the likelihood scores does not influence model selection unless it is a randomly chosen tree. The order in which hypotheses are tested, and the complexity of the initial model in the sequence of tests, influence model selection in some cases. Model fitting in phylogenetics has been suggested for many years, yet many authors still arbitrarily choose their models, often using the default models implemented in standard computer programs for phylogenetic estimation. We show here that a best-fit model can be readily identified. Consequently, given the relevance of models, model fitting should be routine in any phylogenetic analysis that uses models of evolution.

Genomics and Phylogenetic StudiesEvolution and Paleontology StudiesGenetic diversity and population structureAkaike information criterionModel selectionBayesian information criterionSubstitution (logic)Likelihood-ratio testSelection (genetic algorithm)Bayesian probabilityTree (set theory)Statistical modelInformation Criteria

MeSH terms

AlgorithmsBase SequenceBiometryDNAModels, GeneticPhylogenyLikelihood FunctionsEvolution, Molecular

Funding

  • Hong Kong Institute of Education
Citations
856
FWCI
12.96
field-weighted impact
References
57
Percentile
99%
vs. same field & year
Citations per year
Related articles
Selecting the Best-Fit Model of Nucleotide Substitution
Systematic Biology · 2001 · 856 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.