Scinovex
article Open AccessTop 1% cited

Identifying a High Fraction of the Human Genome to be under Selective Constraint Using GERP++

PLoS Computational Biology · 2010 · Vol. 6(12) · pp. e1001025–e1001025
Eugene DavydovDavid L. GoodeMarina SirotaGregory M. CooperArend SidowSerafim Batzoglou

Abstract

Computational efforts to identify functional elements within genomes leverage comparative sequence information by looking for regions that exhibit evidence of selective constraint. One way of detecting constrained elements is to follow a bottom-up approach by computing constraint scores for individual positions of a multiple alignment and then defining constrained elements as segments of contiguous, highly scoring nucleotide positions. Here we present GERP++, a new tool that uses maximum likelihood evolutionary rate estimation for position-specific scoring and, in contrast to previous bottom-up methods, a novel dynamic programming approach to subsequently define constrained elements. GERP++ evaluates a richer set of candidate element breakpoints and ranks them based on statistical significance, eliminating the need for biased heuristic extension techniques. Using GERP++ we identify over 1.3 million constrained elements spanning over 7% of the human genome. We predict a higher fraction than earlier estimates largely due to the annotation of longer constrained elements, which improves one to one correspondence between predicted elements with known functional sequences. GERP++ is an efficient and effective tool to provide both nucleotide- and element-level constraint scores within deep multiple sequence alignments.

Genomics and Phylogenetic StudiesRNA and protein synthesis mechanismsGenomics and Chromatin DynamicsConstraint (computer-aided design)GenomeComputer scienceMultiple sequence alignmentLeverage (statistics)HeuristicHuman genomeSet (abstract data type)Fraction (chemistry)Sequence (biology)

MeSH terms

AlgorithmsAnimalsHumansMammalsModels, GeneticPhylogenySoftwareUser-Computer InterfaceGenome, HumanSequence AlignmentSequence Analysis, DNAGenomics

Funding

  • National Science Foundation
  • National Institutes of Health
  • National Human Genome Research Institute
  • U.S. National Library of Medicine
Citations
1,854
FWCI
14.61
field-weighted impact
References
30
Percentile
99%
vs. same field & year
Citations per year
References
Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1977 · 49,286 citations
Basic local alignment search tool
Journal of Molecular Biology · 1990 · 93,570 citations
Algorithms for Minimization Without Derivatives
Mathematics of Computation · 1974 · 2,934 citations
The Human Genome Browser at UCSC
Genome Research · 2002 · 10,897 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.