Scinovex
article Open AccessTop 1% cited

Repetitive Elements May Comprise Over Two-Thirds of the Human Genome

PLoS Genetics · 2011 · Vol. 7(12) · pp. e1002384–e1002384
A. P. Jason de KoningWanjun GuTodd A. CastoeMark A. BatzerDavid D. Pollock

Abstract

Transposable elements (TEs) are conventionally identified in eukaryotic genomes by alignment to consensus element sequences. Using this approach, about half of the human genome has been previously identified as TEs and low-complexity repeats. We recently developed a highly sensitive alternative de novo strategy, P-clouds, that instead searches for clusters of high-abundance oligonucleotides that are related in sequence space (oligo "clouds"). We show here that P-clouds predicts >840 Mbp of additional repetitive sequences in the human genome, thus suggesting that 66%-69% of the human genome is repetitive or repeat-derived. To investigate this remarkable difference, we conducted detailed analyses of the ability of both P-clouds and a commonly used conventional approach, RepeatMasker (RM), to detect different sized fragments of the highly abundant human Alu and MIR SINEs. RM can have surprisingly low sensitivity for even moderately long fragments, in contrast to P-clouds, which has good sensitivity down to small fragment sizes (∼25 bp). Although short fragments have a high intrinsic probability of being false positives, we performed a probabilistic annotation that reflects this fact. We further developed "element-specific" P-clouds (ESPs) to identify novel Alu and MIR SINE elements, and using it we identified ∼100 Mb of previously unannotated human elements. ESP estimates of new MIR sequences are in good agreement with RM-based predictions of the amount that RM missed. These results highlight the need for combined, probabilistic genome annotation approaches and suggest that the human genome consists of substantially more repetitive sequence than previously believed.

Chromosomal and Genetic VariationsGenomics and Phylogenetic StudiesRNA and protein synthesis mechanismsGenomeHuman genomeBiologyAlu elementFalse positive paradoxComputational biologyTransposable elementGeneticsRepeated sequenceInterspersed repeat

MeSH terms

AlgorithmsDNA Transposable ElementsHumansRepetitive Sequences, Nucleic AcidSoftwareGenome, HumanConsensus SequenceComputational BiologyLong Interspersed Nucleotide ElementsAlu ElementsMolecular Sequence Annotation

Funding

  • National Science Foundation
  • Louisiana Board of Regents
  • National Institutes of Health
Citations
1,132
FWCI
67.59
field-weighted impact
References
41
Percentile
100%
vs. same field & year
Citations per year
References
Tandem repeats finder: a program to analyze DNA sequences
Nucleic Acids Research · 1999 · 9,719 citations
Non-coding RNA
Human Molecular Genetics · 2006 · 2,447 citations
Initial sequencing and analysis of the human genome
Nature · 2001 · 24,452 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.