Scinovex
article Open Access

Automated De Novo Identification of Repeat Sequence Families in Sequenced Genomes

Genome Research · 2002 · Vol. 12(8) · pp. 1269–1276
Zhirong BaoSean R. Eddy

Abstract

Repetitive sequences make up a major part of eukaryotic genomes. We have developed an approach for the de novo identification and classification of repeat sequence families that is based on extensions to the usual approach of single linkage clustering of local pairwise alignments between genomic sequences. Our extensions use multiple alignment information to define the boundaries of individual copies of the repeats and to distinguish homologous but distinct repeat element families. When tested on the human genome, our approach was able to properly identify and group known transposable elements. The program, should be useful for first-pass automatic classification of repeats in newly sequenced genomes.

Genomics and Phylogenetic StudiesChromosomal and Genetic VariationsRNA and protein synthesis mechanismsBiologyGenomeTransposable elementGeneticsIdentification (biology)Computational biologySequence (biology)Human genomeLinkage (software)Pairwise comparison

MeSH terms

AlgorithmsMultigene FamilyHumansGenetic LinkageRepetitive Sequences, Nucleic AcidGenome, HumanCluster AnalysisSequence AlignmentComputational Biology

Funding

  • National Science Foundation
  • Howard Hughes Medical Institute
Citations
1,029
FWCI
field-weighted impact
References
19
Percentile
vs. same field & year
Citations per year
References
The Pfam Protein Families Database
Nucleic Acids Research · 2002 · 14,220 citations
Initial sequencing and analysis of the human genome
Nature · 2001 · 24,452 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.