Scinovex
article Open Access

Gene prediction in novel fungal genomes using an ab initio algorithm with unsupervised training

Genome Research · 2008 · Vol. 18(12) · pp. 1979–1990
Vardges Ter-HovhannisyanAlexandre LomsadzeYury O. ChernoffMark Borodovsky

Abstract

We describe a new ab initio algorithm, GeneMark-ES version 2, that identifies protein-coding genes in fungal genomes. The algorithm does not require a predetermined training set to estimate parameters of the underlying hidden Markov model (HMM). Instead, the anonymous genomic sequence in question is used as an input for iterative unsupervised training. The algorithm extends our previously developed method tested on genomes of Arabidopsis thaliana, Caenorhabditis elegans, and Drosophila melanogaster. To better reflect features of fungal gene organization, we enhanced the intron submodel to accommodate sequences with and without branch point sites. This design enables the algorithm to work equally well for species with the kinds of variations in splicing mechanisms seen in the fungal phyla Ascomycota, Basidiomycota, and Zygomycota. Upon self-training, the intron submodel switches on in several steps to reach its full complexity. We demonstrate that the algorithm accuracy, both at the exon and the whole gene level, is favorably compared to the accuracy of gene finders that employ supervised training. Application of the new method to known fungal genomes indicates substantial improvement over existing annotations. By eliminating the effort necessary to build comprehensive training sets, the new algorithm can streamline and accelerate the process of annotation in a large number of fungal genome sequencing projects.

Genomics and Phylogenetic StudiesGenetic diversity and population structurePlant-Microbe Interactions and ImmunityBiologyGenomeGene predictionComputational biologyGeneGenome projectCaenorhabditis elegansHidden Markov modelGene AnnotationIntron

MeSH terms

AlgorithmsGenes, FungalIntronsPredictive Value of TestsTeachingGenome, FungalSequence Analysis, DNAExpressed Sequence Tags

Funding

  • National Institutes of Health
Citations
1,105
FWCI
field-weighted impact
References
45
Percentile
vs. same field & year
Citations per year
References
Gene finding in novel genomes
BMC Bioinformatics · 2004 · 3,459 citations
Identifying bacterial genes and endosymbiont DNA with Glimmer
Bioinformatics · 2007 · 3,177 citations
GeneWise and Genomewise
Genome Research · 2004 · 3,009 citations
<tt>BLAT</tt>—The <tt>BLAST</tt>-Like Alignment Tool
Genome Research · 2002 · 8,404 citations
Prediction of complete gene structures in human genomic DNA
Journal of Molecular Biology · 1997 · 4,306 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.