Scinovex
articleTop 1% cited

Pfam: A comprehensive database of protein domain families based on seed alignments

Proteins Structure Function and Bioinformatics · 1997 · Vol. 28(3) · pp. 405–420
Erik L. L. SonnhammerSean R. EddyRichard Durbin

Abstract

Databases of multiple sequence alignments are a valuable aid to protein sequence classification and analysis. One of the main challenges when constructing such a database is to simultaneously satisfy the conflicting demands of completeness on the one hand and quality of alignment and domain definitions on the other. The latter properties are best dealt with by manual approaches, whereas completeness in practice is only amenable to automatic methods. Herein we present a database based on hidden Markov model profiles (HMMs), which combines high quality and completeness. Our database, Pfam, consists of parts A and B. Pfam-A is curated and contains well-characterized protein domain families with high quality alignments, which are maintained by using manually checked seed alignments and HMMs to find and align all members. Pfam-B contains sequence families that were generated automatically by applying the Domainer algorithm to cluster and align the remaining protein sequences after removal of Pfam-A domains. By using Pfam, a large number of previously unannotated proteins from the Caenorhabditis elegans genome project were classified. We have also identified many novel family memberships in known proteins, including new kazal, Fibronectin type III, and response regulator receiver domains. Pfam-A families have permanent accession numbers and form a library of HMMs available for searching and automatic annotation of new protein sequences.

Genomics and Phylogenetic StudiesAdvanced Proteomics Techniques and ApplicationsMachine Learning in BioinformaticsAnnotationProtein familyHidden Markov modelSequence alignmentSequence databaseProtein domainComputational biologyProtein sequencingSequence (biology)Computer science

MeSH terms

Amino Acid SequenceMultigene FamilyModels, ChemicalMolecular Sequence DataPlant ProteinsSeedsDatabases, FactualSequence AlignmentSequence Homology, Amino AcidProtein Structure, Tertiary

Funding

  • Wellcome Trust
  • National Institutes of Health
  • Medical Research Council
Citations
1,247
FWCI
14.20
field-weighted impact
References
56
Percentile
99%
vs. same field & year
Citations per year
Cited by
Pfam: The protein families database in 2021
Nucleic Acids Research · 2020 · 7,537 citations
The Pfam Protein Families Database
Nucleic Acids Research · 2002 · 14,220 citations
Machine learning applications in genetics and genomics
Nature Reviews Genetics · 2015 · 1,956 citations
Profile hidden Markov models.
Bioinformatics · 1998 · 5,777 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.