Scinovex
articleTop 1% cited

The Sequence of the Human Genome

Science · 2001 · Vol. 291(5507) · pp. 1304–1351
J. Craig VenterMark D. AdamsEugene W. MyersPeter W. LiRichard MuralGranger G. SuttonHamilton O. SmithMark YandellCheryl EvansRobert A. HoltJeannine D. GocaynePeter G. AmanatidesRichard M. BallewDaniel H. HusonJennifer R. WortmanQing ZhangChinnappa D. KodiraXiangqun Zheng-BradleyLin ChenMarian SkupskiG. SubramanianPaul D. ThomasJinghui ZhangGeorge L. Gabor MiklosCatherine R. NelsonSamuel BroderAndrew G. ClarkJoe NadeauVictor A. McKusickNorton D. ZinderArnold J. LevineRichard J. RobertsMel I. SimonCarolyn W. SlaymanMichael W. HunkapillerRandall BolanosArthur L. DelcherIan DewDaniel FasuloMichael J. FlaniganLiliana FloreaAaron L. HalpernSridhar HannenhalliSaul KravitzSamuel LévyClark MobarryKnut ReinertKarin RemingtonJane Abu-ThreidehEllen M. BeasleyKendra BiddickVivien BonazziRhonda BrandonMichele CargillIshwar ChandramouliswaranRosane CharlabKabir ChaturvediZuoming DengValentina Di FrancescoPatrick DunnKaren EilbeckCarlos EvangelistaAndrei GabrielianWeiniu GanWangmao GeFangcheng GongZhiping GuPing GuanThomas J. HeimanMaureen E. HigginsRui‐Ru JiZhaoxi KeKaren A. KetchumZhongwu LaiYiding LeiZhenya LiJiayin LiYong LiangXiaoying LinFu LuGennady V. MerkulovNatalia V. MilshinaHelen M. MooreAshwinikumar K. NaikVaibhav A. NarayanBeena NeelamDeborah NusskernDouglas B. RuschSteven L. SalzbergWei ShaoBixiong Chris ShueJing‐Tao SunZhen Yuan WangAihui WangXin WangJian WangMinghui WeiRon WidesChunlin XiaoChunhua YanAlison YaoJane J. YeMing ZhanWeiqing ZhangHongyu ZhangQi ZhaoLiansheng ZhengFei ZhongWenyan ZhongShiaoping C. ZhuShaying ZhaoDennis A. GilbertSuzanna BaumhueterGene SpierChristine CarterAnibal CravchikTrevor WoodageFeroze AliHui-Jin AnAderonke AweDanita BaldwinHolly BadenMary BarnsteadIan BarrowKaren BeesonDana BusamAmy CarverMing Lai ChengLiz CurrySteve DanaherLionel B. DavenportRaymond DesiletsSusanne DietzKristina DodsonLisa DoupSteven FerrieraNeha GargAndres GluecksmannBrit J. HartJason HaynesCharles A. HaynesCheryl HeinerSuzanne L. HladunDamon HostinJarrett HouckTimothy J. HowlandChinyere IbegwamJeffery E. JohnsonFrancis KalushLesley KlineShashi KoduruAmy LoveF.H. MannDavid MaySteven McCawleyTina C. McIntoshIvy McMullenMee MoyLinda MoyBrian J. MurphyK. E. NelsonCynthia PfannkochEric C. PrattsVinita PuriHina QureshiMatthew S. ReardonRobert RodriguezYu-Hui RogersDeanna L. RombladBob RuhfelRichard ScottCynthia D. SitterMichelle SmallwoodErin StewartRenee StrongEllen SuhReginald ThomasNi Ni TintSukyee TseClaire VechGary WangJeremy WetterS. WilliamsMonica WilliamsSandra M. WindsorEmily S. Winn-DeenKeriellen WolfeJayshree ZaveriK. ZaveriJosep F. AbrilRoderic GuigóMichael J. CampbellKimmen V. SjölanderBrian KarlakAnish KejariwalHuaiyu MiBetty V. LazarevaThomas W. HattonApurva NarechaniaKaren DiemerAnushya MuruganujanNan GuoShinji SatoVineet BafnaSorin IstrailRoss A. LippertRussell SchwartzBrian P. WalenzShibu YoosephDavid R. AllenAnand BasuJames BaxendaleLouis BlickMarcelo CaminhaJohn Carnes-StineParris M. CaulkYen-Hui ChiangMy D. CoyneCarl DahlkeAnne Deslattes MaysMaria DombroskiMichael DonnellyDale ElyShiva EsparhamCarl FoslerHarold C. GireStephen GlanowskiKenneth GlasserAnna GlodekMark GorokhovKen GrahamBarry GropmanMichael A. HarrisJeremy HeilScott N. HendersonJeffrey P. HooverDonald E. JenningsCatherine JordanJames M. JordanJohn KashaLeonid KaganCheryl KraftAlexander A. LevitskyMark G. LewisXiangjun LiuJohn LopezJ. DanielWilliam H. MajorosJoe W. McDanielSean D. MurphyMatthew NewmanTrung NguyenNgoc B. NguyenMarc NodellSue PanJim PeckMarshall PetersonWilliam RoweRobert D. SandersJohn ScottMichael A. SimpsonThomas J. SmithArlan C. SpragueTimothy B. StockwellRussell TurnerEli VenterMei WangMeiyuan WenDavid WuMitchell M. WuAshley C. XiaAli ZandiehZhu Xiao-hong

Abstract

A 2.91-billion base pair (bp) consensus sequence of the euchromatic portion of the human genome was generated by the whole-genome shotgun sequencing method. The 14.8-billion bp DNA sequence was generated over 9 months from 27,271,853 high-quality sequence reads (5.11-fold coverage of the genome) from both ends of plasmid clones made from the DNA of five individuals. Two assembly strategies-a whole-genome assembly and a regional chromosome assembly-were used, each combining sequence data from Celera and the publicly funded genome effort. The public data were shredded into 550-bp segments to create a 2.9-fold coverage of those genome regions that had been sequenced, without including biases inherent in the cloning and assembly procedure used by the publicly funded group. This brought the effective coverage in the assemblies to eightfold, reducing the number and size of gaps in the final assembly over what would be obtained with 5.11-fold coverage. The two assembly strategies yielded very similar results that largely agree with independent mapping data. The assemblies effectively cover the euchromatic regions of the human chromosomes. More than 90% of the genome is in scaffold assemblies of 100,000 bp or more, and 25% of the genome is in scaffolds of 10 million bp or larger. Analysis of the genome sequence revealed 26,588 protein-encoding transcripts for which there was strong corroborating evidence and an additional approximately 12,000 computationally derived genes with mouse matches or other weak supporting evidence. Although gene-dense clusters are obvious, almost half the genes are dispersed in low G+C sequence separated by large tracts of apparently noncoding sequence. Only 1.1% of the genome is spanned by exons, whereas 24% is in introns, with 75% of the genome being intergenic DNA. Duplications of segmental blocks, ranging in size up to chromosomal lengths, are abundant throughout the genome and reveal a complex evolutionary history. Comparative genomic analysis indicates vertebrate expansions of genes associated with neuronal function, with tissue-specific developmental regulation, and with the hemostasis and immune systems. DNA sequence comparisons between the consensus sequence and publicly funded genome data provided locations of 2.1 million single-nucleotide polymorphisms (SNPs). A random pair of human haploid genomes differed at a rate of 1 bp per 1250 on average, but there was marked heterogeneity in the level of polymorphism across the genome. Less than 1% of all SNPs resulted in variation in proteins, but the task of determining which SNPs have functional consequences remains an open challenge.

Genomics and Phylogenetic StudiesRNA and protein synthesis mechanismsGenomics and Chromatin DynamicsGenomeSequence assemblyBiologyGeneticsHuman genomeReference genomeHybrid genome assemblyGenome projectShotgun sequencingComputational biology

MeSH terms

AlgorithmsAnimalsChromosome BandingChromosome MappingExonsFemaleGenesHumansIntronsMalePhenotypeProteinsPseudogenesRepetitive Sequences, Nucleic AcidSpecies Specificity
Citations
13,619
FWCI
433.96
field-weighted impact
References
189
Percentile
100%
vs. same field & year
Citations per year
Cited by
Phylogenetic study of Lac Insects of Kerria spp. using intron length polymorphism (EPIC-PCR)
Journal of Entomology and Zoology Studies · 2014 · 1 citations
Mechanism and role of PDZ domains in signaling complex assembly
Journal of Cell Science · 2001 · 835 citations
A review of clustering techniques and developments
Neurocomputing · 2017 · 1,308 citations
Molecular Physiology of P2X Receptors
Physiological Reviews · 2002 · 2,895 citations
Counting the Zinc-Proteins Encoded in the Human Genome
Journal of Proteome Research · 2005 · 1,095 citations
Smad regulation in TGF-β signal transduction
Journal of Cell Science · 2001 · 924 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.