Scinovex
articleTop 1% cited

Training Products of Experts by Minimizing Contrastive Divergence

Neural Computation · 2002 · Vol. 14(8) · pp. 1771–1800
Geoffrey E. Hinton

Abstract

It is possible to combine multiple latent-variable models of the same data by multiplying their probability distributions together and then renormalizing. This way of combining individual "expert" models makes it hard to generate samples from the combined model but easy to infer the values of the latent variables of each expert, because the combination rule ensures that the latent variables of different experts are conditionally independent when given the data. A product of experts (PoE) is therefore an interesting candidate for a perceptual system in which rapid inference is vital and generation is unnecessary. Training a PoE by maximizing the likelihood of the data is difficult because it is hard even to approximate the derivatives of the renormalization term in the combination rule. Fortunately, a PoE can be trained using a different objective function called "contrastive divergence" whose derivatives with regard to the parameters can be approximated accurately and efficiently. Examples are presented of contrastive divergence learning using several types of expert on several types of data.

Neural Networks and ApplicationsRemote-Sensing Image ClassificationBayesian Modeling and Causal InferenceLatent variableDivergence (linguistics)InferenceComputer scienceArtificial intelligenceMachine learningConditional independenceLatent variable modelVariable (mathematics)Mathematics

Funding

  • Gatsby Charitable Foundation
Citations
4,959
FWCI
31.82
field-weighted impact
References
24
Percentile
100%
vs. same field & year
Citations per year
Cited by
Fields of Experts
International Journal of Computer Vision · 2009 · 860 citations
Neural Networks and Deep Learning
Machine Learning · 2015 · 920 citations
A Connection Between Score Matching and Denoising Autoencoders
Neural Computation · 2011 · 933 citations
Big Data Deep Learning: Challenges and Perspectives
IEEE Access · 2014 · 1,248 citations
10.1162/153244303322533223
Applied Physics Letters · 2000 · 1,778 citations
Deep Architecture for Traffic Flow Prediction: Deep Belief Networks With Multitask Learning
IEEE Transactions on Intelligent Transportation Systems · 2014 · 1,132 citations
References
The psychology of computer vision
Pattern Recognition · 1976 · 2,574 citations
Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images
IEEE Transactions on Pattern Analysis and Machine Intelligence · 1984 · 17,882 citations
Related articles
Training Products of Experts by Minimizing Contrastive Divergence
Neural Computation · 2002 · 4,959 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.