Scinovex
articleTop 10% cited

Separating Style and Content with Bilinear Models

Neural Computation · 2000 · Vol. 12(6) · pp. 1247–1283
Joshua B. TenenbaumWilliam T. Freeman

Abstract

Perceptual systems routinely separate "content" from "style," classifying familiar words spoken in an unfamiliar accent, identifying a font or handwriting style across letters, or recognizing a familiar face or object seen under unfamiliar viewing conditions. Yet a general and tractable computational model of this ability to untangle the underlying factors of perceptual observations remains elusive (Hofstadter, 1985). Existing factor models (Mardia, Kent, & Bibby, 1979; Hinton & Zemel, 1994; Ghahramani, 1995; Bell & Sejnowski, 1995; Hinton, Dayan, Frey, & Neal, 1995; Dayan, Hinton, Neal, & Zemel, 1995; Hinton & Ghahramani, 1997) are either insufficiently rich to capture the complex interactions of perceptually meaningful factors such as phoneme and speaker accent or letter and font, or do not allow efficient learning algorithms. We present a general framework for learning to solve two-factor tasks using bilinear models, which provide sufficiently expressive representations of factor interactions but can nonetheless be fit to data using efficient algorithms based on the singular value decomposition and expectation-maximization. We report promising results on three different tasks in three different perceptual domains: spoken vowel classification with a benchmark multi-speaker database, extrapolation of fonts to unseen letters, and translation of faces to novel illuminants.

Music and Audio ProcessingSpeech and Audio ProcessingBlind Source Separation TechniquesStress (linguistics)Speech recognitionComputer scienceFactor (programming language)Artificial intelligenceBilinear interpolationStyle (visual arts)PerceptionNatural language processingPsychology

MeSH terms

AlgorithmsFaceHumansLearningPattern Recognition, VisualSpeech PerceptionLinear ModelsNeural Networks, Computer
Citations
874
FWCI
10.36
field-weighted impact
References
61
Percentile
99%
vs. same field & year
Citations per year
Cited by
Face recognition
ACM Computing Surveys · 2003 · 6,131 citations
References
Neural networks for pattern recognition
Choice Reviews Online · 1994 · 18,690 citations
The Helmholtz Machine
Neural Computation · 1995 · 1,207 citations
Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1977 · 49,286 citations
Eigenfaces for Recognition
Journal of Cognitive Neuroscience · 1991 · 13,711 citations
Shape and motion from image streams under orthography: a factorization method
International Journal of Computer Vision · 1992 · 2,865 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.