Scinovex
article Open AccessTop 1% cited

Deep Supervised, but Not Unsupervised, Models May Explain IT Cortical Representation

PLoS Computational Biology · 2014 · Vol. 10(11) · pp. e1003915–e1003915
Seyed‐Mahdi Khaligh‐RazaviNikolaus Kriegeskorte

Abstract

Inferior temporal (IT) cortex in human and nonhuman primates serves visual object recognition. Computational object-vision models, although continually improving, do not yet reach human performance. It is unclear to what extent the internal representations of computational models can explain the IT representation. Here we investigate a wide range of computational model representations (37 in total), testing their categorization performance and their ability to account for the IT representational geometry. The models include well-known neuroscientific object-recognition models (e.g. HMAX, VisNet) along with several models from computer vision (e.g. SIFT, GIST, self-similarity features, and a deep convolutional neural network). We compared the representational dissimilarity matrices (RDMs) of the model representations with the RDMs obtained from human IT (measured with fMRI) and monkey IT (measured with cell recording) for the same set of stimuli (not used in training the models). Better performing models were more similar to IT in that they showed greater clustering of representational patterns by category. In addition, better performing models also more strongly resembled IT in terms of their within-category representational dissimilarities. Representational geometries were significantly correlated between IT and many of the models. However, the categorical clustering observed in IT was largely unexplained by the unsupervised models. The deep convolutional network, which was trained by supervision with over a million category-labeled images, reached the highest categorization performance and also best explained IT, although it did not fully explain the IT data. Combining the features of this model with appropriate weights and adding linear combinations that maximize the margin between animate and inanimate objects and between faces and other objects yielded a representation that fully explained our IT data. Overall, our results suggest that explaining IT requires computational features trained through supervised learning to emphasize the behaviorally important categorical divisions prominently reflected in IT.

Face Recognition and PerceptionVisual Attention and Saliency DetectionNeural dynamics and brain functionCategorizationArtificial intelligenceComputer scienceComputational modelRepresentation (politics)Convolutional neural networkPattern recognition (psychology)Set (abstract data type)Cognitive neuroscience of visual object recognitionCluster analysis

MeSH terms

AnimalsHaplorhiniHumansModels, NeurologicalTemporal LobeComputational BiologySupport Vector Machine

Funding

  • Cambridge Overseas Trust
  • Medical Research Council
Citations
1,349
FWCI
33.96
field-weighted impact
References
98
Percentile
100%
vs. same field & year
Citations per year
References
Stimulus-selective properties of inferior temporal neurons in the macaque
Journal of Neuroscience · 1984 · 1,447 citations
Modeling the Shape of the Scene: A Holistic Representation of the Spatial Envelope
International Journal of Computer Vision · 2001 · 6,378 citations
Shape matching and object recognition using shape contexts
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2002 · 6,295 citations
A Toolbox for Representational Similarity Analysis
PLoS Computational Biology · 2014 · 1,051 citations
Learning Invariance from Transformation Sequences
Neural Computation · 1991 · 680 citations
Receptive fields and functional architecture of monkey striate cortex
The Journal of Physiology · 1968 · 6,597 citations
A Fast Learning Algorithm for Deep Belief Nets
Neural Computation · 2006 · 16,253 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.