Scinovex
articleTop 1% cited

Recognition-by-components: A theory of human image understanding.

Psychological Review · 1987 · Vol. 94(2) · pp. 115–147
Irving Biederman

Abstract

The perceptual recognition of objects is conceptualized to be a process in which the image of the input is segmented at regions of deep concavity into an arrangement of simple geometric components, such as blocks, cylinders, wedges, and cones. The fundamental assumption of the proposed theory, recognition-by-components (RBC), is that a modest set of generalized-cone components, called geons (N £ 36), can be derived from contrasts of five readily detectable properties of edges in a two-dimensiona l image: curvature, collinearity, symmetry, parallelism, and cotermination. The detection of these properties is generally invariant over viewing position an$ image quality and consequently allows robust object perception when the image is projected from a novel viewpoint or is degraded. RBC thus provides a principled account of the heretofore undecided relation between the classic principles of perceptual organization and pattern recognition: The constraints toward regularization (Pragnanz) characterize not the complete object but the object's components. Representational power derives from an allowance of free combinations of the geons. A Principle of Componential Recovery can account for the major phenomena of object recognition: If an arrangement of two or three geons can be recovered from the input, objects can be quickly recognized even when they are occluded, novel, rotated in depth, or extensively degraded. The results from experiments on the perception of briefly presented pictures by human observers provide empirical support for the theory. Any single object can project an infinity of image configurations to the retina. The orientation of the object to the viewer can vary continuously, each giving rise to a different two-dimensional projection. The object can be occluded by other objects or texture fields, as when viewed behind foliage. The object need not be presented as a full-colored textured image but instead can be a simplified line drawing. Moreover, the object can even be missing some of its parts or be a novel exemplar of its particular category. But it is only with rare exceptions that an image fails to be rapidly and readily classified, either as an instance of a familiar object category or as an instance that cannot be so classified (itself a form of classification).

Image Retrieval and Classification TechniquesAdvanced Image and Video Retrieval TechniquesVisual perception and processing mechanismsPsychologyCognitive psychologyImage (mathematics)Cognitive scienceArtificial intelligenceComputer science

MeSH terms

AttentionConcept FormationDiscrimination LearningForm PerceptionHumansOrientationPattern Recognition, Visual
Citations
5,528
FWCI
81.21
field-weighted impact
References
73
Percentile
100%
vs. same field & year
Citations per year
Cited by
Deep Learning for Generic Object Detection: A Survey
International Journal of Computer Vision · 2019 · 2,702 citations
LabelMe: A Database and Web-Based Tool for Image Annotation
International Journal of Computer Vision · 2007 · 4,112 citations
A benchmark for 3D mesh segmentation
ACM Transactions on Graphics · 2009 · 644 citations
Face recognition
ACM Computing Surveys · 2003 · 6,131 citations
What causes the face inversion effect?
Journal of Experimental Psychology Human Perception & Performance · 1995 · 540 citations
References
Visual routines
Cognition · 1984 · 1,036 citations
Parts of recognition
Cognition · 1984 · 1,337 citations
The psychology of computer vision
Pattern Recognition · 1976 · 2,574 citations
Perceptual grouping and attention in visual search for features and for objects.
Journal of Experimental Psychology Human Perception & Performance · 1982 · 880 citations
Uncertainty and Structure as Psychological Concepts
The American Journal of Psychology · 1964 · 1,515 citations
Features of similarity.
Psychological Review · 1977 · 7,238 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.