Scinovex
article Open AccessTop 10% cited

A vector space model for automatic indexing

Communications of the ACM · 1975 · Vol. 18(11) · pp. 613–620
Gerard SaltonAnita M.-Y. WongChul‐Su Yang

Abstract

In a document retrieval, or other pattern matching environment where stored entities (documents) are compared with each other or with incoming patterns (search requests), it appears that the best indexing (property) space is one where each entity lies as far away from the others as possible; in these circumstances the value of an indexing system may be expressible as a function of the density of the object space; in particular, retrieval performance may correlate inversely with space density. An approach based on space density computations is used to choose an optimum indexing vocabulary for a collection of documents. Typical evaluation results are shown, demonstating the usefulness of the model.

Image Retrieval and Classification TechniquesData Management and AlgorithmsData Mining Algorithms and ApplicationsSearch engine indexingVector space modelComputer scienceSpace (punctuation)Information retrievalMatching (statistics)Property (philosophy)ComputationFunction (biology)Data mining
Citations
7,377
FWCI
3.21
field-weighted impact
References
5
Percentile
93%
vs. same field & year
Citations per year
Cited by
Inverted files for text search engines
ACM Computing Surveys · 2006 · 1,067 citations
Automatic text summarization: A comprehensive survey
Expert Systems with Applications · 2020 · 771 citations
Machine learning in automated text categorization
ACM Computing Surveys · 2002 · 7,857 citations
EXPERT SYSTEMS WITH APPLICATIONS
Expert Systems with Applications · 2004 · 1,660 citations
Related articles
A vector space model for automatic indexing
Communications of the ACM · 1975 · 7,377 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.