Scinovex
review Open AccessTop 1% cited

Machine learning in automated text categorization

ACM Computing Surveys · 2002 · Vol. 34(1) · pp. 1–47
Fabrizio Sebastiani

Abstract

The automated categorization (or classification) of texts into predefined categories has witnessed a booming interest in the last 10 years, due to the increased availability of documents in digital form and the ensuing need to organize them. In the research community the dominant approach to this problem is based on machine learning techniques: a general inductive process automatically builds a classifier by learning, from a set of preclassified documents, the characteristics of the categories. The advantages of this approach over the knowledge engineering approach (consisting in the manual definition of a classifier by domain experts) are a very good effectiveness, considerable savings in terms of expert labor power, and straightforward portability to different domains. This survey discusses the main approaches to text categorization that fall within the machine learning paradigm. We will discuss in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.

Text and Document Classification TechnologiesAdvanced Text Analysis TechniquesAlgorithms and Data CompressionComputer scienceCategorizationSoftware portabilityArtificial intelligenceClassifier (UML)Machine learningText categorizationNatural language processing
Citations
7,857
FWCI
293.28
field-weighted impact
References
204
Percentile
100%
vs. same field & year
Citations per year
Cited by
Text-Based Network Industries and Endogenous Product Differentiation
Journal of Political Economy · 2016 · 2,000 citations
Learning multi-label scene classification
Pattern Recognition · 2004 · 2,279 citations
Ensemble of keyword extraction methods and classifiers in text classification
Expert Systems with Applications · 2016 · 621 citations
Metrics for Polyphonic Sound Event Detection
Applied Sciences · 2016 · 561 citations
ML-KNN: A lazy learning approach to multi-label learning
Pattern Recognition · 2007 · 3,495 citations
References
BoosTexter: A Boosting-based System for Text Categorization
Machine Learning · 2000 · 2,181 citations
A vector space model for automatic indexing
Communications of the ACM · 1975 · 7,377 citations
Support vector machines for spam categorization
IEEE Transactions on Neural Networks · 1999 · 1,463 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.