Scinovex
articleTop 1% cited

Support vector machines for spam categorization

IEEE Transactions on Neural Networks · 1999 · Vol. 10(5) · pp. 1048–1054
Harris DruckerDonghui WuVladimir Vapnik

Abstract

We study the use of support vector machines (SVM's) in classifying e-mail as spam or nonspam by comparing it to three other classification algorithms: Ripper, Rocchio, and boosting decision trees. These four algorithms were tested on two different data sets: one data set where the number of features were constrained to the 1000 best features and another data set where the dimensionality was over 7000. SVM's performed best when using binary features. For both data sets, boosting trees and SVM's had acceptable test performance in terms of accuracy and speed. However, SVM's had significantly less training time.

Text and Document Classification TechnologiesSpam and Phishing DetectionFace and Expression RecognitionSupport vector machineBoosting (machine learning)Computer scienceArtificial intelligencePattern recognition (psychology)Curse of dimensionalityDecision treeTraining setMachine learningText categorization
Citations
1,463
FWCI
17.74
field-weighted impact
References
27
Percentile
99%
vs. same field & year
Citations per year
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.