Scinovex
article Open AccessTop 10% cited

A comparison of random forest and its Gini importance with standard chemometric methods for the feature selection and classification of spectral data

BMC Bioinformatics · 2009 · Vol. 10(1) · pp. 213–213
Bjoern MenzeB. Michael KelmRalf MasuchUwe HimmelreichPeter BachertWolfgang PetrichFred A. Hamprecht

Abstract

The Gini importance of the random forest provided superior means for measuring feature relevance on spectral data, but - on an optimal subset of features - the regularized classifiers might be preferable over the random forest classifier, in spite of their limitation to model linear dependencies only. A feature selection based on Gini importance, however, may precede a regularized linear classification to identify this optimal subset of features, and to earn a double benefit of both dimensionality reduction and the elimination of noise from the classification task.

Spectroscopy and Chemometric AnalysesMetabolomics and Mass Spectrometry StudiesGene expression and cancer classificationRandom forestFeature selectionPattern recognition (psychology)Artificial intelligenceLinear discriminant analysisClassifier (UML)Feature (linguistics)MathematicsUnivariatePartial least squares regression

MeSH terms

ClassificationMagnetic Resonance SpectroscopyPattern Recognition, AutomatedRegression AnalysisSpectrophotometry, InfraredModels, StatisticalComputational Biology

Funding

  • Deutsche Forschungsgemeinschaft
Citations
1,281
FWCI
8.88
field-weighted impact
References
47
Percentile
98%
vs. same field & year
Citations per year
References
10.1162/153244303322753616
Applied Physics Letters · 2000 · 3,588 citations
Conditional variable importance for random forests
BMC Bioinformatics · 2008 · 3,188 citations
Random Forests
Machine Learning · 2001 · 121,242 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.