Scinovex
article Open AccessTop 1% cited

Random forest versus logistic regression: a large-scale benchmark experiment

BMC Bioinformatics · 2018 · Vol. 19(1) · pp. 270–270
Raphaël CouronnéPhilipp ProbstAnne‐Laure Boulesteix

Abstract

RF performed better than LR according to the considered accuracy measured in approximately 69% of the datasets. The mean difference between RF and LR was 0.029 (95%-CI =[0.022,0.038]) for the accuracy, 0.041 (95%-CI =[0.031,0.053]) for the Area Under the Curve, and - 0.027 (95%-CI =[-0.034,-0.021]) for the Brier score, all measures thus suggesting a significantly better performance of RF. As a side-result of our benchmarking experiment, we observed that the results were noticeably dependent on the inclusion criteria used to select the example datasets, thus emphasizing the importance of clear statements regarding this dataset selection process. We also stress that neutral studies similar to ours, based on a high number of datasets and carefully designed, will be necessary in the future to evaluate further variants, implementations or parameters of random forests which may yield improved accuracy compared to the original version with default values.

Machine Learning and Data ClassificationMachine Learning and AlgorithmsImbalanced Data Classification TechniquesRandom forestLogistic regressionBenchmark (surveying)Scale (ratio)StatisticsRegressionComputer scienceMachine learningArtificial intelligenceData mining

MeSH terms

AlgorithmsComputer SimulationHumansLogistic ModelsBenchmarkingDatabases as Topic

Funding

  • Deutsche Forschungsgemeinschaft
Citations
824
FWCI
39.43
field-weighted impact
References
43
Percentile
100%
vs. same field & year
Citations per year
References
Greedy function approximation: A gradient boosting machine.
The Annals of Statistics · 2001 · 27,794 citations
Extremely randomized trees
Machine Learning · 2006 · 8,332 citations
Bootstrap Methods and Their Application
Journal of the American Statistical Association · 1999 · 5,530 citations
Random Forests
Machine Learning · 2001 · 121,242 citations
UCI Machine Learning Repository
Medical Entomology and Zoology · 2007 · 24,290 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

Random forest versus logistic regression: a large-scale benchmark experiment · Scinovex