Scinovex
article Open AccessTop 1% cited

Evaluation metrics and statistical tests for machine learning

Scientific Reports · 2024 · Vol. 14(1) · pp. 6086–6086
Oona RainioJarmo TeuhoRiku Klén

Abstract

Research on different machine learning (ML) has become incredibly popular during the past few decades. However, for some researchers not familiar with statistics, it might be difficult to understand how to evaluate the performance of ML models and compare them with each other. Here, we introduce the most common evaluation metrics used for the typical supervised ML tasks including binary, multi-class, and multi-label classification, regression, image segmentation, object detection, and information retrieval. We explain how to choose a suitable statistical test for comparing models, how to obtain enough values of the metric for testing, and how to perform the test and interpret its results. We also present a few practical examples about comparing convolutional neural networks used to classify X-rays with different lung infections and detect cancer tumors in positron emission tomography images.

COVID-19 diagnosis using AIAI in cancer detectionRadiomics and Machine Learning in Medical ImagingComputer scienceArtificial intelligenceMetric (unit)Machine learningBinary classificationConvolutional neural networkStatistical hypothesis testingPattern recognition (psychology)Class (philosophy)Segmentation

MeSH terms

Machine LearningSupervised Machine LearningImage Processing, Computer-AssistedNeural Networks, ComputerPositron-Emission Tomography

Funding

  • Suomen Kulttuurirahasto
  • Jenny ja Antti Wihurin Rahasto
Citations
872
FWCI
352.52
field-weighted impact
References
60
Percentile
100%
vs. same field & year
Citations per year
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.