Scinovex
article Open Access

A comparative study of ensemble-based imputation techniques for handling missing data

Rana Krina DivyeshbhaiPankaj DasTauqueer AhmadAnkur BiwasArpitha TD

Abstract

Missing data is a common and critical issue in census studies as it can cause biased estimates, reduce statistical power and lead to invalid inferences in statistical analysis. This study examines the performance of traditional, machine learning–based, and ensemble imputation techniques using the US Arrests and Swiss Fertility and Socioeconomic Indicators datasets. Artificial missingness was introduced at 5%, 10%, and 15% levels under a Missing Completely at Random mechanism to enable systematic evaluation. Individual imputation methods, including mean, zero, K-nearest neighbours, multiple imputation by chained equations, and random forest, were applied alongside four ensemble-based imputation strategies formed through simple averaging. Imputation accuracy was assessed using root mean squared error and mean absolute error. The results demonstrate that machine learning-based methods outperform traditional approaches, while ensemble strategies combining strong base learners achieve the lowest errors across both datasets. The findings indicate that well-designed ensemble imputation methods can improve robustness and accuracy in handling missing data for population-based statistical analyses.

Statistical Methods and Bayesian Inferencedemographic modeling and climate adaptationSurvey Methodology and NonresponseImputation (statistics)Missing dataRobustness (evolution)Statistical powerEnsemble learningRandom forest
Citations
0
FWCI
0.00
field-weighted impact
References
0
Percentile
26%
vs. same field & year
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

A comparative study of ensemble-based imputation techniques for handling missing data · Scinovex