A comparative study of ensemble-based imputation techniques for handling missing data
Abstract
Missing data is a common and critical issue in census studies as it can cause biased estimates, reduce statistical power and lead to invalid inferences in statistical analysis. This study examines the performance of traditional, machine learning–based, and ensemble imputation techniques using the US Arrests and Swiss Fertility and Socioeconomic Indicators datasets. Artificial missingness was introduced at 5%, 10%, and 15% levels under a Missing Completely at Random mechanism to enable systematic evaluation. Individual imputation methods, including mean, zero, K-nearest neighbours, multiple imputation by chained equations, and random forest, were applied alongside four ensemble-based imputation strategies formed through simple averaging. Imputation accuracy was assessed using root mean squared error and mean absolute error. The results demonstrate that machine learning-based methods outperform traditional approaches, while ensemble strategies combining strong base learners achieve the lowest errors across both datasets. The findings indicate that well-designed ensemble imputation methods can improve robustness and accuracy in handling missing data for population-based statistical analyses.
How this paper connects to the literature. Drag to explore, click any node to open that paper.
