Predictive analysis of psoriasis risk factors using data mining techniques
Abstract
Psoriasis is a long-lasting, immune-driven inflammatory skin condition impacting around 2-3% of worldwide population, resulting in significant physical, psychological, and economic challenges. This research intends to develop a predictive model for psoriasis risk by employing data mining methods on a synthetic dataset containing 350 patient records. The dataset combines genetic markers (e.g., HLA-C*06:02), clinical factors (like BMI and PASI scores), and lifestyle elements (such as smoking and stress). Various machine learning algorithms were utilized specifically Random Forest, Support Vector Machine (SVM), and Neural Networks to evaluate their predictive abilities. Of these, the Random Forest model excelled compared to the others, reaching an accuracy of 86.7% and an AUC of 0.91. Analysis of feature importance indicated that obesity (BMI > 30), family history, and chronic stress were the key predictors, accounting for 23.4%, 19.8%, and 17.5% respectively. Significantly, a combination of obesity and family history raised the risk of psoriasis by 4.7 times (95% CI: 3.5-6.3). This research tackles significant limitations in current psoriasis studies by utilizing various data sources instead of relying exclusively on image-based diagnostics, applying SMOTE oversampling to balance class distributions, and improving model interpretability through SHAP values. Although dependent on synthetic data, the findings highlight the promise of machine learning for early detection and individualized risk evaluation. Future studies will investigate the incorporation of biological indicators like TNF-α along with mobile health apps for ongoing risk assessment. This study enhances precision dermatology by providing a validated, clinically interpretable framework for forecasting psoriasis risk and guiding preventive care measures.
How this paper connects to the literature. Drag to explore, click any node to open that paper.
