An Explainable Random Forest Model for Early Hypertension Risk Prediction: Integrating SHAP and Breakdown Analysis

Authors
Category Primary study
Pre-printResearchSquare
Year 2026
Hypertension remains a leading global cause of cardiovascular mortality, but it often goes undiagnosed due to its silent progression. This study developed and evaluated an explanatory random forest model for the early prediction of hypertension risk using clinical and demographic data. A publicly available dataset of 4,240 medical records with 12 predictor variables (age, gender, blood pressure, body mass index, cholesterol, smoking status, glucose, heart rate, and others) was analyzed. After preprocessing the median imputation of missing values and removing predictors with near-zero variance, the data were divided into 70% training and 30% test sets. A random forest classifier was trained with 5-fold cross-validation to tune the mtry hyperparameter from 2 to 12, optimizing for the area under the receiver operating characteristic curve (AUC). The final model achieved an AUC value of 0.9365, an accuracy of 88.99% (95% CI: 87.13%-90.65%), a sensitivity of 79.45%, a specificity of 93.91%, and an F1 score of 0.8309, with the best mtry value being 2. Systolic and diastolic blood pressure were the most powerful predictors, followed by age, BMI, and heart rate. Model interpretability was improved using SHAP (SHapley Additive Explanations) values and DALEX breakdown plots, providing transparent, single-level explanations. While high specificity indicates excellent ability to exclude non-risk individuals, moderate sensitivity suggests room for improvement in identifying true hypertension cases. This study demonstrates that a tuned random forest model combined with explainable AI can serve as an effective, interpretable screening tool for hypertension risk stratification, data-driven clinical decisions, and early intervention strategies.
Epistemonikos ID: dd7318ad66e2243788d7b6b2fec45c50fa6bb0d7
First added on: May 17, 2026