Julio Warmansyah; Safrial Safrial; Alam Supriatna; Wiwit Thoyyibah
Hypertension is one of the leading non-communicable diseases contributing significantly to cardiovascular morbidity and mortality worldwide. Despite the availability of extensive electronic medical record data in healthcare institutions, these data are often utilized only for administrative reporting rather than predictive analysis. Consequently, opportunities to identify age groups with a higher probability of developing hypertension remain underutilized. This study aims to implement the Naïve Bayes classification algorithm to analyze age distribution and classify the risk of hypertension among patients using healthcare data. The research adopted the Cross-Industry Standard Process for Data Mining (CRISP-DM) methodology, including business understanding, data understanding, data preparation, modeling, evaluation, and deployment. Patient medical record data consisting of demographic and clinical attributes, including age, systolic blood pressure, diastolic blood pressure, body weight, gender, and hypertension status, were processed using the Naïve Bayes algorithm. Model performance was evaluated using a confusion matrix by measuring accuracy, precision, recall, specificity, and balanced accuracy. The implementation demonstrates that the Naïve Bayes algorithm is capable of classifying hypertension risk efficiently while providing probabilistic information regarding age groups with a higher tendency to experience hypertension. The resulting classification model offers an effective decision-support tool for healthcare providers in conducting targeted screening, preventive interventions, and evidence-based health planning. The findings also indicate that data mining techniques can transform routinely collected medical records into valuable clinical knowledge for early hypertension prevention and healthcare decision-making.