Pengelompokan Pasien Penyakit Kronis Diabetes, Hipertensi, dan Gagal Ginjal Menggunakan K-Means Berdasarkan Variabel Klinis
DOI:
10.25047/j-kes.v14i1.664Published:
2026-04-30Downloads
Abstract
The increasing number of patients with chronic diseases poses challenges in delivering effective and efficient healthcare services. The heterogeneity of patients’ clinical characteristics makes risk identification and clinical decision-making processes more complex. This study aimed to classify patients based on similarities in clinical characteristics using the K-Means algorithm to identify specific health patterns. The study employed a quantitative approach with a cross-sectional observational design and data mining techniques. Research data were obtained from a public Kaggle dataset which, after the data cleaning process, resulted in 51 patient records with analytical variables including age, blood pressure, blood glucose level, hemoglobin, diabetes mellitus, hypertension, anemia, and red blood cell (RBC) condition. Cluster analysis was performed using the K-Means algorithm, while cluster validity was evaluated using the Silhouette and Dunn indices. The results showed that the optimal number of clusters was three, with a Silhouette score of 0.537 and a Dunn index of 0.612. Cluster 1 (n=38) represented a low-risk profile characterized by relatively normal blood pressure and blood glucose levels without cases of diabetes mellitus or hypertension. Cluster 2 (n=8) was characterized by a high prevalence of anemia (87.5%) and low hemoglobin levels. Meanwhile, Cluster 3 (n=5) exhibited the most severe metabolic profile, with hypertension prevalence reaching 100%, diabetes mellitus 80%, and the highest average blood glucose level 348.8 mg/dL. These findings indicate that the K-Means method is effective in identifying patient segmentation based on distinct clinical characteristics. However, due to the relatively limited sample size and imbalanced cluster distribution, the findings should be interpreted as an exploratory analysis that requires further validation using larger and more diverse datasets.
Keywords: clinical pattern, data mining, health stratification, patient segmentation, unsupervised learning
License
Copyright (c) 2026 Fitra Tri Damayanti, Muhammad Iffran Ceria Rizal, Sakinah Derajad, Ihksan Kurnia Afnis, Joko Purnomo, Fitriah, Khairunnisa

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish in this journal agree to the following terms:
1. Copyright belongs to the medical journal as a publication
2. The author retains copyright and grants the journal rights to the first publication carried out simultaneously under a Creative Commons Attribution License which allows others to share the work with an acknowledgment of the author's work and initial publication in this journal.
3. Authors may enter into separate additional contractual arrangements for the non-exclusive distribution of the work (eg sending it to an institutional repository or publishing it in a book) with acknowledgment of initial publication in this journal.
4. Authors are permitted and encouraged to post work online (eg in institutional repositories or on their websites) before and during the submission process, as before and larger citations of published work (see Effects of Open Access).
Selengkapnya tentang teks sumber ini



