Cluster Validity Indices for Uncertain Data Based on A Non-Parametric Kernel Method

Authors

  • Changwan Ko Department of Industrial Engineering, Chonnam National University, Gwangju, Korea
  • Behnam Tavakkol Business Analytics Program School of Business, Stockton University, New Jersey, USA
  • Youngseon Jeong Department of Industrial Engineering and Interdisciplinary Program of Arts & Design Technology, Chonnam National University, Gwangju, Korea

DOI:

https://doi.org/10.23055/ijietap.2026.33.4.11487

Abstract

Cluster validity indices are essential for determining the true number of clusters and validating clustering results. Most existing indices represent each data object as a single point and therefore fail to capture uncertainty inherent in the data. Even indices that consider uncertainty often rely on predefined probability distributions, such as Gaussian distributions. In this study, we propose non-parametric cluster validity indices derived from conventional indices, namely the Dunn, Calinski-Harabasz, and Davies-Bouldin. The proposed indices do not require a specific prior probability measure and can handle arbitrary forms, sub-clusters, noise, and high dimensionality. These indices compute compactness (within-cluster dispersion) and separability (between-cluster separation) in a reproducing kernel Hilbert space using a kernel function to assess cluster validity. Experiments on both artificial benchmark and real-world astronomical datasets demonstrate that the proposed indices are superior to existing cluster validity indices for validating the results of clustering problems.

Published

2026-07-28

How to Cite

Ko, C., Tavakkol, B., & Jeong, Y. (2026). Cluster Validity Indices for Uncertain Data Based on A Non-Parametric Kernel Method. International Journal of Industrial Engineering: Theory, Applications and Practice, 33(4). https://doi.org/10.23055/ijietap.2026.33.4.11487

Issue

Section

Statistical Analysis