Establishing Validity for Distance Functions and Internal Clustering Validity Indices in Correlation Space
Degen, Isabella, Abdallah, Zahraa S, Brown, Kate Robson, Reeve, Henry W J
–arXiv.org Artificial Intelligence
Internal clustering validity indices (ICVIs) assess clustering quality without ground truth labels. Comparative studies consistently find that no single ICVI outperforms others across datasets, leaving practitioners without principled ICVI selection. We argue that inconsistent ICVI performance arises because studies evaluate them based on matching human labels rather than measuring the quality of the discovered structure in the data, using datasets without formally quantifying the structure type and quality. Structure type refers to the mathematical organisation in data that clustering aims to discover. Validity theory requires a theoretical definition of clustering quality, which depends on structure type. We demonstrate this through the first validity assessment of clustering quality measures for correlation patterns, a structure type that arises from clustering time series by correlation relationships. We formalise 23 canonical correlation patterns as the theoretical optimal clustering and use synthetic data modelling this structure with controlled perturbations to evaluate validity across content, criterion, construct, and external validity. Our findings show that Silhouette Width Criterion (SWC) and Davies-Bouldin Index (DBI) are valid for correlation patterns, whilst Calinski-Harabasz (VRC) and Pakhira-Bandyopadhyay-Maulik (PBM) indices fail. Simple Lp norm distances achieve validity, whilst correlation-specific functions fail structural, criterion, and external validity. These results differ from previous studies where VRC and PBM performed well, demonstrating that validity depends on structure type. Our structure-type-specific validation method provides both practical guidance (quality thresholds SWC>0.9, DBI<0.15) and a methodological template for establishing validity for other structure types.
arXiv.org Artificial Intelligence
Dec-8-2025
- Country:
- Asia > China
- Jiangsu Province > Nanjing (0.04)
- North America > United States
- California
- Orange County > Irvine (0.04)
- Santa Clara County > Palo Alto (0.04)
- District of Columbia > Washington (0.04)
- California
- Asia > China
- Genre:
- Research Report
- Experimental Study > Negative Result (0.67)
- New Finding (1.00)
- Research Report
- Industry:
- Health & Medicine (0.67)
- Technology: