K-means clustering is widely used, but its iterative and initialization-sensitive nature makes results hard to interpret, compare, and debug, especially when multiple runs are needed to obtain a reliable solution. This paper presents VISPEK, a visual interactive system for progressive ensemble k-means clustering that helps users understand how clustering results evolve and how agreement emerges across runs. By combining Progressive Visual Analytics with an ensemble-based strategy, VISPEK exposes intermediate results, stability and quality metrics, and similarities among runs, enabling users to inspect, explain, and steer the clustering process before convergence. In this way, VISPEK supports the interpretation of consensus formation, highlights uncertainty, and helps users identify promising results early. We validate the approach through two usage scenarios and an expert study with data science and machine learning experts, showing that VISPEK improves analysis transparency while reducing time and computational effort.
VISPEK: a Visual Interactive System for Progressive Ensemble K-Means Clustering / Angelini, M., Blasilli, G., Cazzetta, G., Lenti, S., Palleschi, A., Santucci, G.. - (2026), pp. 1-9. (18th International Conference on Advanced Visual Interfaces, AVI 2026 ita ) [10.1145/3811427.3811436].
VISPEK: a Visual Interactive System for Progressive Ensemble K-Means Clustering
Marco Angelini;Graziano Blasilli;Giorgio Cazzetta;Simone Lenti;Alessia Palleschi;Giuseppe Santucci
2026
Abstract
K-means clustering is widely used, but its iterative and initialization-sensitive nature makes results hard to interpret, compare, and debug, especially when multiple runs are needed to obtain a reliable solution. This paper presents VISPEK, a visual interactive system for progressive ensemble k-means clustering that helps users understand how clustering results evolve and how agreement emerges across runs. By combining Progressive Visual Analytics with an ensemble-based strategy, VISPEK exposes intermediate results, stability and quality metrics, and similarities among runs, enabling users to inspect, explain, and steer the clustering process before convergence. In this way, VISPEK supports the interpretation of consensus formation, highlights uncertainty, and helps users identify promising results early. We validate the approach through two usage scenarios and an expert study with data science and machine learning experts, showing that VISPEK improves analysis transparency while reducing time and computational effort.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


