A New Dimension Reduction Method: Factor Discriminant K-means

Vichi, Maurizio; Rocci, Roberto; Stefano Antonio Gattone,; Rocci, Roberto

doi:10.1007/s00357-011-9085-9

Reduced K-means (RKM) and Factorial K-means (FKM) are two data reduction techniques incorporating principal component analysis and K-means into a unified methodology to obtain a reduced set of components for variables and an optimal partition for objects. RKM finds clusters in a reduced space by maximizing the between-clusters deviance without imposing any condition on the within-clusters deviance, so that clusters are isolated but they might be heterogeneous. On the other hand, FKM identifies clusters in a reduced space by minimizing the within-clusters deviance without imposing any condition on the between-clusters deviance. Thus, clusters are homogeneous, but they might not be isolated. The two techniques give different results because the total deviance in the reduced space for the two methodologies is not constant; hence the minimization of the within-clusters deviance is not equivalent to the maximization of the between-clusters deviance. In this paper a modification of the two techniques is introduced to avoid the afore mentioned weaknesses. It is shown that the two modified methods give the same results, thus merging RKM and FKM into a new methodology. It is called Factor Discriminant K-means (FDKM), because it combines Linear Discriminant Analysis and K-means. The paper examines several theoretical properties of FDKM and its performances with a simulation study. An application on real-world data is presented to show the features of FDKM.

A New Dimension Reduction Method: Factor Discriminant K-means / Vichi, M., Roberto, R., Stefano Antonio, G., Rocci, R.. - In: JOURNAL OF CLASSIFICATION. - ISSN 0176-4268. - STAMPA. - 28:2(2011), pp. 210-226. [10.1007/s00357-011-9085-9]

A New Dimension Reduction Method: Factor Discriminant K-means

VICHI, Maurizio;Roberto Rocci;Stefano Antonio Gattone;ROCCI, Roberto

2011

Abstract

Reduced K-means (RKM) and Factorial K-means (FKM) are two data reduction techniques incorporating principal component analysis and K-means into a unified methodology to obtain a reduced set of components for variables and an optimal partition for objects. RKM finds clusters in a reduced space by maximizing the between-clusters deviance without imposing any condition on the within-clusters deviance, so that clusters are isolated but they might be heterogeneous. On the other hand, FKM identifies clusters in a reduced space by minimizing the within-clusters deviance without imposing any condition on the between-clusters deviance. Thus, clusters are homogeneous, but they might not be isolated. The two techniques give different results because the total deviance in the reduced space for the two methodologies is not constant; hence the minimization of the within-clusters deviance is not equivalent to the maximization of the between-clusters deviance. In this paper a modification of the two techniques is introduced to avoid the afore mentioned weaknesses. It is shown that the two modified methods give the same results, thus merging RKM and FKM into a new methodology. It is called Factor Discriminant K-means (FDKM), because it combines Linear Discriminant Analysis and K-means. The paper examines several theoretical properties of FDKM and its performances with a simulation study. An application on real-world data is presented to show the features of FDKM.

Scheda breve

Scheda completa

	Anno di pubblicazione
	
				2011
			
	Parole chiave
	
				k-means; principal component analysis; dimension reduction; cluster analysis
			
	Tipologia
	
				01 Pubblicazione su rivista::01a Articolo in rivista
			
	Citazione
	
				A New Dimension Reduction Method: Factor Discriminant K-means / Vichi, M., Roberto, R., Stefano Antonio, G., Rocci, R.. - In: JOURNAL OF CLASSIFICATION. - ISSN 0176-4268. - STAMPA. - 28:2(2011), pp. 210-226. [10.1007/s00357-011-9085-9]
			
	Appartiene alla tipologia:
	
				01a Articolo in rivista

File allegati a questo prodotto

Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/377467

Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni

ND

19

18

Catalogo dei prodotti della ricerca