Interrelating Fruchterman-Reingold Graph Visualization and Agglomerative Clustering
Authors: Alexandre Benatti, Luciano da F. Costa
Organizations: Institute of Mathematics and Statistics - DCC University of S˜ao Paulo Rua do Mat˜ao, 1010, S˜ao Paulo, SP 05508-090 Brazil · S˜ao Carlos Institute of Physics - DFCM University of S˜ao Paulo Av. Trabalhador S˜ao-Carlense, 400, S˜ao Carlos, SP 13566-590 Brazil
Graph visualization methods and agglomerative clustering have been frequently considered in data analysis and pattern recognition. Because these approaches are interrelated and complementary, it is of particular interest to investigate their associations. In this work, we study the possible relationship between the Fruchterman-Reingold graph visualization method and four types of agglomerative clustering adopting single- and complete-linkage, average, and Ward's linkage criteria. Three types of datasets have been considered in 2 and 10 dimensions, as well as the PCA projection of the latter to two dimensions. The results obtained suggest that the relationship between the methods considered did not vary much for the three types of data mentioned above. At the same time, the agglomerative methods tended to yield results that are mostly similar to each other, while presenting moderate similarity with the original data. The Fruchterman-Reingold visualization resulted similar to the original data, but exhibited relatively smaller similarity to the agglomerative methods.
Figures & tables
Figure 1 : Flow diagram illustrating the approach adopted for relating Fructerman-Reingold graph visualization and four types of agglomerative clustering.
Figure 2 : Example of 2D dataset considered in our study (a) and respective visualization by using the Fruchterman-Reingold method on the reciprocal of the Euclidean distances.
Figure 3 : Dendrograms obtained by the agglomerative clustering approaches adopting single- (a), complete- (b), average (c), and Ward’s (d) linkage criteria. Observe that different upper limits of the distance axes have been chosen for the sake of improved visualization.
Figure 4 : Coincidence similarity matrix ( D=1 ) obtained for the dataset in Fig. 2 (a).
Figure 5 : Average coincidence similarity matrix ( D=1 ) obtained for 5,000 datasets, considering two-dimensional data.
Figure 6 : Average coincidence similarity matrix ( D=1 ) obtained for 5,000 ten-dimensional datasets (a) and for its two-dimensional PCA projection.
Figure 7 : Correlograms relating the similarities between the visualization and agglomerative clustering obtained for three data sets considering two and ten dimensions, and the 2D PCA projection of the latter. The dashed line indicates the linear data regression.
Khoury College of Computer Sciences, Northeastern University, Seattle · School of Computation, Information and Technology, Technical University of Munich, Heilbronn, Germany