Search code examples
rhclustdendextend

dendextend: color_branches not working for certain hclust methods


I am using R dendextend package to plot hclust tree objects generated by each hclust method from hclust{stats}: "ward.D", "ward.D2", "single", "complete", "average" (= UPGMA), "mcquitty" (= WPGMA), "median" (= WPGMC) or "centroid" (= UPGMC).

I notice the color coding for color_branches fails when I use method = "median" or "centroid".

I tested it with randomly generated matrix and error is replicated for the "median" and "centroid" methods, is there a specific reason for this?

Please see the link for the output plots: fig1. hclust methods (a) ward.D2, (b) median, (c) centroid

library(dendextend)
set.seed(1)
df <- as.data.frame(replicate(10, rnorm(20)))
df.names <- rep(c("black", "red", "blue", "green", "cyan"), 2)
df.col <- rep(c("black", "red", "blue", "green", "cyan"), 2)
colnames(df) <- df.names
df.dist <- dist(t(df), method = "euclidean")

# plotting works for "ward.D", "ward.D2", "single", "complete", "average", "mcquitty"
dend <- as.dendrogram(hclust(df.dist, method = "ward.D2"), labels = df.names)
labels_colors(dend) <- df.col[order.dendrogram(dend)]
dend.colorBranch <- color_branches(dend, k = length(df.names), col = df.col[order.dendrogram(dend)])
dend.colorBranch %>% set("branches_lwd", 3) %>% plot(horiz = TRUE)

# color_branches fails for "median" or "centroid"
dend <- as.dendrogram(hclust(df.dist, method = "median"), labels = df.names)
labels_colors(dend) <- df.col[order.dendrogram(dend)]
dend.colorBranch <- color_branches(dend, k = length(df.names), col = df.col[order.dendrogram(dend)])
dend.colorBranch %>% set("branches_lwd", 3) %>% plot(horiz = TRUE)

dend <- as.dendrogram(hclust(df.dist, method = "centroid"), labels = df.names)
labels_colors(dend) <- df.col[order.dendrogram(dend)]
dend.colorBranch <- color_branches(dend, k = length(df.names), col = df.col[order.dendrogram(dend)])
dend.colorBranch %>% set("branches_lwd", 3) %>% plot(horiz = TRUE)

I am using dendextend_1.4.0. Session Info below:

sessionInfo()
R version 3.3.2 (2016-10-31)
Platform: x86_64-apple-darwin13.4.0 (64-bit)
Running under: macOS Sierra 10.12.3

Thanks.


Solution

  • You can solve this issue using branches_attr_by_clusters (although it could get a bit tricky, see the example below):

    library(dendextend)
    set.seed(1)
    df <- as.data.frame(replicate(10, rnorm(20)))
    df.names <- rep(c("black", "red", "blue", "green", "cyan"), 2)
    df.col <- rep(c("black", "red", "blue", "green", "cyan"), 2)
    colnames(df) <- df.names
    df.dist <- dist(t(df), method = "euclidean")
    
    # plotting works for "ward.D", "ward.D2", "single", "complete", "average", "mcquitty"
    dend <- as.dendrogram(hclust(df.dist, method = "ward.D2"), labels = df.names)
    labels_colors(dend) <- df.col[order.dendrogram(dend)]
    dend.colorBranch <- color_branches(dend, k = length(df.names), col = df.col[order.dendrogram(dend)])
    dend.colorBranch %>% set("branches_lwd", 3) %>% plot(horiz = TRUE)
    
    # color_branches fails for "median" or "centroid"
    dend <- as.dendrogram(hclust(df.dist, method = "median"), labels = df.names)
    aa <- df.col[order.dendrogram(dend)]
    labels_colors(dend) <- aa
    dend.colorBranch <- color_branches(dend, k = length(df.names), col = df.col[order.dendrogram(dend)])
    dend.colorBranch %>% set("branches_lwd", 3) %>% plot(horiz = TRUE)
    
    aa <- factor(aa, levels = unique(aa))
    dend %>% branches_attr_by_clusters(aa, value = levels(aa)) %>% plot
    

    enter image description here