Evidence
Same model and position · retrieved passages supplied · first output
Interpretability remains a central challenge as placement-based analyses grow in scale and complexity. While placement data carry rich phylogenetic signal, the resulting sample comparisons are often difficult to interpret, since the axes of classical ordination plots lack a ready biological reading and the internal nodes of distance-based trees resist intuitive meaning. Methods that explicitly exploit the special structure of placement data, rather than treating placements as generic distance vectors, offer a promising route toward more transparent and interpretable comparisons.
Passages supplied to the Evidence version
Edge principal components and squash clustering: using the special structure of phylogenetic placement data for sample comparison
Principal components (PCA) and hierarchical clustering are two of the most heavily used techniques for analyzing the differences between nucleic acid sequence samples sampled from a given environment. However, a classical application of these techniques to distances computed between samples can lack…
Read full passage excerpt
Principal components (PCA) and hierarchical clustering are two of the most heavily used techniques for analyzing the differences between nucleic acid sequence samples sampled from a given environment. However, a classical application of these techniques to distances computed between samples can lack transparency because there is no ready interpretation of the axes of classical PCA plots, and it is difficult to assign any clear intuitive meaning to either the internal nodes or the edge lengths of trees produced by distance-based hierarchical clustering methods such as UPGMA. We show that more interesting and interpretable results are produced by two new methods that leverage the special structure of phylogenetic placement data.