Computational topology in high-dimensional data clustering and manifold learning
Abstract
The "curse for dimensionality", and complicated geometric structures make high-dimensional data clustering more difficult.This paper offers a framework using topological invariants (Betti numbers), and nonlinear dimensionality reduction through combining computational topology alongside manifold learning to solve these difficulties to assess the interaction between topological characteristics, and clustering efficiency, we create synthetic high-dimensional datasets (Torus, and Sphere manifolds) alongside regulated noise.We show, that topological persistence characteristics (, ) improve cluster separation within high-dimensional spaces through means for t-SNE for manifold learning, and comparative comparison for clustering techniques (K-means, DBSCAN, Hierarchical).Alongside DBSCAN beating other approaches within maintaining topological integrity, our findings indicate strong performance across measures: Adjusted Rand Index (ARI) scores for 0.85-0.87,and Normalized Mutual Information (NMI) scores for 0.78-0.81.This paper provides insights for uses within bioinformatics, image analysis, and network research through linking topological data analysis alongside machine learning.
How this paper connects to the literature. Drag to explore, click any node to open that paper.
