k-means++: the advantages of careful seeding
Explore this paper's citation graph
Summary
By augmenting k-means with a very simple, randomized seeding technique, this work obtains an algorithm that is Θ(logk)-competitive with the optimal clustering.
- Type
- article
- Published
- 2007-01-07
- Cited by
- 10,621
- References
- 27
- OpenAlex
- https://openalex.org/W2073459066
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:1782131
Keywords
Seeding, Cluster analysis, Computer science, Simplicity, Simple (philosophy)
References
- A fast hybrid k-means level set algorithm for segmentation
- Approximation schemes for clustering problems
- The Effectiveness of Lloyd-Type Methods for the k-Means Problem
- How slow is the k-means method?
- Clustering Large Graphs via the Singular Value Decomposition
- Large-scale clustering of cDNA-fingerprinting data.
- How Fast Is the k-Means Method?
- Worst-case and Smoothed Analysis of the ICP Algorithm, with an Application to the k-means Method
- On coresets for k-means and k-median clustering
- Optimal Time Bounds for Approximate Clustering
- Applications of weighted Voronoi diagrams and randomization to variance-based k-clustering: (extended abstract)
- Better streaming algorithms for clustering problems
- A simple linear time (1 + /spl epsiv/)-approximation algorithm for k-means clustering in any dimensions
- k-means projective clustering
- Clustering Data Streams: Theory and Practice
- Least squares quantization in PCM
- Online facility location
- A local search approximation algorithm for k-means clustering
- Worst-case and Smoothed Analysis of the ICP Algorithm, with an Application to the k-means Method
- How fast is κ-means?
Cited by
- A separability index for clustering and classification problems with applications to cluster merging and systematic evaluation of clustering algorithms
- Varieties of capitalist debates: how institutions shape public conflicts on economic liberalization in the U.K., France, and Germany
- A Domain Aware Genetic Algorithm for the p-Median Problem
- Label Transfer by Measuring Compactness
- Adaptive dual control of topic-based information retrieval
- Fully automated registration of vibrational microspectroscopic images in histologically stained tissue sections
- Denoising Autoencoder Self-Organizing Map (DASOM)
- Self-Similarity Constrained Sparse Representation for Hyperspectral Image Super-Resolution
- Acoustic source localization with microphone arrays for remote noise monitoring in an Intensive Care Unit
- A Three-Dimensional Hough Transform-Based Track-Before-Detect Technique for Detecting Extended Targets in Strong Clutter Backgrounds
- Statistical analysis of RNA-seq data from next- generation sequencing technology
- A novel construction of connectivity graphs for clustering and visualization
- Indexation de séquences de descripteurs
- "Fulfilling the Needs of Gray-Sheep Users in Recommender Systems, A Clustering Solution"
- Red tide detection using remotely sensed data: A case study of Sabah, Malaysia
- On the Complexity of Minimum Sum-of-Squares Clustering
- Problems and Mitigation Strategies for Developing and Validating Statistical Cyber Defenses
- Approximation Algorithms and New Models for Clustering and Learning
- Adaptive surrogate models for reliability analysis and reliability-based design optimization
- Center Based Clustering: A Foundational Perspective
Related papers
- UCI Machine Learning Repository
- Deep Residual Learning for Image Recognition
- Visualizing Data using t-SNE
- On Spectral Clustering: Analysis and an algorithm
- An Efficient k-Means Clustering Algorithm: Analysis and Implementation
- Distinctive Image Features from Scale-Invariant Keypoints
- Least squares quantization in PCM
- A tutorial on spectral clustering
- Some methods for classification and analysis of multivariate observations