A Comparison of Document Clustering Techniques
Explore this paper's citation graph
Summary
This paper compares the two main approaches to document clustering, agglomerative hierarchical clustering and K-means, and indicates that the bisecting K-MEans technique is better than the standard K-Means approach and as good or better as the hierarchical approaches that were tested for a variety of cluster evaluation metrics.
- Type
- article
- Published
- 2000-05-23
- Cited by
- 3,319
- References
- 21
- Access
- Open access
- OpenAlex
- https://openalex.org/W1651093245
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:12808608
Keywords
Hierarchical clustering, Cluster analysis, Single-linkage clustering, Brown clustering, Complete-linkage clustering
References
- Information Retrieval Systems: Theory and Implementation, by Gerald Kowalski
- Finding Groups in Data: An Introduction to Cluster Analysis
- Fast and Intuitive Clustering of Web Documents
- Refining Initial Points for K-Means Clustering
- Hierarchically Classifying Documents Using Very Few Words
- Algorithms for Clustering Data
- Scatter/Gather: a cluster-based approach to browsing large document collections
- Incremental clustering and dynamic information retrieval
- On the merits of building categorization systems by supervised clustering
- Projections for efficient document clustering
- Optimization of inverted vector searches
- A practical clustering algorithm for static and dynamic information organization
- Fast and effective text mining using linear-time document clustering
- ROCK: a robust clustering algorithm for categorical attributes
- Comparison of Hierarchie Agglomerative Clustering Methods for Document Retrieval
- WebACE: a Web agent for document categorization and exploration
- TR 00-034 A Comparison of Document Clustering Techniques
Cited by
- A Survey on Clustering Algorithms for web Applications
- Peer Rewiring in Semantic Overlay Networks under Churn - (Short Paper)
- Improving text clustering for functional analysis of genes
- gCLUTO – An Interactive Clustering, Visualization, and Analysis System
- Incorporating Hyperlink Analysis in Web Page Clustering
- A Multiple-Domain Ontology Builder
- Semi-Automatic Web Information Extraction
- Text Clustering Based on Background Knowledge
- Socio‐Environmental Vulnerability Mapping for Environmental and Flood Resilience Assessment: The Case of Ageing and Poverty in the City of Wrocław, Poland
- Distributed Learning from Multiple EHR Databases: Contextual Embedding Models for Medical Events
- Approximation algorithms for clustering streams and large data sets
- Détection de communautés dans les réseaux d'information utilisant liens et attributs. (Community detection in information networks using links and attributes)
- A Roadmap for Web Mining
- A Semantic approach for effective document clustering using WordNet
- Effective and Efficient Correlation Analysis with Application to Market Basket Analysis and Network Community Detection.
- Document Clustering using Word Sense Disambiguation
- Clustering Unstructured Text Documents Using Fading Function
- Exploiting Distribution Skew for Scalable P2P Text Clustering
- Email Grouping Method
- Binary methods in data mining
Related papers
- Criterion Functions for Document Clustering ∗ Experiments and Analysis
- A vector space model for automatic indexing
- Indexing by Latent Semantic Analysis
- Data Mining: Concepts and Techniques
- Some methods for classification and analysis of multivariate observations
- Web document clustering: a feasibility demonstration
- An algorithm for suffix stripping
- Empirical and Theoretical Comparisons of Selected Criterion Functions for Document Clustering
- Frequent term-based text clustering