Exploiting Distribution Skew for Scalable P2P Text Clustering
Explore this paper's citation graph
Summary
A new P2P algorithm which significantly reduces the communication costs involved by exploiting distribution skew, naturally found in text and other datasets, which achieves high clustering quality and requires no synchronization between peers.
- Type
- article
- Published
- 2008-01-01
- Cited by
- 8
- References
- 22
- OpenAlex
- https://openalex.org/W44364808
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:15771875
Keywords
Skew, Cluster analysis, Computer science, Scalability, Data mining
References
- HP2PC: Scalable Hierarchically-Distributed Peer-to-Peer Clustering
- K-Means Clustering Over a Large, Dynamic Network
- Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability; Vol. IV
- Fast and Intuitive Clustering of Web Documents
- Classifying Documents by Distributed P2P Clustering
- Heavy-tailed probability distributions in the World Wide Web
- A Comparison of Document Clustering Techniques
- Distributed data clustering can be efficient and exact
- Scatter/Gather: a cluster-based approach to browsing large document collections
- Block addressing indices for approximate text retrieval
- A practical guide to heavy tails: statistical techniques and applications
- Similarity discovery in structured P2P overlays
- Some methods for classification and analysis of multivariate observations
- A Comparison of Document, Sentence, and Term Event Spaces
- RCV1: A New Benchmark Collection for Text Categorization Research
- On the Bursty Evolution of Blogspace
- Chord: A scalable peer-to-peer lookup service for internet applications
- Shared memory parallelization of data mining algorithms: techniques, programming interface, and performance
- Human behavior and the principle of least effort
- Human behavior and the principle of least effort
Cited by
- Application of K-tree to document clustering
- Document clustering algorithms, representations and evaluation for information retrieval
- Models of distributed data clustering in peer-to-peer environments
- PCIR: Combining DHTs and peer clusters for efficient full-text P2P indexing
- Frequent term based peer-to-peer text clustering
- Approximate algorithms for efficient indexing, clustering, and classification in Peer-to-peer networks
- Decentralized Probabilistic Text Clustering
- Application of K-tree to Document Clustering Masters of IT by Research ( IT 60 )
Related papers
- Survey on text clustering algorithm -Research present situation of text clustering algorithm
- Survey on text clustering algorithm
- An evaluation of Hadoop cluster efficiency in document clustering using parallel K-means
- Data Mining Clustering Algorithm
- An Overview of Text Clustering
- A weighted seeds affinity propagation clustering for efficient document mining
- Peer sampling gossip-based distributed clustering algorithm for unstructured P2P networks
- An effective and efficient grid-based data clustering algorithm using intuitive neighbor relationship for data mining
- Clustering techniques in data mining: A comparison
- Combined Clustering Based on Data Mining