Scalable Techniques for Clustering the Web
Explore this paper's citation graph
Summary
This paper aims to efficiently cluster similar pages on the web, using the technique of Locality-Sensitive Hashing (LSH), in which web pages are hashed in such a way that similar pages have a much higher probability of collision than dissimilar pages.
- Type
- article
- Published
- 2000-01-01
- Cited by
- 189
- References
- 14
- OpenAlex
- https://openalex.org/W192724328
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:46689205
Keywords
Computer science, Cluster analysis, Web page, Locality-sensitive hashing, Information retrieval
References
- Similarity Search in High Dimensions via Hashing
- Introduction to Modern Information Retrieval
- Min-wise independent permutations (extended abstract)
- A small approximately min-wise independent family of hash functions
- Computing Iceberg Queries Efficiently
- Finding interesting associations without support pruning
- A Best Possible Heuristic for the k-Center Problem
- An algorithm for suffix stripping
- Web document clustering: a feasibility demonstration
- WebBase: a repository of Web pages
- On the resemblance and containment of documents
- Syntactic Clustering of the Web
- Learning to Extract Symbolic Knowledge from the World Wide Web
- Min-Wise Independent Permutations
- Approximate nearest neighbors
Cited by
- Optimizing File Replication over Limited-Bandwidth Networks using Remote Differential Compression
- Binary methods in data mining
- Approximate Computation of Object Distances by Locality-Sensitive Hashing
- On Randomly Projected Hierarchical Clustering with Guarantees
- Enhancing Contents-Link Coupled Web Page Clustering and Its Evaluation
- A Sketch-based Sampling Algorithm on Sparse Data
- Parallelizing Data-Centric Programs
- Relevance Ranking for Vertical Search Engines
- Automated subject classification of textual Web pages, for browsing : Thesis for the degree of Licentiate in Philosophy, Swedish intermediate degree between Master’s and Doctoral degrees
- Dimension reduction of streaming data via random projections
- Similarity search and locality sensitive hashing using ternary content addressable memories
- Algorithmic applications of low-distortion geometric embeddings
- Efficient near duplicate document detection for specialized corpora
- An improved algorithm for locality-sensitive hashing
- The Impact of Feature Selection on Signature-Driven Spam Detection
- Decomposing the Web Graph into Parameterized Connected Components
- Web Pages Clustering: A New Approach
- Automated Subject Classification of Textual Documents in the Context of Web-Based Hierarchical Browsing
- Exploiting the Computational Power of Ternary Content Addressable Memory
- Scalable, Behavior-Based Malware Clustering
Related papers
- Approximate nearest neighbors
- On the resemblance and containment of documents
- Syntactic Clustering of the Web
- Similarity Search in High Dimensions via Hashing
- Similarity estimation techniques from rounding algorithms
- Min-Wise Independent Permutations
- Locality-sensitive hashing scheme based on p-stable distributions
- Web document clustering: a feasibility demonstration
- The Anatomy of a Large-Scale Hypertextual Web Search Engine
- Efficient large-scale sequence comparison by locality-sensitive hashing