b-Bit Minwise Hashing for Large-Scale Learning
Explore this paper's citation graph
Summary
It is demonstrated that b-bit minwise hashing can be naturally integrated with linear learning algorithms such as linear SVM and logistic regression, to solve large-scale and high-dimensional statistical learning tasks, especially when the data do not fit in memory.
- Type
- article
- Published
- 2011-12-01
- Cited by
- 4
- References
- 25
- OpenAlex
- https://openalex.org/W79052446
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:16063947
Keywords
Hash function, Computer science, Context (archaeology), Algorithm, Dynamic perfect hashing
References
- Accurate Estimators for Improving Minwise Hashing and b-Bit Minwise Hashing
- b-Bit Minwise Hashing in Practice: Large-Scale Batch and Online Learning and Using GPUs for Fast Preprocessing with Simple Hash Functions
- Finding text reuse on the web
- On compressing social networks
- Efficient detection of large-scale redundancy in enterprise file systems
- Training linear SVMs in linear time
- Near-Optimal Hashing Algorithms for Approximate Nearest Neighbor in High Dimensions
- Extraction and classification of dense implicit communities in the Web graph
- Very sparse random projections
- Feature hashing for large scale multitask learning
- Theory and applications of b-bit minwise hashing
- LIBLINEAR: A Library for Large Linear Classification
- b-Bit Minwise Hashing for Estimating Three-Way Similarities
- On the resemblance and containment of documents
- b-Bit minwise hashing
- Detecting near-duplicates for web crawling
- An axiomatic approach for result diversification
- Syntactic Clustering of the Web
- A dual coordinate descent method for large-scale linear SVM
- A large‐scale study of the evolution of Web pages
Cited by
Related papers
- Sparse multinomial logistic regression: fast algorithms and generalization bounds
- Maximum Entropy Based Associative Regression for Sparse Datasets
- Semi Supervised Logistic Regression
- Efficient Online Learning for Large-Scale Sparse Kernel Logistic Regression
- Sparse Kernel Logistic Regression using Incremental Feature Selection for Text-Independent Speaker Identification
- Unsupervised Supervised Learning I: Estimating Classification and Regression Errors without Labels
- Estimating expected error rates of random forest classifiers: A comparison of cross-validation and bootstrap