Feature hashing for large scale multitask learning
Explore this paper's citation graph
Summary
This paper provides exponential tail bounds for feature hashing and shows that the interaction between random subspaces is negligible with high probability, and demonstrates the feasibility of this approach with experimental results for a new use case --- multitask learning.
- Type
- article
- Published
- 2009-02-12
- Cited by
- 1,105
- References
- 31
- Access
- Open access
- OpenAlex
- https://openalex.org/W2070996757
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:291713
Keywords
Computer science, Feature (linguistics), Scale (ratio), Hash function, Artificial intelligence
References
- Similarity Search in High Dimensions via Hashing
- The concentration of measure phenomenon
- Small Statistical Models by Random Feature Mixing
- THE THEORY OF PROBABILITIES.
- A sparse Johnson: Lindenstrauss transform
- Problems and results in extremal combinatorics--I
- Frustratingly Easy Domain Adaptation
- Random Features for Large-Scale Kernel Machines
- Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform
- The Netflix Prize
- Advances in Neural Information Processing Systems 21
- On variants of the Johnson–Lindenstrauss lemma
- Conditional Random Sampling: A Sketch-based Sampling Technique for Sparse Data
- Database-friendly random projections: Johnson-Lindenstrauss with binary coins
- Dense Fast Random Projections and Lean Walsh Transforms
- Hash Kernels
- An improved data stream summary: the count-min sketch and its applications
- Problems and results in Extremal Combinatorics , Part
- Advances in Neural Information Processing Systems 20
- Electronic Colloquium on Computational Complexity, Report No. 70 (2007) Fast Dimension Reduction Using Rademacher Series on Dual BCH Codes
Cited by
- New features for query dependent sponsored search click prediction
- Streaming and Sketch Algorithms for Large Data NLP
- Collaborative Filtering on a Budget
- Approximation and Relaxation Approaches for Parallel and Distributed Machine Learning
- b-Bit Minwise Hashing for Large-Scale Learning
- Active Multitask Learning Using Both Latent and Supervised Shared Topics
- Combining Hashing and Abstraction in Sparse High Dimensional Feature Spaces
- Large Scale Malware Analysis, Detection and Signature Generation
- Supervised Machine Learning Under Test-Time Resource Constraints: A Trade-off Between Accuracy and Cost
- Predicting Linguistic Structure with Incomplete and Cross-Lingual Supervision
- Interactive Learning for Sequential Decisions and Predictions
- Towards Automatic Software Lineage Inference
- Semi-supervised constraints preserving hashing
- A Randomized Algorithm for CCA
- Recent advances and emerging challenges of feature selection in the context of big data
- Large scale support vector machines algorithms for visual classification
- Learning with single view co-training and marginalized dropout
- Portfolio Allocation for Sellers in Online Advertising
- Models, Inference, and Implementation for Scalable Probabilistic Models of Text
- Contributions to large-scale learning for image classification. (Contributions à l'apprentissage grande échelle pour la classification d'images)
Related papers
- CLIP-HASH: A Lightweight Hashing Network for Cross-Modal Retrieval
- Deep binary constraint hashing for fast image retrieval
- A visual saliency based video hashing algorithm
- Omni-directional Feature Learning for Person Re-identification
- NAPHash: Efficient Image Hash to Reduce Dataset Redundancy
- Deep Self-taught Hashing for Image Retrieval
- DCCH: Deep Continuous Center Hashing for Image Retrieval
- Fast Retrieval Method of Image Data Based on Learning to Hash
- Disturbance Consistent Self-Ensembling for Semi-Supervised Hashing
- Hashing Based Fast Palmprint Identification for Large-Scale Databases