Distributed Asynchronous Online Learning for Natural Language Processing
Explore this paper's citation graph
Summary
This work generalizes existing asynchronous algorithms and experiment extensively with structured prediction problems from NLP, including discriminative, unsupervised, and non-convex learning scenarios, showing asynchronous learning can provide substantial speedups compared to distributed and single-processor mini-batch algorithms.
- Type
- article
- Published
- 2010-07-15
- Cited by
- 44
- References
- 31
- Access
- Open access
- OpenAlex
- https://openalex.org/W1554447113
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:54132
Keywords
Computer science, Asynchronous communication, Exploit, Focus (optics), Artificial intelligence
References
- On-line EM Algorithm for the Normalized Gaussian Network
- Why Doesn’t EM Find Good HMM POS-Taggers?
- Building a Large Annotated Corpus of English: The Penn Treebank
- Log-Linear Models for Wide-Coverage CCG Parsing
- The Mathematics of Statistical Machine Translation: Parameter Estimation
- Online EM for Unsupervised Models
- Learning structured prediction models: a large margin approach
- Large Margin Methods for Structured and Interdependent Output Variables
- Online Large-Margin Training of Syntactic and Structural Translation Features
- Fully distributed EM for very large datasets
- Gradient-based learning applied to document recognition
- Fast, Easy, and Cheap: Construction of Statistical Machine Translation Models with MapReduce
- New Ranking Algorithms for Parsing and Tagging: Kernels over Discrete Structures, and the Voted Perceptron
- A New Perceptron Algorithm for Sequence Labeling with Non-Local Features
- Efficient, Feature-based, Conditional Random Field Parsing
- An Evaluation Exercise for Word Alignment
- Ranking Algorithms for Named Entity Extraction: Boosting and the VotedPerceptron
- Efficient Large-Scale Distributed Training of Conditional Maximum Entropy Models
- Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
- Statistical Phrase-Based Translation
Cited by
- Predicting Linguistic Structure with Incomplete and Cross-Lingual Supervision
- Iterative parameter mixing for distributed large-margin training of structured predictors for natural language processing
- A Named Entity Recognition Method based on Decomposition and Concatenation of Word Chunks
- Efficient mini-batch training for stochastic optimization
- Minibatch and Parallelization for Online Large Margin Structured Learning
- Fast and Adaptive Online Training of Feature-Rich Translation Models
- Optimal Distributed Online Prediction Using Mini-Batches
- Optimal Distributed Online Prediction
- Hope and Fear for Discriminative Training of Statistical Translation Models
- Fast Distributed Asynchronous SGD with Variance Reduction
- A zealous parallel gradient descent algorithm
- A system analysis of improvements in machine learning
- MapReduce/Bigtable for Distributed Optimization
- Word Sense Disambiguation: A Structured Learning Perspective
- Asynchronous Parallel Learning for Neural Networks and Structured Models with Dense Features
- Small Batch or Large Batch?: Gaussian Walk with Rebound Can Teach
- Iteration acceleration for distributed learning systems
- Random Barzilai-Borwein step size for mini-batch algorithms
- Robust Fully Distributed Minibatch Gradient Descent with Privacy Preservation
- Efficient Privacy-preserving Machine Learning in Hierarchical Distributed System
Related papers
- Large Scale Distributed Deep Networks
- Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
- Slow Learners are Fast
- Optimal Distributed Online Prediction Using Mini-Batches
- Parallelized Stochastic Gradient Descent
- Distributed Training Strategies for the Structured Perceptron
- Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
- Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
- Discriminative Training Methods for Hidden Markov Models: Theory and Experiments with Perceptron Algorithms
- Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers