Predictions For Pre-training Language Models
Explore this paper's citation graph
Summary
This paper investigates whether it is still helpful to add the specific task's loss in pre-training step and uses the fine-tuned model to give the user-generated unlabeled data a pseudo-label to pre-train.
- Type
- preprint
- Published
- 2020-11-18
- Cited by
- 0
- References
- 47
- Access
- Open access
- OpenAlex
- https://openalex.org/W3098003361
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:227014523
Keywords
Task (project management), Computer science, Language model, Training set, Training (meteorology)
References
- Improving Translation Model by Monolingual Data
- Semi-Supervised Learning
- Introduction to Semi-Supervised Learning
- Combining labeled and unlabeled data with co-training
- Dropout: a simple way to prevent neural networks from overfitting
- Unsupervised Word Sense Disambiguation Rivaling Supervised Methods
- Probability of error of some adaptive pattern-recognition machines
- Self-Training PCFG Grammars with Latent Annotations Across Languages
- Democratic co-learning
- Class-Based n-gram Models of Natural Language
- A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data
- A Scalable Hierarchical Distributed Language Model
- Tri-training: exploiting unlabeled data using three classifiers
- Self-Training for Enhancement and Domain Adaptation of Statistical Parsers Trained on Small Datasets
- Distributed Representations of Words and Phrases and their Compositionality
- Domain Adaptation with Structural Correspondence Learning
- Word Representations: A Simple and General Method for Semi-Supervised Learning
- GloVe: Global Vectors for Word Representation
- Learning Distributed Representations of Sentences from Unlabelled Data
- Improving Neural Machine Translation Models with Monolingual Data
Cited by
No citing papers recorded for this paper.
Related papers
- Profile Prediction: An Alignment-Based Pre-Training Task for Protein Sequence Models
- FineText: Text Classification via Attention-based Language Model Fine-tuning
- MC-BERT: Efficient Language Pre-Training via a Meta Controller
- On the Transferability of Pre-trained Language Models: A Study from Artificial Datasets
- Revisiting Self-Training for Neural Sequence Generation
- An Empirical Investigation towards Efficient Multi-Domain Language Model Pre-training