Finnish Language Modeling with Deep Transformer Models
Explore this paper's citation graph
Summary
This project investigates the performance of the Transformer architectures-BERT and Transformer-XL for the language modeling task with a sub-word model setting with the Finnish language and compares it to the previous State of the art (SOTA) LSTM model.
- Type
- preprint
- Published
- 2020-03-14
- Cited by
- 0
- References
- 22
- Access
- Open access
- OpenAlex
- https://openalex.org/W3013624417
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:214667211
Keywords
Perplexity, Language model, Transformer, Computer science, Artificial intelligence
References
- Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books
- Neural Machine Translation of Rare Words with Subword Units
- Unsupervised models for morpheme segmentation and morphology learning
- Long Short-Term Memory
- Morfessor 2.0: Toolkit for statistical morphological segmentation
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Investigating Bidirectional Recurrent Neural Network Language Models for Speech Recognition
- Fine-tuned Language Models for Text Classification
- Deep Contextualized Word Representations
- Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
- Modern subword-based models for automatic speech recognition
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- A new algorithm for data compression
- Advances in subword-based HMM-DNN speech recognition across languages
- Attention is All you Need
- A new algorithm for data compression
- Adam: A Method for Stochastic Optimization
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- SGDR: Stochastic Gradient Descent with Restarts
Cited by
Related papers
- Transformer-IC: The Solution to Information Loss
- Learning Deep Transformer Models for Machine Translation
- Syntax-Infused Transformer and BERT models for Machine Translation and Natural Language Understanding
- Re-Transformer: A Self-Attention Based Model for Machine Translation
- Effective Architectures for Low Resource Multilingual Named Entity Transliteration
- Transformers for Low-Resource Languages: Is Féidir Linn!