Deep Contextualized Word Representations
Explore this paper's citation graph
Summary
A new type of deep contextualized word representation is introduced that models both complex characteristics of word use and how these uses vary across linguistic contexts, allowing downstream models to mix different types of semi-supervision signals.
- Type
- preprint
- Published
- 2018-02-15
- Cited by
- 12,252
- References
- 64
- Access
- Open access
- OpenAlex
- https://openalex.org/W2787560479
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:3626819
Keywords
Polysemy, Computer science, Natural language processing, Artificial intelligence, Word (group theory)
References
- ADADELTA: An Adaptive Learning Rate Method
- An Empirical Exploration of Recurrent Network Architectures
- Training Very Deep Networks
- One billion word benchmark for measuring progress in statistical language modeling
- Building a Large Annotated Corpus of English: The Penn Treebank
- A large annotated corpus for learning natural language inference
- Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
- Character-Aware Neural Language Models
- Using a Semantic Concordance for Sense Identification
- Long Short-Term Memory
- Dropout: a simple way to prevent neural networks from overfitting
- Towards Robust Linguistic Analysis using OntoNotes
- Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
- Distributed Representations of Words and Phrases and their Compositionality
- CoNLL-2012 Shared Task: Modeling Multilingual Unrestricted Coreference in OntoNotes
- Word Representations: A Simple and General Method for Semi-Supervised Learning
- The Proposition Bank: An Annotated Corpus of Semantic Roles
- Efficient Non-parametric Estimation of Multiple Embeddings per Word in Vector Space
- Semi-supervised Sequence Learning
- On the Properties of Neural Machine Translation: Encoder–Decoder Approaches
Cited by
- Enhancing Clinical Concept Extraction with Contextual Embedding
- Human and computer estimations of Predictability of words in written language
- Transition-Based Neural Word Segmentation
- Reinforced Mnemonic Reader for Machine Reading Comprehension
- Neural Machine Translation
- DisSent: Sentence Representation Learning from Explicit Discourse Relations
- Large-scale Cloze Test Dataset Designed by Teachers
- Stochastic Answer Networks for Machine Reading Comprehension
- Contextualized Word Representations for Reading Comprehension
- Fine-tuned Language Models for Text Classification
- Span-based Neural Structured Prediction
- Rare Feature Selection in High Dimensions
- Context is Everything: Finding Meaning Statistically in Semantic Spaces
- An Analysis of Neural Language Modeling at Multiple Scales
- AllenNLP: A Deep Semantic Natural Language Processing Platform
- Meta-Learning a Dynamical Language Model
- Simple and Effective Semi-Supervised Question Answering
- Higher-Order Coreference Resolution with Coarse-to-Fine Inference
- Utilizing Neural Networks and Linguistic Metadata for Early Detection of Depression Indications in Text Sequences
- Productivity, Portability, Performance: Data-Centric Python
Related papers
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
- Efficient Estimation of Word Representations in Vector Space
- Attention is All you Need