A Bilingual Generative Transformer for Semantic Sentence Embedding
Explore this paper's citation graph
Summary
A deep latent variable model is proposed that attempts to perform source separation on parallel sentences, isolating what they have in common in a latent semantic vector, and explaining what is left over with language-specific latent vectors.
- Type
- preprint
- Published
- 2019-11-10
- Cited by
- 34
- References
- 51
- Access
- Open access
- OpenAlex
- https://openalex.org/W2988944673
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:207853396
Keywords
Natural language processing, Computer science, Artificial intelligence, Sentence, Semantic similarity
References
- A large annotated corpus for learning natural language inference
- Auto-Encoding Variational Bayes
- Unsupervised Construction of Large Paraphrase Corpora: Exploiting Massively Parallel News Sources
- Long Short-Term Memory
- SemEval-2014 Task 10: Multilingual Semantic Textual Similarity
- SemEval-2015 Task 2: Semantic Textual Similarity, English, Spanish and Pilot on Interpretability
- *SEM 2013 shared task: Semantic Textual Similarity
- Distributed Representations of Words and Phrases and their Compositionality
- Towards Universal Paraphrastic Sentence Embeddings
- Rethinking the Inception Architecture for Computer Vision
- GloVe: Global Vectors for Word Representation
- PPDB: The Paraphrase Database
- SemEval-2012 Task 6: A Pilot on Semantic Textual Similarity
- Learning Distributed Representations of Sentences from Unlabelled Data
- SemEval-2016 Task 1: Semantic Textual Similarity, Monolingual and Cross-Lingual Evaluation
- Charagram: Embedding Words and Sentences via Character n-grams
- Learning Joint Multilingual Sentence Representations with Neural Machine Translation
- An Empirical Analysis of NMT-Derived Interlingual Embeddings and Their Use in Parallel Sentence Identification
- Revisiting Recurrent Networks for Paraphrastic Sentence Embeddings
- Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
Cited by
- Self-training Improves Pre-training for Natural Language Understanding
- Controllable Paraphrasing and Translation with a Syntactic Exemplar
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- Paraphrastic Representations at Scale
- Joint Intent Detection and Slot Filling Based on Continual Learning Model
- Disentangling Semantics and Syntax in Sentence Embeddings with Pre-trained Language Models
- ConvFiT: Conversational Fine-Tuning of Pretrained Language Models
- Exploiting Twitter as Source of Large Corpora of Weakly Similar Pairs for Semantic Sentence Embeddings
- Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity Estimation
- Distilling Word Meaning in Context from Pre-trained Language Models
- QA Is the New KR: Question-Answer Pairs as Knowledge Bases
- Learning High-Order Semantic Representation for Intent Classification and Slot Filling on Low-Resource Language via Hypergraph
- MCSE: Multimodal Contrastive Learning of Sentence Embeddings
- Beyond Contrastive Learning: A Variational Generative Model for Multilingual Retrieval
- Retrofitting Multilingual Sentence Embeddings with Abstract Meaning Representation
- A Comprehensive Survey of Sentence Representations: From the BERT Epoch to the CHATGPT Era and Beyond
- Fine-grained Multi-lingual Disentangled Autoencoder for Language-agnostic Representation Learning
- The Daunting Dilemma with Sentence Encoders: Success on Standard Benchmarks, Failure in Capturing Basic Semantic Properties
- Language Models are Universal Embedders
- Enhancing Sentence Representation with Visually-supervised Multimodal Pre-training
Related papers
- Modeling Sentences in the Latent Space
- Contrasting distinct structured views to learn sentence embeddings
- Neural embeddings: accurate and readable inferences based on semantic kernels
- Generating Contradictory, Neutral, and Entailing Sentences
- A Joint Model for Sentence Semantic Similarity Learning
- Meta-Embedding Sentence Representation for Textual Similarity
- R2-Net: Relation of Relation Learning Network for Sentence Semantic Matching