PPDB: The Paraphrase Database
Explore this paper's citation graph
Summary
The 1.0 release of the paraphrase database, PPDB, contains over 220 million paraphrase pairs, consisting of 73 million phrasal and 8 million lexical paraphrases, as well as 140million paraphrase patterns, which capture many meaning-preserving syntactic transformations.
- Type
- article
- Published
- 2013-06-01
- Cited by
- 803
- References
- 36
- OpenAlex
- https://openalex.org/W2251044566
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:6067240
Keywords
Paraphrase, Natural language processing, Computer science, Artificial intelligence, Sentence
References
- MultiUN: A Multilingual Corpus from United Nation Documents
- From TreeBank to PropBank
- Europarl: A Parallel Corpus for Statistical Machine Translation
- Monolingual Machine Translation for Paraphrase Generation
- PATTY: A Taxonomy of Relational Patterns with Semantic Types
- Building a Large Annotated Corpus of English: The Penn Treebank
- The Theory of Parsing, Translation, and Compiling
- Monolingual Distributional Similarity for Text-to-Text Generation
- Reranking Bilingually Extracted Paraphrases Using Monolingual Distributional Similarity
- DIRT @SBT@discovery of inference rules from text
- Unsupervised Construction of Large Paraphrase Corpora: Exploiting Massively Parallel News Sources
- Web-based models for natural language processing
- TER-Plus: paraphrase, semantic, and alignment enhancements to Translation Edit Rate
- Sentence Compression Beyond Word Deletion
- Findings of the 2009 Workshop on Statistical Machine Translation
- Global Learning of Typed Entailment Rules
- Statistical Machine Translation for Query Expansion in Answer Retrieval
- ParaEval: Using Paraphrases to Evaluate Summaries Automatically
- BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network
- Large Scale Acquisition of Paraphrases for Learning Surface Patterns
Cited by
- Generating and simplifying sentences
- Multiview LSA: Representation Learning via Generalized CCA
- Large-Scale Paraphrasing for Natural Language Understanding
- A Unified Probabilistic Approach for Semantic Clustering of Relational Phrases
- Text Rewriting Improves Semantic Role Labeling
- Event structures in knowledge, pictures and text
- Discriminative methods for statistical spoken dialogue systems
- Knowledge-Based Textual Inference via Parse-Tree Transformations
- On the Proper Treatment of Quantifiers in Probabilistic Logic Semantics
- Encoding Prior Knowledge with Eigenword Embeddings
- Erratum: “From Paraphrase Database to Compositional Paraphrase Model and Back”
- Leveraging Paraphrase Labels to Extract Synonyms from Twitter
- Syntax-Aware Multi-Sense Word Embeddings for Deep Compositional Models of Meaning
- Cross level semantic similarity: an evaluation framework for universal measures of similarity
- Natural Language Semantics using Probabilistic Logic
- Efficient Global Learning of Entailment Graphs
- Document Layout Optimization with Automated Paraphrasing
- Linguistic steganography on Twitter: hierarchical language modeling with manual interaction
- Incorporating Linguistic Knowledge for Learning Distributed Word Representations
- Augmenting FrameNet Via PPDB
Related papers
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Efficient Estimation of Word Representations in Vector Space
- PPDB 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification
- SemEval-2012 Task 6: A Pilot on Semantic Textual Similarity
- Improving Lexical Embeddings with Semantic Knowledge
- GloVe: Global Vectors for Word Representation
- Distributed Representations of Words and Phrases and their Compositionality
- Paraphrasing with Bilingual Parallel Corpora