Unsupervised Construction of Large Paraphrase Corpora: Exploiting Massively Parallel News Sources

Explore this paper's citation graph

Summary

Investigation of unsupervised techniques for acquiring monolingual sentence-level paraphrases from a corpus of temporally and topically clustered news articles collected from thousands of web-based news sources shows that edit distance data is cleaner and more easily-aligned than the heuristic data.

Type
article
Published
2004-08-23
Cited by
908
References
21
Access
Open access

Keywords

Paraphrase, Computer science, Natural language processing, Artificial intelligence, Sentence

References

Cited by

Related papers