Processing highly variant language using incremental model selection
Explore this paper's citation graph
Summary
This dissertation provides an architecture for NLP that allows for better handling of complicated language variation and finds that segmenting language before tagging, and then applying single-language homogeneous language models, is competitive to multilingual heterogeneous tagging models.
- Type
- article
- Published
- 2012-01-01
- Cited by
- 9
- References
- 93
- OpenAlex
- https://openalex.org/W25263363
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:59702535
Keywords
Computer science, Artificial intelligence, Speech recognition, Natural language processing, Pattern recognition (psychology)
References
- A tagging algorithm for mixed language identification in a noisy domain
- Language Identification of Short Text Segments with N-gram Models
- Morfessor and variKN machine learning tools for speech and language technology
- A Concurrent Validity Study of the Raygor Readability Estimate.
- Introduction to the CoNLL-2000 Shared Task Chunking
- Error Correction for Arabic Dictionary Lookup
- Detecting Fake Content with Relative Entropy Scoring
- Language model selection based on the analysis of Japanese spontaneous speech on travel arrangement task
- Jewish and Muslim Dialects of Moroccan Arabic
- SMOG Grading - A New Readability Formula.
- Learning String-Edit Distance
- Language identification: a solved problem suitable for undergraduate instruction
- N-gram-based text categorization
- POS Tagging for German: how important is the Right Context?
- Text Categorization with Support Vector Machines: Learning with Many Relevant Features
- Adaptive Statistical Language Modeling; A Maximum Entropy Approach
- Authorship Attribution with Support Vector Machines
- Building a Large Annotated Corpus of English: The Penn Treebank
- Using Literal and Grammatical Statistics for Authorship Attribution
- Parsers In Tutors
Cited by
- Part of Speech Tagging Bilingual Speech Transcripts with Intrasentential Model Switching
- Code-Mixing in Social Media Text
- Identifying Languages at the Word Level in Code-Mixed Indian Social Media Text
- Simple Tools for Exploring Variation in Code-switching for Linguists
- Evaluation of language identification methods using 285 languages
- Automatic Language Identification in Texts: A Survey
- Language identification in texts
- Italian Language and Dialect Identification and Regional French Variety Detection using Adaptive Naive Bayes
- Tuning HeLI-OTS for Guarani-Spanish Code Switching Analysis
Related papers
- Language identification of individualwords in a multilingual automatic speech recognition system
- An Arabic/English Switch for Audio Indexing and Dialogue Management
- A Hybrid Neural Network for Language Identification from Text
- Support Vector Machines based Part of Speech Tagging for Nepali Text
- Identification of related languages from spoken data: Moving from off-line to on-line scenario
- Document driven machine translation enhanced ASR
- Multi-Dialect Arabic Speech Recognition
- Architectures for Speech-to-Speech Translation Using Finite-state Models
- Similarity Based Language Model Construction for Voice Activated Open-Domain Question Answering