A tagging algorithm for mixed language identification in a noisy domain
Explore this paper's citation graph
Summary
A language neutral language identification approach capable of handling the characteristics of the domain in a robust fashion is discussed.
- Type
- article
- Published
- 2007-08-27
- Cited by
- 31
- References
- 7
- OpenAlex
- https://openalex.org/W74783503
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:20007902
Keywords
Computer science, Identification (biology), Language identification, Domain (mathematical analysis), Natural language processing
References
- Comparing methods for language identification
- Acquaintance: Language-Independent Document Categorization by N-Grams
- N-gram-based text categorization
- Text to Speech Technologies for Mobile Telephony Services
- Book Reviews: Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition
- Speech and language processing: an introduction to natural language processing
- Language Identifier: A Computer Program for Automatic Natural-Language Identification of On-line Tex
Cited by
- Processing highly variant language using incremental model selection
- Learning to Predict Code-Switching Points
- Part-of-Speech Tagging for English-Spanish Code-Switched Text
- Code Mixing: A Challenge for Language Identification in the Language of Social Media
- Foreign Words and the Automatic Processing of Arabic Social Media Text Written in Roman Script
- DCU-UVT: Word-Level Language Classification with Code-Mixed Data
- Code-Mixing in Social Media Text
- Identifying Languages at the Word Level in Code-Mixed Indian Social Media Text
- Demographic Dialectal Variation in Social Media: A Case Study of African-American English
- Applying corpus and computational methods to loanword research : new approaches to Anglicisms in Spanish
- Deep Learning-Based Language Identification in English-Hindi-Bengali Code-Mixed Social Media Corpora
- Automatic Language Identification in Texts: A Survey
- A Dataset for Building Code-Mixed Goal Oriented Conversation Systems
- Automatic Target Recovery for Hindi-English Code Mixed Puns
- Joint Part-of-Speech and Language ID Tagging for Code-Switched Data
- Code-switching in Irish tweets: A preliminary analysis
- Identifying and Modeling Code-Switched Language
- Annotating for Hate Speech: The MaNeCo Corpus and Some Input from Critical Discourse Analysis
- Using automated methods to explore the social stratification of anglicisms in Spanish
- Language Lexicons for Hindi-English Multilingual Text Processing
Related papers
- Remarks on Algorithm 2, Algorithm 3, Algorithm 15, Algorithm 25 and Algorithm 26
- An Anatomization Of Language Detection And Translation Using NLP Techniques
- An Anatomization of Language Detection and Translation using NLP Techniques
- Language model adaptation for language and dialect identification of text
- Language Identification for Multilingual Machine Translation
- Review of language identification techniques
- Text independent root word identification in Hindi language using natural language processing
- Weaponising AI for Natural Language Processing: Novel Perspectives