Automatic Language Identification in Texts: A Survey
Explore this paper's citation graph
Summary
A unified notation is introduced for evaluation methods, applications, as well as off-the-shelf LI systems that do not require training by the end user, to propose future directions for research in LI.
- Type
- preprint
- Published
- 2018-04-22
- Cited by
- 235
- References
- 468
- Access
- Open access
- OpenAlex
- https://openalex.org/W2799012885
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5071247
Keywords
Computer science, Identification (biology), Key (lock), Notation, Natural language processing
References
- Language identification of person names using CF-IOF based weighing function
- Processing highly variant language using incremental model selection
- Stacked generalization
- On text-based language identification for multilingual speech recognition systems
- Subsegmental language detection in Celtic language text
- A tagging algorithm for mixed language identification in a noisy domain
- Text classification and segmentation using minimum cross-entropy
- Language Identification of Short Text Segments with N-gram Models
- Discriminative vs Informative Learning
- The Problems of Language Identification within Hugely Multilingual Data Sets
- Language identification of names with SVMs
- Multilingual document language recognition for creating corpora
- Natural Language Identification using Corpus-Based Models
- Diction for phoneme/syllable/word-category and identification of language using HMM
- Multl-Language Text Indexing for Internet Retrieval
- Special Invited Paper-Additive logistic regression: A statistical view of boosting
- Reconsidering Language Identification for Written Language Resources
- Language identification and IT Addressing problems of linguistic diversity on a global scale
- Accurate Arabic Script Language/Dialect Classification
- Merging Comparable Data Sources for the Discrimination of Similar Languages : The DSL Corpus Collection
Cited by
- Discriminating between Indo-Aryan Languages Using SVM Ensembles
- Classifier Ensembles for Dialect and Language Variety Identification
- Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign
- Neural Machine Translation into Language Varieties
- Computationally efficient discrimination between language varieties with large feature vectors and regularized classifiers
- HeLI-based Experiments in Discriminating Between Dutch and Flemish Subtitles
- Deep Models for Arabic Dialect Identification on Benchmarked Data
- When Simple n-gram Models Outperform Syntactic Approaches: Discriminating between Dutch and Flemish
- HeLI-based Experiments in Swiss German Dialect Identification
- Iterative Language Model Adaptation for Indo-Aryan Language Identification
- Language and Dialect Identification of Cuneiform Texts
- Language model adaptation for language and dialect identification of text
- Experiments in Cuneiform Language Identification
- Language identification in texts
- Semi-supervised Stochastic Multi-Domain Learning using Variational Inference
- Wetin dey with these comments? Modeling Sociolinguistic Factors Affecting Code-switching Behavior in Nigerian Online Discussions
- Document and Word-level Language Identification for Noisy User Generated Text
- Discriminating between Mandarin Chinese and Swiss-German varieties using adaptive language models
- Language Discrimination and Transfer Learning for Similar Languages: Experiments with Feature Combinations and Adaptation
- Improving Cuneiform Language Identification with BERT
Related papers
- Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task
- A Report on the DSL Shared Task 2014
- A Survey of Language Identification Techniques and Applications
- A Systematic Mapping Study of Language Features Identification from Large Text Collection
- Evaluation of Natural Language Processors
- Structure Discovery in Natural Language
- Enhancing Text Retrieval by Using Advanced Stylistic Techniques
- Language Model Techniques in Machine Translation