Unsupervised Discovery of Morphemes
Explore this paper's citation graph
Summary
Two methods for unsupervised segmentation of words into morpheme-like units are presented based on the Minimum Description Length (MDL) principle and Maximum Likelihood (ML) optimization is used.
- Type
- article
- Published
- 2002-05-21
- Cited by
- 422
- References
- 15
- Access
- Open access
- OpenAlex
- https://openalex.org/W2117621558
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:5133576
Keywords
Morpheme, Computer science, Artificial intelligence, Minimum description length, Segmentation
References
- Unsupervised Word Induction Using MDL Criterion
- Unsupervised Learning of Word Boundary with Description Length Gain
- Stochastic Complexity in Statistical Inquiry
- Redundancy Reduction as a Strategy for Unsupervised Learning
- Morphemes as Necessary Concept for Structures Discovery from Untagged Corpora
- An Efficient, Probabilistically Sound Algorithm for Segmentation and Word Discovery
- Inference of variable-length linguistic and acoustic units by multigrams
- A General Computational Model for Word-Form Recognition and Production
- Unsupervised Learning of the Morphology of a Natural Language
- Exploiting latent semantic information in statistical language modeling
- Linguistic Structure as Composition and Perturbation
- Knowledge-Free Induction of Morphology Using Latent Semantic Analysis
- Two-Level Morphology: A General Computational Model for Word-Form Recognition and Production
- The Unsupervised Acquisition of a Lexicon from Continuous Speech
- MASSACHUSETTS INSTITUTE OF TECHNOLOGY ARTIFICIAL INTELLIGENCE LABORATORY and CENTER FOR BIOLOGICAL AND COMPUTATIONAL LEARNING DEPARTMENT OF BRAIN AND COGNITIVE SCIENCES
Cited by
- On lexicon creation for turkish LVCSR
- Language models for automatic speech recognition : construction and complexity control
- Unsupervised segmentation of words into morphemes - morpho challenge 2005 application to automatic speech recognition
- Morfessor and Hutmegs: Unsupervised Morpheme Segmentation for Highly-Inflecting and Compounding Languages
- Selección de unidades léxicas para reconocimiento automático del habla continua en euskera
- Data driven subword unit modeling for speech recognition and its application to interactive reading tutors
- Unsupervised morphological analysis of small corpora: First experiments with Kilivila
- To recover from speech recognition errors in spoken document retrieval
- Apprentissage de connaissances morphologiques pour l'acquisition automatique de ressources lexicales. (Unsupervised learning of morphological knowledge for the automatic acquisition of lexical resources)
- A Markov model for the acquisition of morphological structure
- Morfessor and variKN machine learning tools for speech and language technology
- Unsupervised Learning of Morphology by using Syntactic Categories
- Comparison of ML, MAP, and VB based acoustic models in large vocabulary speech recognition
- Computational models relating properties of visual neurons to natural stimulus statistics
- Structures and distributions in morphology learning
- INDUCING THE MORPHOLOGICAL LEXICON OF A NATURAL LANGUAGE FROM UNANNOTATED TEXT
- Applying Morphological Decompositions to Statistical Machine Translation
- QuickAssist Extensive Reading for Learners of German Using CALL Technologies
- Morfessor 2.0: Python Implementation and Extensions for Morfessor Baseline
- Contributions à la description de la structure morphologique du lexique et à l'approche extensive en morphologie
Related papers
- Exceptions vs. Non-exceptions in Sound Changes: Morphological Condition and Frequency
- On the Shift of Non-morpheme to Morpheme
- The Multifunctional Morphemes in Kurdish Language the Morpheme Le as an example
- MORPHEME ANALYSIS OF ENGLISH LANGUAGE
- On Zhou Boqi’s View of Chinese Morphology
- A New Method for Narrowband Signal Source-Number Detection and DOA Estimation