Class-Based n-gram Models of Natural Language
Explore this paper's citation graph
Summary
This work addresses the problem of predicting a word from previous words in a sample of text and discusses n-gram models based on classes of words, finding that these models are able to extract classes that have the flavor of either syntactically based groupings or semanticallybased groupings, depending on the nature of the underlying statistics.
- Type
- article
- Published
- 1992-12-01
- Cited by
- 3,681
- References
- 14
- OpenAlex
- https://openalex.org/W2121227244
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:10986188
Keywords
n-gram, Computer science, Class (philosophy), Natural language processing, Word (group theory)
References
- An inequality and associated maximization technique in statistical estimation of probabilistic functions of a Markov process
- Interpolated estimation of Markov source parameters from sparse data
- Experiments with the Tangora 20,000 word speech recognizer
- A Maximum Likelihood Approach to Continuous Speech Recognition
- Information Theory and Reliable Communication
- Computational analysis of present-day American English
- Context based spelling correction
- Maximum likelihood from incomplete data via the EM - algorithm plus discussions on the paper
- An Introduction to Probability Theory and Its Applications, Volume II
- THE POPULATION FREQUENCIES OF SPECIES AND THE ESTIMATION OF POPULATION PARAMETERS
- A Statistical Approach to Machine Translation
- Information Theory and Reliable Communication
- An Introduction to Probability Theory and its Applications, Volume I
- An Introduction to Probability Theory and Its Applications
- An introduction to probability theory and its applications
Cited by
- Knowledge Acquisition for Knowledge Management
- Unsupervised Part of Speech Tagging Supporting Supervised Methods
- Language models for automatic speech recognition : construction and complexity control
- Language Models With Meta-information
- Within and across sentence boundary language model
- The Detection of Fraudulent Financial Statements: an Integrated Language Model
- Fast hierarchical grammar optimization algorithm toward time and space efficiency
- Class-based variable memory length Markov model
- Method for Improving Automatic Word Categorization
- On the use of morphological constraints in n-gram statistical language model
- Informations morpho-syntaxiques et adaptation thématique pour améliorer la reconnaissance de la parole
- Transducer-based speech recognition with dynamic language models
- Augmented context features for Arabic speech recognition
- Language model adaptation using word clustering
- Towards a Bootstrapping Framework for Corpus Semantic Tagging
- MSRA-USTC-SJTU at TRECVID 2007: High-Level Feature Extraction and Search
- Improving Pronunciation Accuracy of Proper Names with Language Origin Classes
- De-identification of clinical notes in French: towards a protocol for reference corpus development
- Comparing and Combining Generative and Posterior Probability Models: Some Advances in Sentence Boundary Detection in Speech
- "Rate My Therapist": Automated Detection of Empathy in Drug and Alcohol Counseling via Speech and Language Processing
Related papers
- GloVe: Global Vectors for Word Representation
- Natural Language Processing (Almost) from Scratch
- An empirical study of smoothing techniques for language modeling
- Word Representations: A Simple and General Method for Semi-Supervised Learning
- Semi-Supervised Learning for Natural Language
- Distributed Representations of Words and Phrases and their Compositionality
- Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
- Indexing by Latent Semantic Analysis
- Estimation of probabilities from sparse data for the language model component of a speech recognizer
- A neural probabilistic language model