Statistical Language Models within the Algebra of Weighted Rational Languages
Explore this paper's citation graph
Summary
This paper presents a novel way of creating N-gram language models using weighted finite automata and makes use of five special constant weighted transductions which rely only on the alphabet and the model parameter N.
- Type
- article
- Published
- 2009-02-01
- Cited by
- 4
- References
- 36
- Access
- Open access
- OpenAlex
- https://openalex.org/W32677852
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:19310263
Keywords
Alphabet, Computer science, Finite-state machine, Automaton, Implementation
References
- A weight pushing algorithm for large vocabulary speech recognition
- Semirings, Automata and Languages
- Transducer composition for context-dependent network expansion
- On a combinatorical problem
- Statistical methods for speech recognition
- Interpolated estimation of Markov source parameters from sparse data
- Semirings, Automata, Languages
- A Maximum Likelihood Approach to Continuous Speech Recognition
- Computational analysis of present-day American English
- Minimization algorithms for sequential transducers
- On transformations of formal power series
- Minimisation of Acyclic Deterministic Automata in Linear Time
- Simpler and More General Minimization for Weighted Finite-State Automata
- Prediction and Entropy of Printed English
- An Empirical Study of Smoothing Techniques for Language Modeling
- Efficient string matching
- The zero-frequency problem: Estimating the probabilities of novel events in adaptive text compression
- Semiring Frameworks and Algorithms for Shortest-Distance Problems
- Generalized Algorithms for Constructing Statistical Language Models
- Estimation of probabilities from sparse data for the language model component of a speech recognizer
Cited by
Related papers
- Context-Free Recognition with Weighted Automata
- Handbook of Weighted Automata
- Rational transductions and complexity of counting problems
- Automata Learning: An Algebraic Approach
- Developments in Language Theory, 10th International Conference, DLT 2006, Santa Barbara, CA, USA, June 26-29, 2006, Proceedings
- The Chomsky-SCHüTzenberger Theorem for Quantitative Context-Free Languages
- Exploring the Relationship between the Structural and the Actual Similarities of Automata