Glossary
Searchable terminology from accessibility, web standards, and related fields.
5 results found in automatic speech recognition.
- Character Error Rate (CER)
- A metric for evaluating automatic speech recognition (ASR) and optical character recognition (OCR) accuracy, measuring the minimum number of character-level edits (insertions, deletions, substitutions…
- Forced Alignment (Phonetic Alignment, Phone-Level Alignment)
- Forced alignment is an automatic speech processing technique that aligns a speech recording with its known transcription at the phoneme or word level. Unlike free speech recognition which determines t…
- Perceptual Linear Prediction (PLP, PLP Coefficients)
- Perceptual Linear Prediction (PLP) is an acoustic feature extraction technique used in speech processing that models human auditory perception. PLP analysis applies psychoacoustic principles including…
- Universal Background Model (UBM)
- A Universal Background Model (UBM) is a large Gaussian Mixture Model trained on speech from many speakers to represent speaker-independent acoustic characteristics. The UBM serves as a reference distr…
- Whisper (OpenAI Whisper, Whisper ASR)
- An open-source automatic speech recognition (ASR) model released by OpenAI in 2022, trained on 680,000 hours of multilingual and multitask supervised audio data. Whisper supports transcription in doze…