Glossary

Searchable terminology from accessibility, web standards, and related fields.

5 results found in automatic speech recognition.

Character Error Rate (CER)
A metric for evaluating automatic speech recognition (ASR) and optical character recognition (OCR) accuracy, measuring the minimum number of character-level edits (insertions, deletions, substitutions…
Forced Alignment (Phonetic Alignment, Phone-Level Alignment)
Forced alignment is an automatic speech processing technique that aligns a speech recording with its known transcription at the phoneme or word level. Unlike free speech recognition which determines t…
Perceptual Linear Prediction (PLP, PLP Coefficients)
Perceptual Linear Prediction (PLP) is an acoustic feature extraction technique used in speech processing that models human auditory perception. PLP analysis applies psychoacoustic principles including…
Universal Background Model (UBM)
A Universal Background Model (UBM) is a large Gaussian Mixture Model trained on speech from many speakers to represent speaker-independent acoustic characteristics. The UBM serves as a reference distr…
Whisper (OpenAI Whisper, Whisper ASR)
An open-source automatic speech recognition (ASR) model released by OpenAI in 2022, trained on 680,000 hours of multilingual and multitask supervised audio data. Whisper supports transcription in doze…