Glossary

Searchable terminology from accessibility, web standards, and related fields.

17 results found in speech processing.

Accelerated Speech (Time-Compressed Speech, Speed-Up Speech)
Audio output played at faster than normal speaking rate, commonly used by people with visual impairments when interacting with screen readers and other audio-based assistive technologies. Research sho…
Computer-Assisted Language Learning (CALL, Computer-Aided Language Learning)
Computer-Assisted Language Learning (CALL) refers to the use of computers and digital technology to support language education and pronunciation training. CALL systems often incorporate automatic spee…
Deaf Speech (Deaf Accent, Deaf Voice)
Accented speech produced by many individuals who are deaf or significantly hard of hearing, resulting from incomplete acoustic feedback from their own voices. Because deaf speakers cannot fully hear t…
Forced Alignment (Phonetic Alignment, Phone-Level Alignment)
Forced alignment is an automatic speech processing technique that aligns a speech recording with its known transcription at the phoneme or word level. Unlike free speech recognition which determines t…
Formant (Vocal Formant, Formant Frequency)
A concentration of acoustic energy around a particular frequency in the speech signal, produced by the resonance of the vocal tract. Formants are labeled sequentially (F1, F2, F3, etc.) from lowest to…
Fundamental Frequency (F0, Pitch Frequency, Voice Pitch)
The lowest frequency of a periodic sound wave, corresponding to the rate at which the vocal folds vibrate during voiced speech. Fundamental frequency (F0) is perceived by listeners as pitch and is a p…
Gaussian Mixture Model (GMM)
A Gaussian Mixture Model (GMM) is a probabilistic model that represents data as a weighted combination of multiple Gaussian (normal) distributions. Each component Gaussian has its own mean and covaria…
Hyperarticulation (Clear Speech, Over-Articulation)
A speaking style in which a person exaggerates the clarity of their pronunciation by moving their tongue and mouth to more extreme positions, producing more distinct vowel and consonant sounds. Hypera…
Iterative Crowdsourcing (Iterative Human Computation, Multi-Round Crowdsourcing)
A human computation workflow in which multiple rounds of crowd workers build iteratively upon each other's responses to collectively achieve higher quality results than any individual worker could pro…
Perceptual Linear Prediction (PLP, PLP Coefficients)
Perceptual Linear Prediction (PLP) is an acoustic feature extraction technique used in speech processing that models human auditory perception. PLP analysis applies psychoacoustic principles including…
Speech Synthesis (Synthetic Speech, TTS Engine)
The artificial production of human speech by computer, most commonly used in text-to-speech (TTS) systems that convert written text into spoken audio. Speech synthesis is foundational to screen reader…
Stammering (Stuttering, Stammer, Stutter)
A neurological condition that affects the rhythmic flow of speech, causing involuntary repetitions, prolongations, or blocks of sounds, syllables, or words. Blocking describes audible or silent moment…
Supervector (GMM Supervector)
A supervector is a high-dimensional feature representation created by concatenating the mean vectors from all components of a Gaussian Mixture Model (GMM) adapted to a specific speaker or utterance. T…
UAspeech Database (UAspeech, UA-Speech, Universal Access Speech)
The UAspeech Database is a standardized corpus of dysarthric speech recordings created for research in accessible speech technology. It contains isolated word recordings from speakers with cerebral pa…
Universal Background Model (UBM)
A Universal Background Model (UBM) is a large Gaussian Mixture Model trained on speech from many speakers to represent speaker-independent acoustic characteristics. The UBM serves as a reference distr…
Voice Activity Detection (VAD, Speech Detection)
A signal processing technique that automatically determines whether a segment of audio contains human speech or not. In accessibility applications, voice activity detection is used in audio descriptio…
Voice Interface (Speech Interface, Voice User Interface, VUI, Conversational Interface)
An interface that allows users to interact with a system using spoken natural language commands rather than keyboard, mouse, or touch input. Voice interfaces range from simple command-and-control syst…