Glossary
Searchable terminology from accessibility, web standards, and related fields.
17 results found in speech processing.
- Accelerated Speech (Time-Compressed Speech, Speed-Up Speech)
- Audio output played at faster than normal speaking rate, commonly used by people with visual impairments when interacting with screen readers and other audio-based assistive technologies. Research sho…
- Computer-Assisted Language Learning (CALL, Computer-Aided Language Learning)
- Computer-Assisted Language Learning (CALL) refers to the use of computers and digital technology to support language education and pronunciation training. CALL systems often incorporate automatic spee…
- Deaf Speech (Deaf Accent, Deaf Voice)
- Accented speech produced by many individuals who are deaf or significantly hard of hearing, resulting from incomplete acoustic feedback from their own voices. Because deaf speakers cannot fully hear t…
- Forced Alignment (Phonetic Alignment, Phone-Level Alignment)
- Forced alignment is an automatic speech processing technique that aligns a speech recording with its known transcription at the phoneme or word level. Unlike free speech recognition which determines t…
- Formant (Vocal Formant, Formant Frequency)
- A concentration of acoustic energy around a particular frequency in the speech signal, produced by the resonance of the vocal tract. Formants are labeled sequentially (F1, F2, F3, etc.) from lowest to…
- Fundamental Frequency (F0, Pitch Frequency, Voice Pitch)
- The lowest frequency of a periodic sound wave, corresponding to the rate at which the vocal folds vibrate during voiced speech. Fundamental frequency (F0) is perceived by listeners as pitch and is a p…
- Gaussian Mixture Model (GMM)
- A Gaussian Mixture Model (GMM) is a probabilistic model that represents data as a weighted combination of multiple Gaussian (normal) distributions. Each component Gaussian has its own mean and covaria…
- Hyperarticulation (Clear Speech, Over-Articulation)
- A speaking style in which a person exaggerates the clarity of their pronunciation by moving their tongue and mouth to more extreme positions, producing more distinct vowel and consonant sounds. Hypera…
- Iterative Crowdsourcing (Iterative Human Computation, Multi-Round Crowdsourcing)
- A human computation workflow in which multiple rounds of crowd workers build iteratively upon each other's responses to collectively achieve higher quality results than any individual worker could pro…
- Perceptual Linear Prediction (PLP, PLP Coefficients)
- Perceptual Linear Prediction (PLP) is an acoustic feature extraction technique used in speech processing that models human auditory perception. PLP analysis applies psychoacoustic principles including…
- Speech Synthesis (Synthetic Speech, TTS Engine)
- The artificial production of human speech by computer, most commonly used in text-to-speech (TTS) systems that convert written text into spoken audio. Speech synthesis is foundational to screen reader…
- Stammering (Stuttering, Stammer, Stutter)
- A neurological condition that affects the rhythmic flow of speech, causing involuntary repetitions, prolongations, or blocks of sounds, syllables, or words. Blocking describes audible or silent moment…
- Supervector (GMM Supervector)
- A supervector is a high-dimensional feature representation created by concatenating the mean vectors from all components of a Gaussian Mixture Model (GMM) adapted to a specific speaker or utterance. T…
- UAspeech Database (UAspeech, UA-Speech, Universal Access Speech)
- The UAspeech Database is a standardized corpus of dysarthric speech recordings created for research in accessible speech technology. It contains isolated word recordings from speakers with cerebral pa…
- Universal Background Model (UBM)
- A Universal Background Model (UBM) is a large Gaussian Mixture Model trained on speech from many speakers to represent speaker-independent acoustic characteristics. The UBM serves as a reference distr…
- Voice Activity Detection (VAD, Speech Detection)
- A signal processing technique that automatically determines whether a segment of audio contains human speech or not. In accessibility applications, voice activity detection is used in audio descriptio…
- Voice Interface (Speech Interface, Voice User Interface, VUI, Conversational Interface)
- An interface that allows users to interact with a system using spoken natural language commands rather than keyboard, mouse, or touch input. Voice interfaces range from simple command-and-control syst…