Glossary

Searchable terminology from accessibility, web standards, and related fields.

62 results found in Speech Technology.

Grid-Based Navigation (Grid Navigation, Grid Cursor Control)
A speech-controlled cursor positioning technique that divides the screen into numbered regions, allowing users to select progressively smaller areas by speaking numbers until the cursor reaches the ta…
JSML (Java Speech Markup Language)
An XML-based markup language developed by Sun Microsystems that provides directives for controlling the output of speech synthesis engines. JSML allows developers to specify pronunciation details incl…
Landmark Detection (Acoustic Landmark Detection, Stevens Landmark Theory)
Landmark detection is a speech analysis method based on Kenneth Stevens' acoustic model of speech production, which identifies perceptually significant points in the acoustic signal where listeners ex…
Listening Window
The interval during which a voice assistant or speech-recognition system actively captures user audio after being activated (by wake word or button press). A short or fixed listening window causes pre…
Math-to-Speech (Mathematical Speech Generation, Math Speech)
The process of converting mathematical notation into spoken language that can be rendered by text-to-speech engines or read aloud by screen readers. Math-to-speech is significantly more complex than r…
Mispronunciation Detection (Pronunciation Error Detection, Mispronunciation Diagnosis)
Mispronunciation detection is the automated process of identifying errors in a speaker's pronunciation by comparing their speech production against a target or expected utterance. In assistive technol…
Natural Speech Output (Recorded Speech, Digitized Speech)
Speech output produced from digital recordings of actual human speakers, as opposed to artificially generated synthetic speech. Natural speech output preserves the prosody, intonation, emotion, and vo…
Neural Vocoder
A deep-learning model that synthesises audio waveforms from intermediate acoustic representations such as mel-spectrograms or discrete speech units. Examples include HiFi-GAN, WaveNet, WaveGlow, and S…
Non-Verbal Vocalization (Non-Speech Vocalization, Vocal Gesture, Non-speech Vocalisation, Non-Speech Vocalisation)
A sound produced by the voice that is not a spoken word, such as a sustained vowel sound ("Ahhhhh"), hum, or other vocal noise. In assistive technology and alternative input contexts, non-verbal vocal…
Perceptual Linear Prediction (PLP, PLP Coefficients)
Perceptual Linear Prediction (PLP) is an acoustic feature extraction technique used in speech processing that models human auditory perception. PLP analysis applies psychoacoustic principles including…
Re-speaking (Respeaking, Speech-to-Text Relay)
A captioning technique in which a trained operator listens to a speaker and repeats (re-speaks) their words clearly into a high-quality microphone in a controlled environment, allowing automatic speec…
Recurrent Neural Network (RNN)
A recurrent neural network (RNN) is a type of artificial neural network designed to process sequential data by maintaining an internal state (memory) that captures information from previous inputs in …
Repair Mechanism (Conversational Repair)
In conversational interface design, a feature that helps the user and the system recover from misrecognition, ambiguity, or misunderstanding — for example, clarification prompts ("Did you mean the [X]…
SAPI (Speech Application Programming Interface, Microsoft SAPI)
The Speech Application Programming Interface (SAPI) is a Microsoft Windows API that enables applications to use speech recognition and text-to-speech synthesis. SAPI provides a standardized interface …
Semantically Unpredictable Sentences (SUS, SUS Test)
A standardised method for evaluating speech intelligibility in which listeners are presented with sentences that are grammatically correct but semantically meaningless, such as "A polite art jumps ben…
Speaker Adaptation (Voice Adaptation, Speaker-Adaptive Training, Voice Personalization)
Speaker adaptation is the process of adjusting an existing automatic speech recognition (ASR) system — usually one trained on a large, demographically broad corpus of able-bodied speakers — to a parti…
Speaker Diarisation (Speaker Diarization, Speaker Segmentation)
The automatic process of segmenting an audio recording by speaker identity — answering "who spoke when" — and labelling each segment. A critical pre-requisite for accessible transcripts of multi-voice…
Spectrogram (Sonogram, Spectral Display)
A spectrogram is a visual representation of the frequency spectrum of a signal as it varies over time, typically showing time on the horizontal axis, frequency on the vertical axis, and intensity repr…
Speech Composer (Speech Generation, Message Composition Engine)
A software component in AAC (Augmentative and Alternative Communication) systems that takes user input — whether typed text, selected symbols, or telegraphic phrases — and processes it for spoken outp…
Speech Diversity (Diverse Speech, Non-Typical Speech)
The full range of ways human speech varies from the narrow 'typical' speech on which most speech-AI systems are trained and benchmarked. Speech diversity includes people who stutter, d/Deaf and Hard-o…