Glossary

Searchable terminology from accessibility, web standards, and related fields.

62 results found in Speech Technology.

Speech Language Model (SLM, Audio Language Model, Speech Foundation Model)
A class of large neural models that processes both speech and text in a single end-to-end framework, integrating tasks — automatic speech recognition, spoken language understanding, dialogue, speech g…
Speech Neuroprosthesis (Speech BCI, Speech Brain-Computer Interface)
A brain-computer interface that decodes neural activity associated with attempted or imagined speech and converts it into text, synthesized voice, or both. Speech neuroprostheses are designed for peop…
Speech Prosodics (Prosodic Features, Suprasegmental Features)
Speech prosodics refers to the nonverbal acoustic features of speech that convey meaning beyond the words themselves, including pitch (fundamental frequency), rhythm, stress, intonation patterns, paus…
Speech Rate (Speaking Rate, Articulation Rate)
The speed at which speech is produced, typically measured in words per minute (WPM) or syllables per second. Normal conversational speech ranges from 120-180 WPM, while screen reader users often confi…
Speech Visualization (Visual Speech Display, Speech-to-Visual Display)
Speech visualization refers to techniques that convert spoken language into visual representations to aid comprehension, particularly for individuals who are deaf or hard of hearing. These displays ca…
Speech-Generating Device (SGD, Voice Output Communication Aid, VOCA)
An electronic AAC device that produces spoken output from text or symbol input, enabling people with speech disabilities to communicate verbally with others. Speech-generating devices range from dedic…
Speech-to-Speech (S2S, Speech-to-Speech Conversion)
A class of systems that transform one speech signal directly into another — for example, converting atypical input (whispered, dysarthric, accented, or cross-lingual speech) into clear, intelligible o…
Spoken Dialogue System (SDS, Voice Dialogue System)
A computer system that communicates with users through spoken natural language, allowing them to interact via voice rather than visual or manual interfaces. Spoken dialogue systems are used in telecar…
Supervector (GMM Supervector)
A supervector is a high-dimensional feature representation created by concatenating the mean vectors from all components of a Gaussian Mixture Model (GMM) adapted to a specific speaker or utterance. T…
Synthesized Speech (Synthetic Speech, Speech Synthesis, TTS Output)
Computer-generated speech produced by text-to-speech (TTS) engines that convert written text into spoken audio output. Synthesized speech is the primary means by which screen readers convey on-screen …
Synthesized Video Description (TTS Video Description, Text-to-Speech Description, Synthesized Audio Description)
An audio description for video content that is generated using text-to-speech (TTS) technology rather than recorded by a human narrator. A describer writes a text script describing the visual elements…
Synthetic Speech (Artificial Speech, Computer-generated Speech)
Speech that is artificially produced by computer systems rather than recorded from human speakers. Synthetic speech is the output of text-to-speech systems and is fundamental to screen readers and voi…
Talking Head (Virtual Talking Head, Animated Face, 3D Talking Head, Virtual Speech Tutor)
A talking head is a computer-generated 3D or 2D animated representation of a human face and articulatory system that produces visible speech movements synchronised with audio output. In accessibility …
Text-to-Speech (TTS, Speech Synthesis)
Technology that converts written text into spoken audio output. Text-to-speech is a fundamental component of many assistive technologies, including screen readers, audio description tools, and communi…
Time-compressed Speech (Accelerated Speech, Speed-altered Speech)
Speech that has been digitally processed to play at a faster rate than it was originally recorded or synthesized, while preserving pitch. Unlike simply increasing playback speed (which raises pitch), …
Unit Selection Synthesis (Concatenative Unit Selection, Unit Selection TTS)
A text-to-speech synthesis approach that generates speech by selecting and concatenating variable-length segments of pre-recorded human speech from a large database to match the input text. Unit selec…
Universal Background Model (UBM)
A Universal Background Model (UBM) is a large Gaussian Mixture Model trained on speech from many speakers to represent speaker-independent acoustic characteristics. The UBM serves as a reference distr…
Visual Feedback (Visual Biofeedback)
A method of providing real-time visual information to a user about their actions, performance, or physiological state. In speech therapy and assistive technology, visual feedback systems display graph…
Visual Speech Aid (Speech Reading Aid, Visual Communication Aid)
A visual speech aid is an assistive device or system that converts auditory speech information into visual form to help individuals with hearing impairments follow spoken conversation. These aids may …
Vocalization Analysis (Vocal Analysis, Infant Vocalization Analysis)
Vocalization analysis is the systematic study and measurement of vocal productions, including speech, pre-speech sounds, and non-speech vocalizations. In developmental and clinical contexts, vocalizat…