Glossary
Searchable terminology from accessibility, web standards, and related fields.
62 results found in Speech Technology.
- Speech Language Model (SLM, Audio Language Model, Speech Foundation Model)
- A class of large neural models that processes both speech and text in a single end-to-end framework, integrating tasks — automatic speech recognition, spoken language understanding, dialogue, speech g…
- Speech Neuroprosthesis (Speech BCI, Speech Brain-Computer Interface)
- A brain-computer interface that decodes neural activity associated with attempted or imagined speech and converts it into text, synthesized voice, or both. Speech neuroprostheses are designed for peop…
- Speech Prosodics (Prosodic Features, Suprasegmental Features)
- Speech prosodics refers to the nonverbal acoustic features of speech that convey meaning beyond the words themselves, including pitch (fundamental frequency), rhythm, stress, intonation patterns, paus…
- Speech Rate (Speaking Rate, Articulation Rate)
- The speed at which speech is produced, typically measured in words per minute (WPM) or syllables per second. Normal conversational speech ranges from 120-180 WPM, while screen reader users often confi…
- Speech Visualization (Visual Speech Display, Speech-to-Visual Display)
- Speech visualization refers to techniques that convert spoken language into visual representations to aid comprehension, particularly for individuals who are deaf or hard of hearing. These displays ca…
- Speech-Generating Device (SGD, Voice Output Communication Aid, VOCA)
- An electronic AAC device that produces spoken output from text or symbol input, enabling people with speech disabilities to communicate verbally with others. Speech-generating devices range from dedic…
- Speech-to-Speech (S2S, Speech-to-Speech Conversion)
- A class of systems that transform one speech signal directly into another — for example, converting atypical input (whispered, dysarthric, accented, or cross-lingual speech) into clear, intelligible o…
- Spoken Dialogue System (SDS, Voice Dialogue System)
- A computer system that communicates with users through spoken natural language, allowing them to interact via voice rather than visual or manual interfaces. Spoken dialogue systems are used in telecar…
- Supervector (GMM Supervector)
- A supervector is a high-dimensional feature representation created by concatenating the mean vectors from all components of a Gaussian Mixture Model (GMM) adapted to a specific speaker or utterance. T…
- Synthesized Speech (Synthetic Speech, Speech Synthesis, TTS Output)
- Computer-generated speech produced by text-to-speech (TTS) engines that convert written text into spoken audio output. Synthesized speech is the primary means by which screen readers convey on-screen …
- Synthesized Video Description (TTS Video Description, Text-to-Speech Description, Synthesized Audio Description)
- An audio description for video content that is generated using text-to-speech (TTS) technology rather than recorded by a human narrator. A describer writes a text script describing the visual elements…
- Synthetic Speech (Artificial Speech, Computer-generated Speech)
- Speech that is artificially produced by computer systems rather than recorded from human speakers. Synthetic speech is the output of text-to-speech systems and is fundamental to screen readers and voi…
- Talking Head (Virtual Talking Head, Animated Face, 3D Talking Head, Virtual Speech Tutor)
- A talking head is a computer-generated 3D or 2D animated representation of a human face and articulatory system that produces visible speech movements synchronised with audio output. In accessibility …
- Text-to-Speech (TTS, Speech Synthesis)
- Technology that converts written text into spoken audio output. Text-to-speech is a fundamental component of many assistive technologies, including screen readers, audio description tools, and communi…
- Time-compressed Speech (Accelerated Speech, Speed-altered Speech)
- Speech that has been digitally processed to play at a faster rate than it was originally recorded or synthesized, while preserving pitch. Unlike simply increasing playback speed (which raises pitch), …
- Unit Selection Synthesis (Concatenative Unit Selection, Unit Selection TTS)
- A text-to-speech synthesis approach that generates speech by selecting and concatenating variable-length segments of pre-recorded human speech from a large database to match the input text. Unit selec…
- Universal Background Model (UBM)
- A Universal Background Model (UBM) is a large Gaussian Mixture Model trained on speech from many speakers to represent speaker-independent acoustic characteristics. The UBM serves as a reference distr…
- Visual Feedback (Visual Biofeedback)
- A method of providing real-time visual information to a user about their actions, performance, or physiological state. In speech therapy and assistive technology, visual feedback systems display graph…
- Visual Speech Aid (Speech Reading Aid, Visual Communication Aid)
- A visual speech aid is an assistive device or system that converts auditory speech information into visual form to help individuals with hearing impairments follow spoken conversation. These aids may …
- Vocalization Analysis (Vocal Analysis, Infant Vocalization Analysis)
- Vocalization analysis is the systematic study and measurement of vocal productions, including speech, pre-speech sounds, and non-speech vocalizations. In developmental and clinical contexts, vocalizat…