Glossary

Searchable terminology from accessibility, web standards, and related fields.

29 results found in Audio.

Speaker Diarisation (Speaker Diarization, Speaker Segmentation)
The automatic process of segmenting an audio recording by speaker identity — answering "who spoke when" — and labelling each segment. A critical pre-requisite for accessible transcripts of multi-voice…
Speech Gap (Dialogue Gap, Audio Gap)
A pause or silence between spoken dialogue in a video or film where audio descriptions can be inserted without overlapping with the original soundtrack. Identifying speech gaps is a critical first ste…
Structured Audio (Structured Digital Audio)
Structured audio refers to digital audio content that has been encoded with hierarchical markers and metadata, allowing non-sequential access to specific segments such as chapters, sections, paragraph…
Synthesized Video Description (TTS Video Description, Text-to-Speech Description, Synthesized Audio Description)
An audio description for video content that is generated using text-to-speech (TTS) technology rather than recorded by a human narrator. A describer writes a text script describing the visual elements…
Tempo (BPM, Beats Per Minute)
The speed or pace of a musical piece, typically measured in beats per minute (BPM). Tempo is one of the primary features that shapes emotional perception of music — fast tempos (130+ BPM) are associat…
Text-to-Audio (Text-to-Audio Generation, TTA)
A class of generative AI models that synthesise non-speech sound (environmental sounds, sound effects, music stems) from a text prompt - for example producing the sound of 'leaves rustling in wind' or…
Text-to-Sound (Text-to-Audio, TTA, Sound Generation from Text)
A class of generative AI models that synthesize non-speech audio - sound effects, ambient environments, foley, or short music clips - from a natural-language description such as 'a door creaking shut'…
Timbre (Tone Colour, Tone Color)
The perceived quality or 'colour' of a sound that distinguishes different sources playing the same pitch and loudness — for instance, the difference between a flute, a violin, and a human voice singin…
Voicemark (Voice Bookmark, Audio Bookmark)
A navigable audio marker or bookmark that allows users to quickly locate and access specific sections of web content or documents through speech or keyboard interaction. Voicemarks are created by anal…