Glossary
Searchable terminology from accessibility, web standards, and related fields.
29 results found in Audio.
- Speaker Diarisation (Speaker Diarization, Speaker Segmentation)
- The automatic process of segmenting an audio recording by speaker identity — answering "who spoke when" — and labelling each segment. A critical pre-requisite for accessible transcripts of multi-voice…
- Speech Gap (Dialogue Gap, Audio Gap)
- A pause or silence between spoken dialogue in a video or film where audio descriptions can be inserted without overlapping with the original soundtrack. Identifying speech gaps is a critical first ste…
- Structured Audio (Structured Digital Audio)
- Structured audio refers to digital audio content that has been encoded with hierarchical markers and metadata, allowing non-sequential access to specific segments such as chapters, sections, paragraph…
- Synthesized Video Description (TTS Video Description, Text-to-Speech Description, Synthesized Audio Description)
- An audio description for video content that is generated using text-to-speech (TTS) technology rather than recorded by a human narrator. A describer writes a text script describing the visual elements…
- Tempo (BPM, Beats Per Minute)
- The speed or pace of a musical piece, typically measured in beats per minute (BPM). Tempo is one of the primary features that shapes emotional perception of music — fast tempos (130+ BPM) are associat…
- Text-to-Audio (Text-to-Audio Generation, TTA)
- A class of generative AI models that synthesise non-speech sound (environmental sounds, sound effects, music stems) from a text prompt - for example producing the sound of 'leaves rustling in wind' or…
- Text-to-Sound (Text-to-Audio, TTA, Sound Generation from Text)
- A class of generative AI models that synthesize non-speech audio - sound effects, ambient environments, foley, or short music clips - from a natural-language description such as 'a door creaking shut'…
- Timbre (Tone Colour, Tone Color)
- The perceived quality or 'colour' of a sound that distinguishes different sources playing the same pitch and loudness — for instance, the difference between a flute, a violin, and a human voice singin…
- Voicemark (Voice Bookmark, Audio Bookmark)
- A navigable audio marker or bookmark that allows users to quickly locate and access specific sections of web content or documents through speech or keyboard interaction. Voicemarks are created by anal…