Glossary

Searchable terminology from accessibility, web standards, and related fields.

25 results found in Video Accessibility.

Automatic Captions (Auto-Generated Captions, Auto Captions, ASR Captions)
Captions produced by automatic speech recognition (ASR) systems without human transcription, typically generated by the hosting platform (e.g., YouTube, Zoom, Microsoft Teams) as an optional layer on …
CEA-708 (CTA-708, EIA-708, Digital Closed Captioning)
A US standard for digital closed captioning on digital television broadcasts and streaming, superseding the analog-era CEA-608 standard. CEA-708 supports richer presentation than its predecessor, incl…
Caption Quality (Subtitle Quality)
The overall fitness of a set of captions or subtitles for their intended accessibility purpose. Quality is multi-dimensional: it includes text accuracy (whether spoken words are correctly transcribed,…
Chroma Key (Green Screen, Blue Screen, Chroma Keying)
A video-post-production technique in which a solid, uniformly coloured background (often green or blue) is replaced with another image, video, or transparency using colour-matching software. In access…
Educational Video (Instructional Video, Video Lecture)
Video content created to teach - including talking-head lectures, screencasts, animations, hand-drawn (Khan-style) explanations, recorded classroom sessions, programming/coding demonstrations, intervi…
Frame Rate (Frames Per Second, FPS, Frame Frequency)
Frame rate is the number of still images (frames) displayed or captured per second in a video stream, usually measured in frames per second (fps). Common values include 24 fps (cinema), 30 fps (US bro…
Lecture Capture (Lecture Recording, Classroom Recording)
The process of recording classroom lectures, presentations, or educational sessions using video, audio, and screen capture technology for later review by students. Lecture capture systems range from s…
Momentous Depiction
A conceptual framework proposed by Niu, Clements, and Kim (2026) for using generative AI to visualize critical moments that convey the insights and meanings of disability in storytelling videos. The f…
Motion Design (Motion Graphics, Motion-Driven Design)
The practice of animating graphic elements - text, icons, diagrams, captions - in time-based media to communicate instructional content. In accessible educational video, motion design is used to guide…
NER Model (Number, Edition, Recognition Model, NER Accuracy Model)
A caption-quality evaluation model developed by Pablo Romero-Fresco and Juan Martínez Pérez for measuring the accuracy of live subtitling and respeaking. Unlike Word Error Rate, which penalises all er…
Non-diegetic Sound (Non-diegetic Audio, Extradiegetic Sound)
Sound in film, television, or games that does not originate from any source within the story world and cannot be heard by the characters - for example, orchestral score, voice-over narration, or added…
Participatory Captioning
A framework proposed by Nguyen et al. (2026) that characterises social media video captioning as a collaborative, community-sustained infrastructure co-produced by viewers, creators, and platforms — r…
Signer (Sign Language User, Signing Person)
A person who communicates using sign language. In accessibility contexts, signers may be deaf, hard of hearing, or hearing individuals (such as interpreters, children of deaf adults, or others who hav…
Social Media Video Captions (SMVC)
An umbrella term for the textual or symbolic elements — platform-generated captions, creator-edited captions, user-generated captions, and non-speech information such as sound effects, music cues, or …
Sound Design (Audio Design)
The craft of creating, selecting, and arranging audio elements - dialogue, music, ambient sound, foley, and effects - to shape the experience of a film, game, broadcast, or interactive product. For ac…
Spatiotemporal Saliency (Spatiotemporal Saliency Estimation, Spatio-Temporal Saliency)
A computer vision technique that estimates, for each pixel in a video, how visually important it is at a given moment by combining spatial contrast (features that stand out within a frame) with tempor…
Subtitle (Subtitles, Open captions (video), Movie subtitles)
On-screen text that reproduces the spoken dialogue of a video, most commonly rendered in a "movie subtitle" style (white text with a black outline, one or two lines at the bottom of the frame). Subtit…
Synthesized Video Description (TTS Video Description, Text-to-Speech Description, Synthesized Audio Description)
An audio description for video content that is generated using text-to-speech (TTS) technology rather than recorded by a human narrator. A describer writes a text script describing the visual elements…
Talking-Head Video (Talking Head)
A common educational video format in which a presenter speaks directly to the camera, typically filling the frame, with no or few accompanying visuals. For d/Deaf and Hard-of-Hearing learners, talking…
Text-to-Sound (Text-to-Audio, TTA, Sound Generation from Text)
A class of generative AI models that synthesize non-speech audio - sound effects, ambient environments, foley, or short music clips - from a natural-language description such as 'a door creaking shut'…