Glossary

Searchable terminology from accessibility, web standards, and related fields.

40 results found in Captioning.

Real-Time Captioning (CART, Communication Access Realtime Translation, Live Captioning, Real-Time Text)
The instant conversion of spoken language into text displayed simultaneously as speech occurs, provided either by a trained human captioner or through automatic speech recognition (ASR) technology. Re…
Real-Time Captioning (Live Captioning, Live Transcription)
The process of converting spoken language into text simultaneously as it is being spoken, displayed with minimal delay. Real-time captioning is essential for deaf and hard of hearing individuals to pa…
Real-Time Captioning (Live Captioning, Live Speech-to-Text)
The process of converting spoken language to text simultaneously or with minimal delay as the speech occurs. Real-time captioning can be produced by human transcriptionists (CART, C-Print, TypeWell), …
Remote Captioning (Remote CART, Remote Real-Time Captioning)
A live captioning service delivered at a distance, in which a human captioner (CART provider) or automatic speech recognition system receives an audio feed from a meeting, classroom, or event over the…
Respeaking (Speech-to-Speech Captioning, Voice Writing)
A real-time captioning method in which a trained operator listens to speech and repeats it clearly into a speech recognition system optimized for their voice, producing captions. Respeaking is commonl…
SRT (SubRip, SubRip Text, SRT Subtitle Format)
SRT (SubRip Text) is a widely used plain-text subtitle file format originally created by the SubRip software for extracting subtitles from DVDs. An SRT file contains sequentially numbered subtitle ent…
Social Media Video Captions (SMVC)
An umbrella term for the textual or symbolic elements — platform-generated captions, creator-edited captions, user-generated captions, and non-speech information such as sound effects, music cues, or …
Stenographic Keyboard (Steno Machine, Stenotype, Shorthand Keyboard)
A specialized keyboard used by CART captioners and court reporters that allows simultaneous pressing of multiple keys to represent syllables, words, or phrases in a single stroke, enabling transcripti…
Subtitle (Subtitles, Open captions (video), Movie subtitles)
On-screen text that reproduces the spoken dialogue of a video, most commonly rendered in a "movie subtitle" style (white text with a black outline, one or two lines at the bottom of the frame). Subtit…
Subtitles (Captions, Closed Captions, CC)
Text displayed on screen that represents the spoken dialogue and other relevant audio information in video content. Subtitles (called captions in North America) are essential for deaf and hard of hear…
Tactile Captions (Haptic Captions, Vibrotactile Captions)
An enhanced captioning approach that supplements traditional text-based captions with vibrotactile feedback, allowing deaf and hard of hearing viewers to feel non-speech sounds (such as phone rings, d…
Text Alignment (Sequence Alignment, Transcript Alignment)
The process of matching corresponding segments between two or more text sequences that represent the same content but may differ in timing, wording, or structure. In captioning systems, text alignment…
Tracked Captions (Speaker-following captions, Dynamic captions)
Captions that move dynamically within the video frame to stay near the current speaker's face or mouth, rather than remaining anchored at a fixed position (typically the bottom of the video). Tracked …
Transcript (Text Transcript, Video Transcript, Audio Transcript)
A written document containing the complete text of spoken content from a video or audio recording, presented separately from the media rather than synchronized with it. Unlike captions, which appear o…
Transcripts (Transcript, Text Transcript)
A written, text-based representation of spoken audio or audiovisual content. WCAG 2.1 success criterion 1.2.1 (Audio-only and Video-only Prerecorded) requires an alternative for time-based media — typ…
User-Generated Captions (UGC captions)
Captions created and added to video content by non-professional contributors — typically the video's own creator or community members — rather than by professional captioners or fully automated system…
Wav2Vec (Wav2Vec2, Wav2Vec 2.0)
A family of self-supervised speech representation models from Meta AI that learn rich acoustic embeddings directly from raw waveform audio without requiring transcribed training data. Wav2Vec 2.0, int…
WebVTT (Web Video Text Tracks, Web Video Text Tracks Format)
WebVTT (Web Video Text Tracks) is the W3C standard text format for providing timed text tracks — including captions, subtitles, descriptions, chapters, and metadata — synchronized with HTML5 <video> a…
Word Error Rate (WER)
A metric used to evaluate the accuracy of automatic speech recognition (ASR) and captioning systems, calculated as the number of word-level errors (insertions, deletions, and substitutions) divided by…
Word Error Rate (WER)
A standard metric for evaluating speech recognition and captioning accuracy, calculated as the number of insertions, deletions, and substitutions needed to transform the transcribed text into the refe…