Literature Reviews
Reviewed research papers, articles, and publications relevant to digital accessibility.
3 results found tagged multimodal large language models.
-
Sonic Stage: Automatically Generating an Interactive Spatial Soundscape to Facilitate Dialogue Video Comprehension for Blind and Low Vision Viewers
Xu and colleagues (HKUST, Columbia, Aalto, Rochester) tackle a well-known but largely unsolved problem in video accessibility: standard audio description (AD) is constrained not to overlap with dialogue, so dialogue-heavy scenes in films and TV - whe…
-
ViDscribe: Multimodal AI for Customizing Audio Description and Question Answering in Online Videos
Cheema and colleagues (Arizona State University and Saarland University) present ViDscribe, a web platform that layers AI-generated audio description (AD) and conversational visual question answering (VQA) on top of arbitrary YouTube videos for blind…
-
How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People
This CHI 2026 paper reports a two-week diary study with 20 Blind and Low Vision (BLV) participants (ages 19–75, 11 female/9 male, 13 blind/7 low vision) investigating how multimodal large language models (MLLMs) support real-world access to visual in…