Literature Reviews
Reviewed research papers, articles, and publications relevant to digital accessibility.
11 results found tagged multimodal AI.
-
Mnemonic Tracing: Using Eye Gaze to Search for Visual Memories
Mnemonic Tracing is a non-verbal image-retrieval interaction in which a user, wearing eye-tracking glasses, deliberately retraces the contents of a remembered image with their gaze on a blank surface. The paper builds on gaze-reinstatement research, …
-
From Struggle to Success: Context-Aware Guidance for Screen Reader Users in Computer Use
Chen, Lu, Wang, Qiu, Chen and Yang present AskEase, an NVDA add-on that delivers on-demand, step-by-step, screen-reader-friendly guidance for blind and low-vision computer users tackling unfamiliar desktop software. The work responds to a persistent …
-
SceneScout: Towards AI-Driven Access to Street Level Imagery for Blind Users
Jain, Findlater and Gleason present SceneScout, a prototype web interface that uses a multimodal large language model (GPT-4o) to make street level imagery — the panoramic pedestrian-height photography behind Apple Maps Look Around and Google Street …
-
Expanding Perspectives to Improve Access to Visual Archives through Multimodal Image Enrichment
This paper addresses a pervasive challenge in the Galleries, Libraries, Archives and Museums (GLAM) sector: large-scale visual collections that have been digitised but remain undiscoverable because they lack descriptive metadata. The authors, from th…
-
Multi-Perspective Visual Contrastive Decoding for Reliable Assistance
This technical paper presents MPVCD (Multi-Perspective Visual Contrastive Decoding), a framework designed to address the reliability of AI-generated visual descriptions for people who are blind or have low vision (BLV). The core problem it tackles: w…
-
Making Accessible Movies Easily: An Intelligent Tool for Authoring and Integrating Audio Descriptions to Movies
This paper introduces EasyAD, an intelligent tool that automates the process of authoring and integrating audio descriptions (AD) into movies for blind and visually impaired (BVI) users. The traditional AD production workflow is highly labor-intensiv…
-
AccessMenu: Enhancing Usability of Online Restaurant Menus for Screen Reader Users
This paper addresses the significant accessibility barriers that blind and visually impaired (BVI) screen reader users face when trying to access online restaurant menus, which are typically presented as images or PDFs. The research proceeds in two p…
-
DescribePro: Collaborative Audio Description with Human-AI Interaction
This paper presents DescribePro, a web-based platform that combines human expertise with AI capabilities to create and refine audio descriptions (AD) for video content. The system addresses the fundamental tension in AD production: human-crafted desc…
-
Temp access: Reflecting on multimodal GAI as an accessibility technology for temporary disability
This paper presents an autoethnographic account of using multimodal generative AI (GAI) tools as accessibility technology during a period of temporary disability. The author, an accessibility researcher, experienced an illness that simultaneously imp…
-
Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
This paper addresses a critical safety problem in AI-powered visual access technology: multimodal large language models (MLLMs) like GPT-4o, Gemini, and Claude produce fluent, confident image descriptions that can contain fabricated content, misinter…