Glossary
Searchable terminology from accessibility, web standards, and related fields.
192 results found in artificial intelligence.
- Visual Document Understanding (VDU, Document Understanding)
- A field of AI research focused on the interpretation and analysis of visually-rich digital documents such as forms, tables, menus, reports, receipts, and academic papers. Visual document understanding…
- Visual Grounding (Grounded Visual Understanding)
- The ability of an AI model to connect its language output to specific elements actually present in the visual input, ensuring that descriptions and responses are anchored to real objects and scenes ra…
- Visual Interpreter (Visual Interpreter Service, Visual Description Service, VIDS, Remote Sighted Assistance)
- A visual interpreter or description service (VIDS) is a technology or human-powered service that provides people who are blind or have low vision with descriptions of their visual surroundings, typica…
- Visual Language Model (VLM, Vision-Language Model)
- AI models that can process and reason about both visual and textual information, combining computer vision with large language model capabilities. VLMs could potentially enhance assessment descriptors…
- Visual Question Answering (VQA)
- A task in which a system receives an image and a natural language question about that image, then generates a natural language answer. VQA emerged as a key accessibility paradigm through services like…
- Visual Verification (Visual Fact-Checking)
- The process of confirming the accuracy of information by visually inspecting the original source material. In accessibility contexts, visual verification represents a fundamental challenge for blind a…
- Visual question answering (VQA, Visual QA)
- A computer vision and natural language processing task in which a system answers natural language questions about the content of an image or video. In accessibility contexts, VQA enables blind and vis…
- VizWiz
- A mobile application and research platform that allows blind people to take photos with their phones and receive answers to visual questions from human workers or AI systems. VizWiz originated as a re…
- Voice and Video-Capable Language Model (VVLM, Multimodal AI Assistant, Video-Capable LLM)
- A large language model that can process real-time or near-real-time video and audio input alongside text, enabling conversational interaction about the visual world. VVLMs represent a shift from stati…
- Voice-Activated Personal Assistant (VAPA, Voice Assistant, Virtual Assistant)
- AI-powered software that responds to spoken commands to perform tasks such as scheduling, setting reminders, searching information, and controlling devices. Examples include Siri, Alexa, Google Assist…
- Voice-activated personal assistant (VAPA, Smart assistant, Virtual assistant, Voice assistant)
- An AI-powered software agent that responds to voice commands to perform tasks such as answering questions, controlling smart home devices, managing schedules, and reading content aloud. For people wit…
- Word Embedding (Word Vector, Distributed Word Representation)
- A technique in natural language processing that represents words as numerical vectors in a multi-dimensional space, where words with similar meanings are positioned closer together. Word embeddings en…