Glossary
Searchable terminology from accessibility, web standards, and related fields.
33 results found in evaluation.
- Accessibility Metrics (Web Accessibility Metrics, Accessibility Scores, Accessibility Measurement)
- Quantitative methods for measuring and scoring the accessibility level of websites or digital content. Accessibility metrics typically work by evaluating web pages against checkpoints derived from sta…
- Agreement Rate (AR)
- A statistical measure used in end-user gesture elicitation studies to quantify how much consensus participants show when proposing gestures or interactions for a given task (referent). Agreement rate …
- Back-translation (Reverse translation)
- A validation technique in cross-linguistic instrument translation where an independently translated version (e.g., ASL video) is translated back into the source language (e.g., English) by someone who…
- Benchmark dataset (Evaluation dataset, Test benchmark)
- A standardized dataset used to evaluate and compare the performance of AI models, algorithms, or systems against established baselines. In accessibility, the absence of benchmark datasets that include…
- Caption quality metric (ACE metric, Caption evaluation metric)
- A measure designed to predict how understandable automatically generated captions are for Deaf and Hard-of-Hearing users, as an alternative to standard Word Error Rate which correlates poorly with act…
- Cloze Test (Cloze Procedure, Cloze Deletion Test)
- A reading comprehension assessment method in which words are systematically deleted from a text and the reader must fill in the missing words based on context. Developed by Wilson Taylor in 1953, cloz…
- Cognitive Walkthrough (Expert Walkthrough)
- An accessibility and usability evaluation method in which one or more experts step through a series of tasks from the perspective of a target user, identifying potential barriers and difficulties at e…
- Cross-syndrome comparison (Cross-disability comparison)
- A research methodology that evaluates a technology or intervention with participants from multiple disability groups to determine whether findings and design principles generalize across conditions. C…
- Ecological validity (Real-world validity)
- The degree to which research findings from controlled laboratory settings accurately reflect behaviour and performance in real-world everyday contexts. In accessibility research, ecological validity i…
- Error-spread modelling (Error propagation modelling, Error radiation)
- An approach to evaluating the impact of speech recognition errors that accounts for how a single misrecognized word degrades comprehension of its neighbouring words, not just the word itself. For exam…
- Formative Evaluation (Formative Usability Testing, Formative Assessment)
- Usability evaluation conducted early in the design process using prototypes, mockups, or wireframes to identify design problems and inform improvements. Formative testing is qualitative and iterative,…
- GOMS (Goals, Operators, Methods, and Selection rules, KLM, Keystroke-Level Model, CPM-GOMS, CMN-GOMS)
- A family of human-computer interaction models used to predict how long it will take a user to complete a task with a given interface. GOMS stands for Goals, Operators, Methods, and Selection rules — t…
- Gamified evaluation (Game-based assessment, Gamified testing)
- A research methodology that incorporates game design elements — such as challenges, scoring, progressive difficulty, and rewards — into the evaluation of technology or user performance, to increase pa…
- Heuristic evaluation (Expert review, Usability inspection)
- A usability and accessibility evaluation method where trained evaluators systematically assess an interface against a set of recognized principles or guidelines (heuristics) to identify potential prob…
- Intrinsic Motivation Inventory (IMI)
- A standardized psychometric instrument used to assess participants' subjective experience during activities, measuring dimensions such as interest/enjoyment, perceived competence, effort/importance, v…
- NASA-TLX (NASA Task Load Index, Task Load Index, NASA TLX, Raw-TLX, Raw TLX, RTLX)
- A widely used subjective workload assessment tool developed by NASA that measures perceived workload across six dimensions: mental demand, physical demand, temporal demand, performance, effort, and fr…
- Participant pool bias (Sampling bias, Recruitment bias)
- Systematic distortion in research findings caused by the demographic characteristics and backgrounds of study participants, rather than by the technology or intervention being evaluated. In accessibil…
- Participatory Evaluation (PE)
- A research approach in which the people affected by a program, technology, or intervention are actively involved in evaluating it, rather than being passive subjects of assessment. In accessibility re…
- Perturbation testing (Counterfactual testing, Template-based testing)
- A bias evaluation methodology for NLP models that systematically substitutes identity-related terms (e.g., disability phrases) in otherwise identical sentences to measure whether the model produces di…
- Psychometric validation (Psychometric evaluation, Instrument validation)
- The process of establishing that a measurement instrument (such as a questionnaire or scale) possesses adequate reliability (consistency of measurement), criterion validity (correlation with establish…