Glossary

Searchable terminology from accessibility, web standards, and related fields.

33 results found in evaluation.

Accessibility Metrics (Web Accessibility Metrics, Accessibility Scores, Accessibility Measurement)
Quantitative methods for measuring and scoring the accessibility level of websites or digital content. Accessibility metrics typically work by evaluating web pages against checkpoints derived from sta…
Agreement Rate (AR)
A statistical measure used in end-user gesture elicitation studies to quantify how much consensus participants show when proposing gestures or interactions for a given task (referent). Agreement rate …
Back-translation (Reverse translation)
A validation technique in cross-linguistic instrument translation where an independently translated version (e.g., ASL video) is translated back into the source language (e.g., English) by someone who…
Benchmark dataset (Evaluation dataset, Test benchmark)
A standardized dataset used to evaluate and compare the performance of AI models, algorithms, or systems against established baselines. In accessibility, the absence of benchmark datasets that include…
Caption quality metric (ACE metric, Caption evaluation metric)
A measure designed to predict how understandable automatically generated captions are for Deaf and Hard-of-Hearing users, as an alternative to standard Word Error Rate which correlates poorly with act…
Cloze Test (Cloze Procedure, Cloze Deletion Test)
A reading comprehension assessment method in which words are systematically deleted from a text and the reader must fill in the missing words based on context. Developed by Wilson Taylor in 1953, cloz…
Cognitive Walkthrough (Expert Walkthrough)
An accessibility and usability evaluation method in which one or more experts step through a series of tasks from the perspective of a target user, identifying potential barriers and difficulties at e…
Cross-syndrome comparison (Cross-disability comparison)
A research methodology that evaluates a technology or intervention with participants from multiple disability groups to determine whether findings and design principles generalize across conditions. C…
Ecological validity (Real-world validity)
The degree to which research findings from controlled laboratory settings accurately reflect behaviour and performance in real-world everyday contexts. In accessibility research, ecological validity i…
Error-spread modelling (Error propagation modelling, Error radiation)
An approach to evaluating the impact of speech recognition errors that accounts for how a single misrecognized word degrades comprehension of its neighbouring words, not just the word itself. For exam…
Formative Evaluation (Formative Usability Testing, Formative Assessment)
Usability evaluation conducted early in the design process using prototypes, mockups, or wireframes to identify design problems and inform improvements. Formative testing is qualitative and iterative,…
GOMS (Goals, Operators, Methods, and Selection rules, KLM, Keystroke-Level Model, CPM-GOMS, CMN-GOMS)
A family of human-computer interaction models used to predict how long it will take a user to complete a task with a given interface. GOMS stands for Goals, Operators, Methods, and Selection rules — t…
Gamified evaluation (Game-based assessment, Gamified testing)
A research methodology that incorporates game design elements — such as challenges, scoring, progressive difficulty, and rewards — into the evaluation of technology or user performance, to increase pa…
Heuristic evaluation (Expert review, Usability inspection)
A usability and accessibility evaluation method where trained evaluators systematically assess an interface against a set of recognized principles or guidelines (heuristics) to identify potential prob…
Intrinsic Motivation Inventory (IMI)
A standardized psychometric instrument used to assess participants' subjective experience during activities, measuring dimensions such as interest/enjoyment, perceived competence, effort/importance, v…
NASA-TLX (NASA Task Load Index, Task Load Index, NASA TLX, Raw-TLX, Raw TLX, RTLX)
A widely used subjective workload assessment tool developed by NASA that measures perceived workload across six dimensions: mental demand, physical demand, temporal demand, performance, effort, and fr…
Participant pool bias (Sampling bias, Recruitment bias)
Systematic distortion in research findings caused by the demographic characteristics and backgrounds of study participants, rather than by the technology or intervention being evaluated. In accessibil…
Participatory Evaluation (PE)
A research approach in which the people affected by a program, technology, or intervention are actively involved in evaluating it, rather than being passive subjects of assessment. In accessibility re…
Perturbation testing (Counterfactual testing, Template-based testing)
A bias evaluation methodology for NLP models that systematically substitutes identity-related terms (e.g., disability phrases) in otherwise identical sentences to measure whether the model produces di…
Psychometric validation (Psychometric evaluation, Instrument validation)
The process of establishing that a measurement instrument (such as a questionnaire or scale) possesses adequate reliability (consistency of measurement), criterion validity (correlation with establish…