Glossary

Searchable terminology from accessibility, web standards, and related fields.

7 results found in Datasets.

Data Annotation (Data labeling, AI labeling)
The process of attaching labels, transcriptions, bounding boxes, or other structured metadata to raw data so that it can be used to train, evaluate, or benchmark machine-learning models. Annotation is…
Dataset Collection (Data Collection Protocol)
The process of gathering, curating, and documenting data used to train, evaluate, or benchmark machine learning systems. In accessibility contexts, dataset collection decisions — who contributes, what…
Disability-First Dataset (Disability-first AI dataset)
An approach to AI dataset creation, articulated by Theodorou et al. and others, that treats serving a disability community as the primary objective rather than collecting disability data as a minority…
Ground Truth (Gold standard, Reference labels)
In machine learning, the labels treated as authoritative when training or evaluating a model - typically produced by human annotators or expert consensus and assumed to represent the 'correct' answer.…
Inter-Annotator Agreement (IAA, Inter-rater agreement, Inter-coder agreement)
A statistical measure of how consistently two or more human annotators assign the same label to the same data item, widely used in NLP, computer vision, and AI dataset construction as a proxy for labe…
ORBIT Dataset (Object Recognition for Blind Image Training)
A disability-first machine learning dataset for teachable object recognition, contributed by people who are blind or have low vision. The original ORBIT dataset (Massiceti et al., 2021) contains 3,822…
Sign Language Corpus (ASL Corpus, Signed Language Corpus)
A structured collection of recorded signed-language performances — typically video, and increasingly motion-capture data — annotated by expert signers with time-stamped linguistic information such as …