Validating and Refining Measurements for Generative AI Evaluation Via Stakeholder Engagement

The 2026 ACM Conference on Fairness, Accountability, and Transparency |

Generative AI systems are notoriously difficult to evaluate, in part because definitions of their capabilities, behaviors, and impacts can be contested across use cases, cultures, and languages. To address this, machine learning researchers have begun to draw on measurement theory from the social sciences to develop systematic frameworks for the measurement tasks involved in generative AI evaluation. In this tradition, the first step in tackling a measurement task is to precisely define or systematize the concept to be measured. Systematization creates an opportunity to include stakeholders—including those who will use or be impacted by a system—in conceptual debates about the proposed definitions and boundaries of a concept, ultimately leading to measurements that are more reflective of stakeholder needs and values. In this paper, we explore how to validate and refine systematized concepts via stakeholder engagement. We situate our study in the context of measuring erasure, engaging stakeholders in validating and refining the systematized concept of erasure developed by Corvi et al. [13]. We conducted six workshops with 23 participants. Participants’ understandings of erasure largely aligned with Corvi et al.’s systematized concept, but also surfaced boundary tensions requiring refinement and gaps suggesting conceptual expansion. We provide suggestions for how the systematized concept could be refined to better reflect stakeholder perspectives. We reflect on learnings from our study to derive recommendations for how to better engage stakeholders in the measurement process. Content Warning: This paper includes examples of language related to erasure that may be harmful to readers.