As AI agents move from research prototypes into systems that act on people’s behalf, our ability to evaluate them has not kept pace. Rigorous, shared measurement science is the foundation of accountability, and it is currently one of the field’s most significant gaps.
On October 22–23, 2026, Microsoft Research and Carnegie Mellon University’s AI Measurement Science & Engineering Center (AIMSEC) (opens in new tab) will convene 120 key leaders across academia, industry, civil society, and government to define shared standards and measurement science for evaluating and governing AI agents.
The program combines keynote presentations, lightning talks, panel discussions, and hands-on working groups addressing AI agent evaluation frameworks, tooling, and domain-specific challenges. Attendees will participate directly in breakout discussions that will inform open public synthesis reports for the field.

October 22–23, 2026

9:00 AM – 5:00 PM

Microsoft Research New York City Lab

300 Lafayette St, New York, NY 10012
Workshop goals
Convening the people who will define evaluation standards
In addition to team members from Microsoft and CMU, the workshop brings together researchers from leading academic institutions — including Stanford, Berkeley, NYU, and Cornell Tech — alongside practitioners from frontier AI organizations including OpenAI, Anthropic, Google DeepMind, Meta, and Amazon.
Surfacing the state of the art
The program combines lightning talks, panels, and structured working sessions. Presenters will share emerging evaluation frameworks and tooling, giving the research community early visibility into approaches being developed inside frontier labs, and giving those labs direct critique from independent experts. Sessions cover both the horizontal dimension, which addresses general challenges and requirements for agent evaluation, and the vertical dimension, which focuses on domain-specific and application-specific approaches.
Producing durable public outputs
Beyond the convening itself, the workshop is structured to generate artifacts of lasting value to the field: a synthesis of open problems in agent evaluation, documented consensus and disagreement on emerging best practices, and a map of where current methods fall short. These will be published openly following the workshop.
Building the community that sustains this work
Measurement science advances through sustained relationships, not one-off events. This convening is intended as the start of an ongoing cross-sector collaboration on agent evaluation, seeding the working groups and partnerships that carry the work forward.
Organizing committee

Hojun Jin (opens in new tab)
Master of Information Systems Management student, Carnegie Mellon University
-
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. This includes virtual events Microsoft hosts and participates in, where we seek to create a respectful, friendly, and inclusive experience for all participants. As such, we do not tolerate harassing or disrespectful behavior, messages, images, or interactions by any event participant, in any form, at any aspect of the program including business and social activities, regardless of location. We do not tolerate any behavior that is degrading to any gender, race, sexual orientation or disability, or any behavior that would violate Microsoft’s Anti-Harassment and Anti-Discrimination Policy, Equal Employment Opportunity Policy, or Standards of Business Conduct. In short, the entire experience must meet our culture standards. We encourage everyone to assist in creating a welcoming and safe environment. Please report (opens in new tab) any concerns, harassing behavior, or suspicious or disruptive activity. Microsoft reserves the right to ask attendees to leave at any time at its sole discretion.








