Portrait of Saleema Amershi

Saleema Amershi

Partner Research Manager

About

Saleema Amershi is a Partner Research Manager at Microsoft and a founding member of AI Frontiers, Microsoft’s advanced AI research lab focused on next generation agentic AI models and systems. She leads the Magentic Team, an interdisciplinary group of researchers and engineers building AI agents, multi-agent platforms, and technologies for collective intelligence.

In collaboration with partners across AI Frontiers, her team has created some of Microsoft’s most influential open-source AI technologies, including:

Today, Saleema’s team focuses on multi-agent networks, building the foundations for AI agents to work together with people and other agents. Their work spans agent-to-agent (A2A) harnesses and platforms, emergent risks and behaviors in agent networks, benchmarking and evaluation of collaborative capabilities, and model post-training for collaboration and social reasoning (selected references below).

Previously, Saleema led Microsoft’s Human-AI Experiences (HAX) team, focused on human-AI interaction and responsible AI. She is the author of the widely used Guidelines for Human-AI Interaction (opens in new tab) and HAX Toolkit (opens in new tab), whose guidance and best practices helped shape Microsoft’s Responsible AI Standard (opens in new tab). Saleema also serves on Microsoft’s advisory Committee on AI, Engineering, and Ethics, helping guide responsible AI development across the company.

Saleema holds a PhD in Computer Science & Engineering from the University of Washington’s Paul G. Allen School (opens in new tab) and an MSc in Computer Science and a BSc in Computer Science & Mathematics from the University of British Columbia (opens in new tab).

Selected Work

  1. Magentic Marketplace: an open-source simulation environment for studying agentic markets – Microsoft Research
  2. Red-teaming a network of agents: Understanding what breaks when AI agents interact at scale – Microsoft Research
  3. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests – Microsoft Research;
  4. The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation | OpenReview (opens in new tab);
  5. [2606.05342] SentinelBench: A Benchmark for Long-Running Monitoring Agents
  6. From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models – Microsoft Research