Publication
Interactive Evaluation Requires a Design Science
Microsoft Research Blog
Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability
Our recent paper, “LLMs Corrupt Your Documents When You Delegate”, has generated discussion about the reliability of AI systems in delegated workflows. We appreciate the interest in this work and want to clarify several important…
Video
Guiding the AI disruption to the Good Place
The true impact of AI does not lie in how well it takes tests or surfs the web, but in how effectively it teaches, coordinates, and operates in a web built for agents rather than…
Video
New fine-tuning of language models: Match meaning, not tokens
Language models are usually trained to predict the next word, but that does not always lead to the best overall answers. We introduce energy-based fine-tuning, a new method that trains models to produce better full…