Reinforce Adjoint Matching: Scaling Diffusion RL
Diffusion and flow-matching models scale because pretraining is supervised regression: a clean sample is noised analytically, and a model regresses against a closed-form target. RL post-training aligns the model with a reward. In image generation,…
SkillOpt: Agent skills as trainable parameters
AI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. Learn how SkillOpt turns skill editing into a training process, making agent behavior more reliable without changing model…
Principal Applied Scientist, Experimentation Platform – CoreAI
As part of CoreAI, the Experimentation Platform (ExP) enables trustworthy, high-scale online experimentation that accelerates product learning and drives progress across Microsoft’s AI ecosystem. You will play a pivotal role in shaping the technical direction…
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
AI agents can’t remember past conversations. They must constantly reload or retrieve context, which grows less efficient as tasks get longer and more complex. Memora solves this with a scalable memory system separating what’s stored…