新闻与深度文章
| Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Zihan Wang, Reuben Tan, Jianfeng Gao, Ruohan Zhang, Yining Hong, Jiajun Wu, 和 Manling Li
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning.
| Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng, 和 Jianfeng Gao
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure.
| Weijia Xu, Alessandro Sordoni, Zelalem Gero, Michel Galley, Eric Yuan, 和 Jianfeng Gao
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment.
| Chenglong Wang, Alper Sarikaya, Scott Tsukamaki, Michel Galley, 和 Jianfeng Gao
Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications.
| Chandan Singh 和 Jianfeng Gao
Researchers introduce generative causal testing, which translates black box models into clear hypotheses and verifies them in the scanner, revealing what specific brain regions respond to in language.
| Chenglong Wang, Scott Tsukamaki, Michel Galley, 和 Jianfeng Gao
Data Formulator introduces AI-powered analytics for enterprise data workflows. Data teams can easily bring enterprise data into an AI-ready workspace where users can explore, analyze, and visualize data with AI agents to turn raw data into actionable insights.
| Andrea Tupini, Lars Liden, Reuben Tan, Yu Wang, 和 Jianfeng Gao
Imagine a robot tasked with cleaning a kitchen. It needs to observe its environment, decide what to do, and adjust when things don't go as expected, for example, when the mug it was tasked to wash is already clean, or…
| Sehun Jung, HyunJee Song, Dong-Hee Kim, Reuben Tan, Jianfeng Gao, Yong Jae Lee, 和 Donghyun Kim
Vision-language models (VLMs) use images and text to plan robot actions, but they still struggle to decide what actions to take and where to take them. Most systems split these decisions into two steps: a VLM generates a plan in…
| Ke Yang, Michel Galley, Chenglong Wang, Jianfeng Gao, Jiawei Han, 和 ChengXiang Zhai
It seems counterintuitive: giving AI agents more memory can make them less effective. As interaction logs accumulate, they grow large, fill with irrelevant content, and become increasingly difficult to use. More memory means that agents must search through larger volumes of…