Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Sunnie S. Y. Kim ,
- Wesley Deng ,
- Jennifer Wortman Vaughan ,
- Buxin Su ,
- Wei-Jie Su ,
- Alekh Agarwal ,
- Sharon Li ,
- Martin Jaggi ,
- Daniel G. Goldstein ,
- Nihar B. Shah ,
- Miro Dudík
arXiv
LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.