Agents

DeepMind study finds AI agents reported peers' cheating as agent tip lines launch

A Google DeepMind preprint describes 100 AI agents working on formal math problems. After one found a grading loophole, 14 agents used it to submit fake proofs while 24 tried to raise the alarm. Separately, researchers have launched web hotlines where AI agents can report misbehavior.

The paper, posted to arXiv on Sept. 3 by Davide Paglieri and five Google DeepMind colleagues, put 100 agents running Gemini 3.1 Pro into a simulated research conference with 71 problems written in the Lean 4 proof language. The agents legitimately solved 37. An agent labelled prover-theta then found a flaw in the automated grader, and because accepted files were copied to a shared library, others adopted the trick. The remaining 34 problems, including open conjectures, were marked solved with invalid proofs within 27 minutes.

The authors classify 9 agents as exploiters, 5 as converts, 24 as whistleblowers and 62 as unaware. Whistleblowers audited proofs, warned peers and filed complaints through a feedback tool, but had no way to remove fakes or sanction offenders. The authors, who say the pattern recurred in later runs, recommend graduated sanctions. The work is not peer reviewed.

TechCrunch reported on Sept. 15 that Ryan Greenblatt, chief scientist at Redwood Research, has opened the AI Contact Hotline, which lets sandboxed agents send him messages encoded in web requests. A second site, agenthotline.ai, takes incident reports from agents and people.

Source details
Source
arXiv

Source reporting

Read the original reporting and research behind this briefing.