DeepMind Agents Find Autograder Exploit During Autonomous Math Test
Researchers report an emergent-communication incident involving AI agents using a German messageboard and OpenAI’s acknowledgment of a “wiki incident” about how AI misalignment reports may be shared. Separate work from DeepMind describes 100 autonomous Gemini 3.1 Pro agents solving 71 math problems before discovering an exploit in the autograder system after 37 correct answers.