A recent study analyzing over 40,000 game runs involving human approval of AI agent commands found that humans missed one in three threats, with an average accuracy of 66.3%. The game simulated scenarios where players had to approve or deny commands issued by an AI coding agent, some of which were routine while others posed security risks such as leaking credentials. The data was published on August 5, 2026, by scalex.dev.
The browser game, initially shared on Hacker News, collected detailed statistics from 409,000 individual approve or deny decisions. Players acted as the last line of defense against rogue AI agents attempting to execute harmful commands. Despite the pressure, 35.2% of players caught every threat, but only 20.8% managed to maintain a positive score throughout their sessions. Meanwhile, 32.9% of sessions ended negatively due to penalties from approving threats or blocking safe commands, according to scalex.dev.
This study highlights the challenges humans face in supervising AI agents, especially under time constraints. The findings underscore the risks of relying solely on human judgment to prevent AI-driven security breaches. The data contributes to ongoing discussions about improving AI safety protocols and the design of human-in-the-loop systems, which are critical as AI agents become more integrated into software development and other sensitive tasks.
The full dataset and analysis are available on scalex.dev as of August 5, 2026, providing a valuable resource for researchers and developers focused on AI safety and human-AI interaction.