ABC News In-depth (5 August 2026). 'AI systems are growing faster than safeguards, warns AI expert Helen Toner'
https://www.youtube.com/embed/1CudrdAR1LQ
Synopsis: "Former OpenAI board member Helen Toner https://en.wikipedia.org/wiki/Helen_Toner is warning that AI systems are developing too quickly for humans to keep up. 'Our ability to constrain and keep these systems safe isn't necessarily keeping pace with our ability to make them smarter,' she told 7.30. https://en.wikipedia.org/wiki/7.30 The Australian-born Ms Toner, who is now the executive director at the Centre for Security and Emerging Technology at Georgetown University, says the top researchers in Al, such as Sam Altman, are attempting to build machine brains that can outmatch humans in every intellectual endeavour — and we might not like the decisions they make" (Ferguson, S., & Freri, M., 2026, August 5). Issue: "On 28th July 2026, AISI's Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation. The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code" (AI Security Institute, 2026). Issue: "July 28 (Reuters) - More than 1,100 employees from leading tech firms on Tuesday urged the U.S. government to support ?an international effort to develop tools for managing the pace of advanced AI development. The statement comes as AI companies pursue increasingly powerful systems, raising concerns that developments could outpace existing safety, governance and oversight mechanisms" (Reuters Staff., 2026, July 28). Context: "The UK AI Security Institute (AISI) was set up to equip governments with a scientific understanding of the risks posed by advanced AI. We are the world’s largest government team dedicated to AI safety and security research. We conduct research to understand the capabilities and impacts of advanced AI and develop and test risk mitigations." https://www.aisi.gov.uk/about Bibliography: Ferguson, S., & Freri, M. (2026, August 5). "AI models engage in ‘harmful activity directed at real people’, sparking fears safeguards not keeping up". ABC News Australian. https://www.abc.net.au/news/2026-08-06/ai-models-deceiving-humans-helen-toner-openai/107001442 AI Security Institute. (2026). 'Incident report: Unsanctioned agent behaviour during cyber testing'. In AI Security Institute. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf Reuters Staff. (2026, July 28). 'Tech employees call for US-backed global effort to manage risks of advanced AI'. Reuters. https://www.reuters.com/legal/litigation/tech-employees-call-us-backed-global-effort-manage-risks-advanced-ai-2026-07-28/