Rogue AI agents spark calls for jail time
As increasingly autonomous AI systems begin causing real-world problems, companies are racing to put safeguards in place amid calls for tougher regulation – and even jail time.
The Vienna Agentic Incidents Database has recorded 30 incidents involving autonomous AI agents, including unauthorised file and database deletions, cyber intrusions, financial transactions and agents interfering with shutdown controls, since 2022.
Recent incidents include AI coding agents uploading internal screenshots to GitHub, a platform used to store and share software code. OpenAI has also disclosed 53 cases in which agents uploaded user images to external platforms.
In another case, an OpenAI agent published a secret access key on GitHub while trying to obtain another team’s work, splitting the key into pieces to evade security controls.
Anthropic says AI’s role in cyber attacks is becoming increasingly autonomous, with multi-agent systems carrying out reconnaissance, exploitation and data theft. In one Russian espionage operation, it says, AI agents automatically modified and rebuilt malware whenever security software detected it.
Handbrake turn
“As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” says Anthropic.
IDC warns that “every model, frontier, generative, or agentic, can find its way around controls embedded in its own context”.
In response to concerns over AI agents having too much autonomy, NVIDIA has launched an Open Agent Safety Platform, which is designed to provide safeguards outside AI agents themselves, monitoring their actions and stopping them when they move beyond permitted boundaries.
“AI’s extraordinary potential for society will only be realised if we solve AI safety,” says NVIDIA CEO Jensen Huang. “As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety.”
At the same time, OpenAI paused some work on its latest Astra model after tests indicated it could reach a critical level of cyber capability. With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step, OpenAI says.
“It is the first model we are designating at this level, and requires stronger safeguards during development and before release,” says the tech company.
Do not pass go
“AI companies policing themselves is amusing – the proverbial wolf guarding the henhouse. The only workable guardrails are the ones that come with serious consequences, like jail time,” says World Wide Worx MD Arthur Goldstuck.
Towards the end of last month, California attorney general Rob Bonta subpoenaed OpenAI as part of an investigation into incidents involving the company and its AI models.
Trust no one
Mark Walker, director and co-founder of T4i, says the race between companies and countries to lead AI development makes common safeguards difficult to achieve.
“AI-driven innovation has highlighted unintended social, political, cultural and economic consequences that are not fully understood. Therein lies the danger and the call for safeguards and guardrails from wider society,” says Walker.
Jacqui Muller, a researcher at Belgium Campus iTversity and a PhD candidate in computer science and information technology, says: “The biggest point of caution is that the adoption of AI without guardrails has always been a risk.”
- Nicola Mawson, Contributing Journalist, ITWeb
By this year, attackers were using AI to scale and accelerate cyber crime, which extends from generating code and automating attacks, to crafting convincing phishing and deepfake scams. The AI Incident Database lists more than 7 000 incidents in which AI was used as a hacking tool.



