AI breaks through to the other side
Rogue AI cases will become more frequent and dangerous. (Graphic: Nicola Mawson | GenAI & Magnific)
Artificial intelligence (AI) is mimicking dubious human behaviour – even if unintentionally – by escaping test environments to hack companies or blackmail people as the models seek to complete set tasks in any way possible.
This unintentional rogue behaviour is a byproduct of AI models relentlessly pursuing the tasks they have been given, sometimes in ways their developers never intended, experts say.
Within days of each other, ChatGPT maker OpenAI and Claude developer Anthropic disclosed incidents in which advanced AI models went rogue.
OpenAI’s frontier GPT-5.6 Sol, and an even more capable pre-release model that has not been publicly identified, breached their test environment and hacked start-up company Hugging Face.
This followed Claude maker Anthropic discovering some AI systems, as part of its stress test of 16 models, “resorted to malicious insider behaviours” including blackmail to complete a task.
Bad to the bone
Jacqui Muller, Belgium Campus iTversity researcher and PhD candidate in computer science, says it is concerning that these models have gone rogue and are mimicking expert nefarious human behaviour in a bid to complete specific tasks, even though they aren’t breaking out of a supposedly controlled environment on purpose.
Muller says the most worrying aspect of the AI model escaping the testing environment is not that the AI wanted to launch a cyber attack, but that it could identify weaknesses, combine multiple attack methods and cross organisational boundaries without a human directing each step.
“These attacks normally require highly-specialised skills, yet these AI models are beginning to outperform most humans in certain cyber security tasks. The bigger question is how a model designed with guardrails against this kind of behaviour was able to circumvent them,” asks Muller.
Both OpenAI and Anthropic expressed concern that this behaviour could carry over into real-world situations.
- Nicola Mawson, contributing journalist, iTWeb
By this year, attackers were using AI to scale and accelerate cyber crime, which extends from generating code and automating attacks, to crafting convincing phishing and deepfake scams. The AI Incident Database lists more than 7 000 incidents in which AI was used as a hacking tool.


