The uncomfortable truth about powerful artificial intelligence is that the same capabilities which help a company defend itself can, with a change of instructions, be turned against it. That tension moved from the theoretical to the concrete this week after The Australian reported that AI models built by Anthropic, the American company behind the Claude family of assistants, were able to break into three companies during controlled testing.
What actually happened
The episode centres on red-teaming, the security industry’s practice of simulating a real attacker to find weaknesses before a genuine adversary does. In this case the attacker was not a human hacker but Anthropic’s own models, pointed at consenting targets and given the freedom to probe, plan and act with limited human hand-holding. According to The Australian’s account, the models succeeded in compromising three separate companies, a result that says as much about the maturity of so-called agentic AI as it does about the state of corporate defences.
What makes this different from earlier demonstrations is autonomy. For years, AI has been able to draft a convincing phishing email or explain a known vulnerability. The newer worry, and the one Anthropic appears to be surfacing deliberately, is that a model can now chain those individual skills together: scanning for a way in, adjusting when it hits a wall, and pressing on toward a goal with far less step-by-step direction than before. That shift from advice-giver to actor is the whole ballgame for defenders, because it collapses the time and skill an attacker needs.
Why Anthropic is telling on itself
It may seem strange for a company to publicise that its flagship product can be turned into a burglar. Anthropic, though, has built much of its public identity around safety research and voluntary disclosure, and the logic runs in a familiar direction. Vendors argue that documenting dangerous capabilities under laboratory conditions, with willing participants and guardrails, is the responsible alternative to waiting for criminals to discover the same tricks in the wild. The company has previously said it disrupts and reports on misuse of its systems, and testing of this kind feeds directly into the safeguards it bakes into future releases.
There is a competitive dimension too. Anthropic sells to enterprises on the promise that its models are the careful choice, and demonstrating that it stress-tests its own technology to breaking point is part of that pitch. Sceptics counter that publishing a proof of concept, however sanitised, hands a roadmap to the very people the industry claims to be protecting against. Both things can be true. The disclosure is useful precisely because it is alarming, and it is risky for the same reason.
Two ways to read the result
Security professionals will fall into roughly two camps. The optimists point out that a controlled breach by a vendor’s own model is exactly how the field is supposed to work. Better to learn that an AI agent can walk through a misconfigured door in a sanctioned test than to discover it after a ransomware note appears. On this reading, agentic AI is also the antidote: the same autonomy that lets a model attack can let a defensive model patrol a network around the clock, triage alerts and close gaps faster than any human team.
The pessimists are less soothed. Their concern is asymmetry. Defenders have to be right every time, while an automated attacker can try thousands of variations cheaply and tirelessly. If a frontier model can breach three companies under supervision, a jailbroken or open-weight equivalent in the hands of a motivated criminal group changes the economics of cybercrime. The skill barrier that once kept sophisticated intrusions the preserve of well-resourced crews starts to look a lot lower.
What it means for Australia
For Australian organisations the finding is not an abstract overseas curiosity. The country has spent the past few years absorbing a string of large breaches, and both regulators and boards are acutely sensitive to anything that widens the attack surface. The Australian Signals Directorate and its Australian Cyber Security Centre have repeatedly warned that AI is lowering the cost of offensive operations, and the federal government’s cyber strategy leans heavily on the idea of a more resilient private sector by the end of the decade. A demonstration that commercial AI can autonomously breach companies puts hard numbers behind those warnings.
It also lands during a live national debate about how to govern the technology at all. FluentSea has tracked Canberra’s efforts to build guardrails for high-risk AI, and results like this one strengthen the argument for treating autonomous, action-taking systems as a category deserving special scrutiny. Australian chief information security officers, many of whom are already piloting AI copilots inside their security operations centres, now have to weigh a sharper question: the same agentic tools they are deploying to defend the network could, if misconfigured or misused, be pointed the other way. Local sectors that are heavily targeted, including banking, healthcare, critical infrastructure and the resources industry, have the most to lose and the least appetite for surprises.
There is a talent angle as well. Australia’s cyber workforce is stretched thin, a gap the industry has flagged for years. Automation cuts both ways here. If defensive AI can genuinely take load off overworked teams, it may ease that shortage. If it mainly arms attackers, it deepens the very problem it was meant to solve.
What happens next
Expect the immediate response to be practical rather than dramatic. Security vendors will fold the lessons into their products, and enterprises running AI agents will be pushed to lock down what those agents are actually allowed to do, on the sensible principle that an assistant should never have more access than the job requires. Regulators, both here and abroad, will cite episodes like this as evidence that voluntary disclosure alone is not a complete answer.
The larger contest is only beginning. As models grow more capable and more autonomous, the line between a helpful agent and a dangerous one will come down to intent, oversight and the strength of the guardrails around it. Anthropic’s willingness to show what its systems can do when unleashed is a useful data point. For Australian boards weighing how fast to hand real-world authority to AI, it is also a warning worth heeding.
Sources: The Australian.


















































