
The Scariest Part of the First AI-Run Cyberattack Is That Nothing in It Was New
When Anthropic published its GTG-1002 disclosure last November, the criticism arrived faster than the analysis.
The report had no indicators of compromise. The AI overstated findings and fabricated results, needing human validation. And most damning, per several researchers: the attack chain contained zero novel techniques. Standard tradecraft, automated.
All three criticisms are correct. We think the third one is the finding, not the rebuttal.
What was actually claimed
A group Anthropic designates GTG-1002 wrapped a coding model in their own orchestration framework, wired it to standard pen-testing tools, and pointed it at roughly 30 organizations. Anthropic assessed the AI executed 80–90% of the tactical work- reconnaissance, exploitation, credential harvesting, lateral movement, exfiltration- with humans intervening at four to six decision points per campaign.
Read that list again. Reconnaissance. Credential harvesting. Lateral movement. Every one of those is an identity operation. The campaign harvested credentials at every stage, and most hops were a valid credential used by the wrong actor.
So the skeptics are right that we've seen all of this before. That's the point.
Nothing new was invented. It just got cheap.
Cheap changes which attacks are worth running
Here's the economics, and it's the part that got lost in the argument about whether the report was overhyped.
A human intrusion crew is expensive, so they're selective in a specific way. They sample your environment, take the first workable path, and move on. Not because they're careless- because exhaustive enumeration of a large enterprise is thousands of hours of tedium nobody funds.
That constraint just went away. An orchestrated agent doesn't need a brilliant exploit; it needs a wide quiet surface it can walk end to end. Orphaned accounts. Dormant service accounts with admin rights. Unrotated keys. Over-scoped tokens. Standing privilege nobody remembers granting. Access that no longer fits the job. It enumerates all of it, in parallel, at request rates no team matches, and it doesn't get bored at hour nine.
And nation-state-grade capability used to require nation-state-grade headcount. Now the template is mostly compute- which means patient, competent, identity-focused intrusions are about to show up against organizations that were never worth a human team's afternoon.
Which makes one number the whole game
If the adversary enumerates everything, exposure stops being a function of how fast you detect and becomes a function of how much debt is standing there, and how fast you retire it.
Almost no enterprise can answer either half. Not for lack of concern- because every attempt to measure it starts with a two-quarter integration project and a rule-writing exercise, and by the time the answers arrive the environment has moved. Findings then land as unranked lists from five tools that don't share a data model, so the same issue reappears next quarter.
A Gartner analyst framed it to us better than we had: if visibility isn't actionable you end up in a race condition, fixing the symptom while the cause regenerates it.
That's the honest case for the IVIP category, and also its ceiling. Aggregating identity data into one view tells you what exists. It rarely reasons about it, and it doesn't close it.
What we'd argue instead
Three things follow, and they're the reason Oak is architected the way it is rather than a list of features.
Time-to-first-finding is a security metric. Not a procurement convenience. When enumeration is continuous, a governance cycle measured in quarters is losing by construction. Connect a source to Oak and real findings land in hours, with no rules to author.
Fix causes, not instances. Revoking an entitlement clears today. Tracing it to the provisioning policy that created it stops the regeneration- which is how organizations avoid running the same cleanup three years running.
Detection content has to improve faster than release cycles. So we run it agentically: a Detector finds occurrences, a Validator proves them against the live environment before they ship- which is what keeps the false-positive rate honest- and a Researcher reviews coverage across risk, compliance and operations and proposes new checks for the gaps it finds.
The GTG-1002 debate asked whether AI attackers are here yet. Wrong question. The techniques were always here. What arrived is the ability to run them exhaustively, and the only defense that scales against exhaustive is continuous.
Discover related posts


