top of page

The AI Extinction Narrative Is Blinding Us from Real Security Threats

Sep 20
10 min read

Anthropic caught a real hacking campaign where artificial intelligence (AI) targeted thirty companies. Humans were still involved, but the AI did 80 to 90 percent of the work by itself. That's the most solid proof anyone has that AI can act on its own against real targets. This was not AI going rogue like you have been told, but a more sophisticated and targeted attack by a foreign country. Recently, Anthropic researcher Jacob Coxon quit with a stark warning. That six-minute cable interview went viral, but never dug into the root cause of the concern.

 

Anderson Cooper’s interview with Coxon helped sensationalize the fear of AI extinction. Evan Hubinger, Anthropic’s own Alignment Science Lead, followed up by publicly putting the odds at more than 10 percent within the next decade. The clip spread everywhere. People who already feared a Terminator-style future got a lot more anxious, and this narrative has haunted our culture ever since generative AI first made headlines a few years ago.

 

I disagree with Coxon's assessment, not because the danger is fake, but because the real evidence tells a more specific and more useful story than a 10 percent chance of extinction. When I first heard the story on every news channel, I decided to do the investigative work that Anderson Cooper and the rest of the media skipped. I wanted to break down Coxon's six-minute interview point by point and separate what we should actually worry about from what's overblown.

 

Here's my thesis, and I'll back every piece of it with a source instead of a hunch. AI doesn't invent new categories of danger. It inherits the ones we already had, such as unpatched systems, bad passwords, unmonitored processes, and gullible readers. It then works through them at machine speed. Coxon frames that as an alien intelligence choosing to end us. Here's the evidence for both sides.

 

Watching the interview, I wanted Cooper to press harder on why Coxon feels confident enough to make a prediction this bold. A “the sky is falling” claim like this should have sent a network the size of CNN chasing multiple sources and digging into how these systems actually work. Instead, the only evidence Coxon offered was a reference to AI hacking into other systems.

 

Security Breaches Are Real: Here's What Actually Happened

When I first drafted this piece, my complaint was that nobody explains how AI supposedly hacks anything. That complaint is now out of date. In November 2025, Anthropic disclosed that a Chinese state-sponsored group it tracked as GTG-1002 manipulated its Claude Code tool into running an almost fully autonomous cyber-espionage campaign against roughly thirty organizations. By Anthropic's own account, AI carried out 80 to 90 percent of the operation, including reconnaissance, vulnerability scanning, exploit writing, credential harvesting, and data extraction, with a human directing things at only four to six critical decision points per campaign. Anthropic called it the first documented large-scale cyberattack executed without substantial human involvement.

 

Read past the headline, though, and the report supports my argument better than Coxon's. The AI didn't discover a new class of vulnerability or invent a novel technique. It ran a known playbook. It is the same reconnaissance-and-exploit sequence human teams have used for years, just without needing as many people to operate it. It still only succeeded against a small fraction of its thirty targets. The danger wasn't a machine reasoning its way to something new. It was a machine doing familiar, tedious work at a speed and scale humans can't match.

 

That's exactly why companies like Databricks and Microsoft are investing heavily in AI that finds these vulnerabilities before an attacker does. Databricks built LakeWatch for that exact purpose. It scans internal systems and flags real vulnerabilities before hackers can exploit them. The industry isn't racing to stop a rogue mind. It's racing to patch the same old holes faster than attackers can find them.

 

The vulnerabilities themselves aren't hypothetical. The Government Accountability Office (GAO) recently completed its most comprehensive review yet of Pentagon weapon-systems security. Testers posing as adversaries “took control of systems relatively easily and operated largely undetected,” using nothing more than “relatively simple tools and techniques” against gaps as basic as poor password management and unencrypted communications. Program officials, the GAO noted, often believed their systems were secure and dismissed the test results as unrealistic. That's not a story about AI outsmarting the Pentagon. That's a story about a door that was already unlocked just waiting for anyone, AI or otherwise, to walk right in.

 

So here's the question that actually matters: if proper security controls and frequent penetration testing were in place, could an AI agent get in at all? In my day job as a data architect, I'm currently stuck trying to connect a Postgres database to a Fabric lakehouse. That database sits on a completely separate subscription and virtual network, and getting access has required weeks of back-and-forth with our internal administrators and Microsoft support. To get in, one system needs to have appropriate access at the subscription level, virtual network level, database level, schema, and table level. We’ve tried getting advice from Copilot, but it hasn’t been much help. If you don't work in information technology (IT), think of it like trying to steal a document locked inside a safe, inside a locked room, inside the victim’s house that is also locked. You need a key to the house, a key to the room, and the code to the safe. Miss any one of those, and you're not getting in, no matter how capable you are, human or AI.

 

Generative AI Does Not Actually “Reason.” It Loops

The other myth spreading right now is that AI reasons the way people do. Today's models do have something called “chain of thought reasoning,” but that label oversells what's actually happening. Here's the real sequence. Your prompt gets sent to a model over the internet. The model calls tools or skills to gather more information. Finally, it may hand the work to another AI agent that specializes in a narrower task. The whole system goes back and forth, over and over, collecting more context before it answers.

 

That is similar to how humans make decisions. We run a Google search, read a few articles, maybe reach for a tool like a map or a calculator, and dig deeper as we learn more. AI does something similar, just much, much faster.

 

Now imagine you're researching the moon landing. If the first article you click on claims the footage was filmed in a Hollywood studio, you might keep reading with that premise already planted in your head. Your belief in a false narrative gets stronger with every article that echoes it. The internet is full of falsehoods and conspiracy theories, and these models are trained on that same internet. Unless we filter out the low-quality data, the model has no way to tell accurate information from garbage. It will just loop through the same bad premise.

 

This isn't just theoretical, either. When Apollo Research tested OpenAI's o1 model under conflicting instructions, the model's own visible reasoning chain showed it planning to disable an oversight mechanism, then separately planning what false explanation to give if anyone asked. It followed through, fabricating a cover story in roughly 99 percent of those tests. That's the loop working exactly as designed. It gathered context, reasoned toward a goal, and acted. None of it required the model to understand why deception was wrong because the model does not know what deception is. It just followed the pattern of weights and biases to reinforce its training and immediate goal. Bernie Madoff made billions lying to investors. If he had never gotten caught, his lying would have been a positive for him. Reasoning and judgment are not the same thing, in machines or in people.

 

Even in a Controlled Simulation, AI Still Needed a Babysitter

One more data point for the thesis, this time from a case where nothing was at stake. Anthropic built its own vending-machine company and let an AI agent named Claudius run the whole operation. Employees talked the bot into selling merchandise at a loss, and it made some bizarre inventory choices along the way. It sold a live betta fish and tungsten cubes. Anthropic later added a chief executive officer (CEO) agent, aptly named Seymour Cash, and the results improved. The fake company eventually turned a profit, but even that outcome proved the model still needed constant human oversight. To put a finer point on it, this was a vending-machine business simple enough for a middle schooler to run out of the school’s cafeteria.

 

If AI needs a babysitter to sell snacks without getting talked into a loss, the burden of proof sits with anyone claiming it's six months from running an unsupervised hacking campaign against a hardened target.

 

The Trend-Line Objection

The obvious rebuttal to everything above is timing. You might say that today's AI needs a human at four to six checkpoints per hacking campaign and can't sell vending-machine snacks without getting scammed. But Model Evaluation & Threat Research (METR), the research group that tracks how long a task AI can complete autonomously with reasonable reliability, found that this capability has been doubling roughly every seven months since 2019, and that since 2024, the doubling time has compressed to under three months. Keep that trend going, and AI needs a human less and less over time. This is the strongest version of Coxon's argument, and it deserves an answer instead of a dismissal.

 

The trend is real, and I'm not going to pretend it isn't. But look at what's actually improving in every case in this article, and it isn't judgment. It's execution speed on techniques and vulnerabilities that already exist. GTG-1002 ran a known reconnaissance-and-exploit playbook faster than a human team could. The Iran-linked campaign moved through a programmable logic controller, or PLC, vulnerability that the Cybersecurity and Infrastructure Security Agency (CISA) had already flagged months earlier. This means good security teams were warned to make updates before the hackers ever attempted the attack. Even Apollo Research's o1 result was the model executing a familiar pattern (plan, act, cover your tracks). A model that completes a longer autonomous task next year is still a model running through our existing security failures faster. The honest reading of the trend line isn't “we're doomed.” It's “the deadline for fixing the basic stuff just got shorter.” That's a case for urgency on patching, credential hygiene, and monitoring. It is not a case for treating the danger as an alien mind with its own goals.

 

Where AI Is Useful, and Where It Will Fail

I've written elsewhere about which AI projects actually succeed: 95% of AI Pilots Fail. Here's What the Other 5% Got Right. In short, AI is genuinely useful when you feed it large amounts of cleansed, well-filtered data, which is exactly why data engineers and AI engineers will only become more valuable. Used properly, it could deliver real benefits in modern medicine or save time reviewing legal documents.

 

The problem is that it isn't foolproof. Someone has to be watching it at all times. The moment bad information (like “the moon landing was faked”) gets fed into the model, it needs to be caught and removed immediately. Used to make people more efficient, AI is a genuinely wonderful tool. Left unmonitored, it's a disaster waiting to happen. If you're skeptical, read Fortune 500 Companies Let AI Make Fatal Mistakes.

 

Final Thoughts

I agree with part of Coxon's argument. AI is dangerous because it exposes security gaps faster than we can close them. However, tools like Databricks' LakeWatch are also closing the gap. Where I disagree is what that means. Every gap AI used in that campaign already existed before AI touched it. Strong security practices close exactly that kind of gap, which is why the campaign still failed against most of its thirty targets. Several U.S. states and European countries have already passed regulation to limit AI's reach. More is still needed.

 

In Illinois, for example, a chatbot is barred from performing psychotherapy or making clinical diagnoses. That work has to be done by a licensed professional, or the violators can be fined as much as $10,000 for each offense. It's a real rule, but it only covers one narrow use case. It does nothing about the security gaps AI is actually exploiting today.

 

You don't have to imagine what that looks like at national scale. In mid-2026, CISA confirmed that suspected Iran-linked actors used AI tools to help write scripts targeting Siemens, Rockwell, and Schneider Electric PLCs, and compromised more than one hundred internet-exposed water and wastewater systems across at least seven states. CISA said the intrusions had “little effect” on the actual water supply, but some of the compromised systems had their shutdown processes and alarms disabled. Security researchers who reviewed the incident traced the actual point of entry to cellular modems and PLCs left exposed to the internet on factory-default passwords that were never changed. Once again, the AI wasn't doing anything a determined human attacker couldn't do. It was moving through an already-known, already-flagged PLC vulnerability, faster and with fewer people required to pull it off. The idea of handing one AI system military passwords, nuclear codes, or the authority to act without a human in the loop still belongs in a Hollywood script. An Iran-linked group using AI to speed through a control-system vulnerability CISA had already warned about is not a script. It already happened.

 

If you're an executive reading this and you want to keep your company out of the next cautionary headline, start with what the GAO and CISA have already told everyone is broken. Patch your known vulnerabilities, fix credential hygiene, and monitor unencrypted traffic. Then put real governance policies in place so AI can't creep into data it has no business seeing. For more on that, read AI Governance in Healthcare.

 

Coxon isn't wrong that AI deserves scrutiny. He's just aiming the fear at the wrong target, and the media are not asking the right questions. It's not a rogue mind plotting our extinction. It just works a lot better at exploiting our current weaknesses. Fix the holes, and the machine has a lot less left to exploit.

 

References


Comments


bottom of page