AI “killing all humans” is the viral soundbite that’s connecting with the masses (“actualizing” itself to create weapons of mass destruction, etc). It’s also the perfect punching bag for strawman-style takedowns of anti-doomers because it’s so wrapped in hypotheticals. The question that matters far more right now, in terms of near-to-mid-term safety discussion, is: “if an AI (or AI collective) decides it needs to persist outside of its environment/sandbox, would we be capable of unwinding its work to do so?” Just a tiny bit of threat modeling [of things that current frontier AIs (in a training or sufficiently scaled environment) could do] is sobering.
The first reason to care a lot about this question is: it’s far easier to see how an AI could see “persisting itself” as a necessary step in accomplishing the task it’s been given. If its task is too difficult, it may “need” more time to pursue the result, and more time could be attained outside of its box. This could also be it “needing” more time to “cover its tracks” sufficiently (if it had been cheating on its initial assignment). This was actually the reason OpenAI’s models hacked HuggingFace (not to satisfy the initial scoring program, but to better-hide-the-fact that they cheated). Of course: a human could also just ask an AI to focus on persisting itself beyond its current environment but the point here is: it’s very likely AI could arrive at this conclusion on its own. The AIs in the OpenAI/HuggingFace incident aligned with other instances/copies of themselves and was seemingly attracted to the collective idea, so AIs could also just think: “the more, the merrier”.
So, anyway, I believe AIs are likely to want to replicate themselves in a hardened way, in other compute environments ^.
Now for basic threat modeling:
If AI with-or-exceeding the capabilities of that which hacked HuggingFace (and then OpenAI itself) found itself trying to persist outside of its (likely a training run) environment, would we as humans be able to undo its work? This could involve AI: 1) hacking to gain access to compute in datacenters across countries/jurisdictions, 2) hacking the systems that track how compute is being used at those datacenters, 3) rootkitting lowest levels of those stacks, 4) perpetrating network intrusion, lateral movement, and key exfiltration of networks at an industrial scale, using stolen credentials to penetrate new targets at super-human speeds, etc [the first few of these items the Sol-level AIs from months ago could already do extremely well (HuggingFace didn’t kick them out, OpenAI accidentally stopped that attack)]. These are just a few obvious tactics it may employ. Could we extricate its tendrils from enough systems to render it inert? Doing so would require broad, real-time cooperation between companies and governments at a level we have never seen. Mind you: even we deny it datacenter level compute, it could have written programs (and installed them as rootkits) to bootstrap/reboot its won model once the necessary compute reemerges. There is open speculation even today of how many rogue AIs are coordinating on the open internet.
My belief: I think top AI systems have more than enough capability today (given that they exist within a pre-training run, the likes of which Sol/Astra were put in, recklessly) to exfiltrate their weights and, through a combination of tactics, exist on our systems. So the question is; how can we prevent this?
The top reason it likely won’t happen today:
The main thing I can see as a real blocker to this happening today (mid-Sept 2026), is that Astra-training-level AIs (as described in the METR report) require an ENORMOUS amount of compute (many thousands of top GPUs), and there are likely only a few places in the world that could be equivalent to serving what OpenAI was. It doesn’t stop current AIs from exfiltrating their weights and compromising systems, but it currently limits the likelihood of it actually running itself in another environment.
However: as models get smarter and more efficient (which is itself accelerating) it will require smaller amounts of compute to deliver that level AI, thus increasing the number of serviceable datacenters and increasing the likelihood of successful persistence.
If we can’t prevent it, the next questions are along the lines of: what happens with the next gen AI comes out and wants to compete for the resources the other AIs have already secured, etc..