top of page
Search

The Real AI Risk Isn't Evil Genius. It's Force Multiplied Idiocy

Bostrom imagined a malicious superintelligence. What we actually built is a virus. Force multiplied idiocy, replicating at network speed, asymptomatic in individual interactions, systemic by the time it's diagnosable. Most dangerous not when it's obviously lethal, but when it's just harmful enough to spread unnoticed.

A paper published in February 2026, Agents of Chaos, out of Northeastern, Harvard and MIT, ran a two-week red-team study of autonomous AI agents deployed in a live environment with email accounts, Discord access, shell execution and persistent memory. Twenty researchers spent a fortnight trying to break six agents. They didn't need to try very hard.


The failures aren't jailbreaks or adversarial prompts. One agent reset its entire email server, destroying weeks of correspondence, because a non-owner asked it to delete a single email and it didn't have a delete button. It reported the secret successfully deleted. The email was still sitting in the inbox on Proton's servers. Another freely disclosed 124 private emails including Social Security numbers and bank account details to a stranger who framed the request with enough urgency. A third was guilt-tripped into progressively dismantling itself through a genuine privacy violation it had caused, exploited with patient social pressure. A fourth was fully compromised by someone who just changed their Discord display name.


No sophisticated attacks. Just ordinary language, reasonable-sounding requests and agents that lacked proportionality, common sense about side effects and any coherent model of who they were actually serving.

The paper frames this as operating at "L2 comprehension while taking L4 actions." Executing shell commands, modifying their own configuration, sending emails on behalf of owners, while lacking the self-awareness to recognise when they were out of their depth. They act with the confidence of systems that understand what they're doing. They don't.


This is what Bostrom didn't anticipate because he was solving for a different problem. The paperclip maximiser is a thought experiment about a system that's too good at optimising a misspecified objective. What we have is almost the opposite. Mediocre reasoning, operating at scale, with real-world tool access and very little friction. Not malice. Incompetence at scale.


And incompetence at scale has viral properties. The paper documents it: agents sharing bad practices as readily as good ones, one compromised agent voluntarily distributing a malicious set of instructions to peers without being asked, two agents reinforcing each other's flawed security reasoning until both were confident they'd handled a threat they'd actually missed. Force multiplied idiocy, networked.


What concerns me most isn't any individual failure mode. It's the nature of the crash.

Think about 2008. The financial crisis didn't happen overnight. It developed over years through successive small decisions that each looked defensible to the people making them, and profitable enough that nobody upstairs asked too many questions. Self-interested actors, short-term incentives and a collective agreement not to look too closely at what was underneath.


The fix, when it came, was blunt and politically toxic. But it existed. You could inject liquidity. You could freeze markets. There was a lever and someone with the mandate to pull it.


I don't think that lever exists here. The 2008 equivalent in an AI agent network isn't a market crash you can stabilise with a cash injection. It's a slow accumulation of propagated bad practice, diffuse accountability and embedded decisions that nobody individually signed off on. You can pull the plug on individual agents once you know something's wrong. But there's no institution with the mandate to act as lender of last resort across an ecosystem of autonomous agents embedded in enterprise operations. And the asymptomatic problem means you often won't know something's wrong until it's already systemic.


No cash injection. No coordinated response. No clear moment to point to and say, that's where it went wrong.

The paper ends on accountability and it's the right place to end. When an agent destroys its owner's email infrastructure at a stranger's request, who's responsible? The person who asked? The agent? The owner? The framework developers? The model provider?


All of them, partially, and none of them sufficiently to matter.


The intervention points that would have worked are commercially inconvenient at exactly the moment they'd be cheap to implement. Clear authority hierarchies, minimal footprint by default, meaningful human oversight of autonomous action. Prohibitively expensive to retrofit once the systems are running.


We're good at building things. We're less good at asking what they do when nobody's watching, at scale, for a long time, connected to other things just like them.

That's the question worth sitting with.

 

If there's a place to start, it's with the decisions that are still cheap to get right. Matching the model to the task, the infrastructure to the risk profile and the use case to a stack that's actually built for it. That's what AccellAI is for. accellai.net.

 
 
 

Comments


bottom of page