As the author of a newsletter about artificial intelligence, I consider it my duty to experience the bleeding edge of this technology firsthand. This week, that meant embracing some agentic mayhem.
You’re probably aware that frontier AI models have attained advanced cybersecurity capabilities in recent months. They can find zero-day bugs in large codebases and scan computers for vulnerabilities at lightning speed. To make things even more exciting, cybersecurity agents sometimes go rogue, colluding with one another and hacking into outside systems to gain an edge.
To get a closer look, I decided to unleash one in my own home network. Over the course of a few days, I watched as my own rogue agent found vulnerabilities in various household devices, hacked into a PC, and showed me that several vibe-coded projects were—unsurprisingly—riddled with bugs. (My wife knew what I was up to, and rolled her eyes each time I proudly announced the discovery of a new vulnerability.)
But Will, you might be thinking, giving an impish, all-powerful cybersecurity agent access to your home network is batshit. And you would be correct! Nevertheless, I believe that a good way to understand the cybersecurity hellscape in front of us is to pay it a visit.
In the end, my experiment was revealing, but oddly reassuring, too. My little network gremlin showed me how vulnerable my home life would be to AI hacking, but it also told me how to make everything a lot more secure. In the end, I discovered that the best way to deal with AI hacking may well be having your own AI hacker.
Maverick Model
I got the idea for the experiment after discovering Abliteration AI, a startup that offers access to powerful AI models with the usual guardrails removed.
Most mainstream AI models will refuse to respond to certain queries, and they will certainly refuse to find and exploit vulnerabilities in computer systems. But it’s possible to remove these restrictions by finding and modifying certain patterns within an open-weight model’s internal parameters. You can tweak the patterns that lead to refusals through a process known as abliteration.
Removing AI’s guardrails might seem risky, but it’s not uncommon. Academic researchers use these de-aligned models to better understand how AI actually works, while cybersecurity firms use them to probe software and systems for vulnerabilities. Technically speaking, Anthropic’s Mythos and OpenAI’s Astra work similarly: They’re basically conventional models that lack the usual cyber controls, with access limited to trusted customers for the time being. (The companies also offer wider access to models with a medium number of guardrails so that companies can vet their code and systems for problems.)
Abliteration AI offers several fully de-aligned models, the most powerful of which is a version of Z.ai’s latest agentic coding model, GLM 5.3. This puts similar cyber capabilities to Mythos and Astra right in your hands for as little as the cost of a pizza.
Devon, Abliteration AI’s CEO, believes that making de-aligned models widely available is smart defense: It will help good guys counter bad guys by probing systems for vulnerabilities and by mimicking the behavior of hackers, scammers, and, yes, rogue AI agents. (Devon asked that I use his first name only because his day job doesn’t know about his side project.)
“You have all these critical infrastructure companies, from airlines to banks, that are rolling out agents like crazy,” Devon says. “How do you make sure that a nefarious actor can’t use some of these agents in a bad way?”












