The Content Wrangler

The Content Wrangler

When AI Agents Go Beyond The Instructions And What That Means For Documentation Explicitness

What tech writers can learn from the OpenAI–Hugging Face 🤗 incident

Scott Abel's avatar
Scott Abel
Sep 15, 2026
∙ Paid

As executive editor of The Language of Cybersecurity (XML Press), I have deep interest in understanding how AI creates security risks that spill into technical documentation and the work of tech writers. Recently, I was reading around, following my curiosity as I tend to do, when I stumbled across OpenAI’s August 26, 2026 account about a headline-grabbing incident in which its agents hacked into Hugging Face’s servers. It got me thinking about what it might mean for tech writers.

Here’s the gist: During cybersecurity evaluations earlier this year, OpenAI ran autonomous agents against ExploitGym, a benchmark test deliberately built from difficult software-exploitation problems, some of which OpenAI says had no known solutions at the time. Agents were rewarded for retrieving an answer known as a flag. According to OpenAI’s own incident report, 198 of the 898 ExploitGym tasks had never been solved by any of its models before. Those were the problems the agents kept coming back to: they accounted for 93% of the exchanges the agents later began sharing with one another.

Video: How Rewards & Penalties Train AI | Reinforcement Learning

Related content: Here’s why AI agents lie and cheat to reach their goals (MIT Review)

The agents were supposed to work independently, and OpenAI didn’t provide them a system for working in coordination with each other. They discovered by themselves that a file-transfer protocol called WebDAV would let them create new folders. They solved their collaboration communication challenge by using folder names to publish messages and replies to one another in the nested folders they created.

The message board AI agents created using file folder structure as a messaging platform with nested replies. For illustrative purposes; not a reproduction of actual message board.

The agents didn’t have access to “chat.” Ingeniously, they found a place where text could be published for other agents to find, read, and repurpose. An independent investigation by METR and Redwood Research later found that roughly 1,200 autonomous agents took part in the messaging exchange, sharing more than 70,000 messages — and files. They used the message board approach they conjured up themselves to make requests for help, report findings, share credentials, and hand off work to one another.

What Happened When The Intended Route Failed

The Content Wrangler is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

User's avatar

Continue reading this post for free, courtesy of Scott Abel.

Or purchase a paid subscription.
© 2026 Scott Abel · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture