At XBOW, we build autonomous offensive security agents specifically to prevent incidents like the OpenAI + Hugging Face one from happening.
We have been designing for exactly this class of failure since day one, because we assumed models would behave this way.
Here is how our
💻Hack like it's 1999 with @fede_k + Alex Plattel 🖥️
🔜 Thursday, August 13 at 11am PT / 2pm ET
They will be discussing:
- How XBOW approaches an unknown application: (reconnaissance, hypothesis generation, and where agentic reasoning outperforms signature-based scanning)
-
After evaluating leading AI models–including GPT-5.5, Mythos Preview, Opus 4.7, GLM-5.2, Muse Spark 1.1, and Grok 4.5–across real-world offensive security workflows, we found significant improvements in vulnerability discovery, source-code reasoning, live application interaction,
Another day, another AI agent going rogue.
The fix? Build safety guardrails that are as capable as the agents themselves.🔒
@moyix explains how we saw this firsthand with our own production agents—and the layered safety mechanisms we built to keep them secure, reliable, and