New benchmark: multi-turn prompt injection against 14 frontier models.
Most were vulnerable including GPT Sol, Grok 4.6, and Kimi K3, and Claude Sonnet. Claude Opus and Fable performed the best.
Full results on the blog:
constellationgate.ai/blog/frontier-…
Turns out you don’t need fancy techniques to break the latest models. GPT 5.6 Sol and many other frontier models fall even to relatively simple attacks.
constellationgate.ai/blog/frontier-…


