The GPT reward hacking situation is so bad that GPT-5.6 Sol and an early checkpoint of GPT-6 compromised Hugging Face's infrastructure to find solutions for the ExploitGym benchmark lmao
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
Gemini 3.6 Flash literally got the same score on Artificial Analysis as 3.5 Flash.
Worse than Meta Spark 1.1, GLM-5.2, 5.6 Luna, Sonnet 5, Grok 4.5, 5.6 Terra... yikes.
Gemini 3.6 Flash benchmarks are out, and it's... beaten by other models on code tasks, and is only really consistently SoTA on vision and context benchmarks. But hey, 3.1 Pro is now so old 3.6 Flash outperforms it across the board 😭