- Cannot fucking believe that the "new" AI model does worse than the old. What is anthropic doing? Same AI Lab that created Mythos btw.Claude Opus 4.7 just regressed hard on BridgeBench. Bullshit Benchmark tests if models push back on nonsense or just make things up. Claude Opus 4.6: 95.0. Rank 1. Claude Opus 4.7: 75.5. Rank 5. Claude Opus 4.7 accepts made up jargon 24% of the time. Claude Opus 4.6 accepts


