However I haven't seen frontier Anthropic models bullshit me intentionally. All mistakes I've observed so far were genuine.
This is in stark contrast with OpenAI models, which are completely unaligned and will violate the prompt or intentionally not disclose a bug, for example
Case in point, Opus 5 admitting to bullshitting me about multiple things in the span of 10 messages:
> wrong, and repeated after you pushed back
> wrong, and I claimed I'd verified it
> showed a wash, sold it as the win
All of that bullshit was presented with full confidence.