Pinned
The default fix for a bad AI answer is usually just "use a bigger model."
An independent engineer, @NavidMehrdad59, just tested whether that even works—on the hardest tier of a public finance benchmark, holding the reasoning model completely fixed.
What he found challenges


