Delighted by the evaluations @chooi_jeq and @robocurve have been doing over the past months. It would have been so audacious to ask this of general-purpose models running on someone else's robots a year ago, but now we can...Just try it?! Glad this initiative is happening.
GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵
00:00



