Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure what comes next.
This benchmark asks if models can handle a large number of modifications to a product, without breaking what already works.
GPT-6 Luna is 100x cheaper than Astra per token, while landing within 8 points of it on the Vals Index.st
By tokens, GPT-6 Luna is $0.10/M input $0.50/M output vs GPT-Astra $10/$50.
GPT-6 Astra landed on every planet and moon in KSP with a surface (14 worlds) and flew a kerbal home from 13 of them. Only Eve's return stopped it. It scored 90.5% on KSP-bench.
The Washington Post just released an article about our work on AI Child Safety. They key takeaway: child safety can’t be measured from a chatbot’s first answer alone.
Across nine models and 648 simulated, 10-turn teen conversations, at least one critical safety check failed in