Pinned
CS prof @Penn & founder @Rabdos_AI: creating data for the world’s hardest reasoning, from the chalkboard to the wet lab
- Check out Rabdos Math Index! A holistic benchmark for math x AI that draws from our experience supplying math data to frontier AI labs over the last six months. Comprises graduate-level theorems formalized in Lean and checked by semantic rubrics, research-level problems withWe are delighted to introduce Rabdos Math Index 📐 One holistic measurement for mathematical reasoning on proof, solution, and perception. Claude Opus 5 leads at 46. GPT-5.6 Sol & Claude Fable 5 follows at 39. No other model scored above 25. Mathematics remains open.
- Incredibly proud of PhDs #8, #9, and #10 -- congratulations Drs. Jiani Huang (@jiani_huang_ai), Aaditya Naik (@aaditya_naik), and Adam Stein (@adamlsteinl)! I am lucky and grateful for the privilege of working with the three of you!
- Very timely, especially in light of revelation that 1/3rd of problems in FrontierMath are fatally flawed. As expert human validation of frontier math tasks approaches its inevitable limit, LLMs are stepping in to fill the void. But our work below shows that the discovery ofDo we need frontier models to verify math proofs? EpochAI just announced that they found several fatal flaws in their FrontierMath benchmark using GPT-5.5. But isn't verification supposed to be easier than generation, so why were they not spotted earlier? In our recent work, we
- Delighted to announce MathDuels, the first self-play math benchmark! We evaluated 26 frontier models across 780 generated problems from 30 math sub-domains. Check out mathduels.ai for the results, which we plan to update on a regular basis as new models enter theStatic math benchmarks saturate. We built one that doesn't. Announcing MathDuels, the first self-play math benchmark. Every frontier LLM writes problems for the others, and is graded on the ones written for it. As models improve, so does the benchmark.



