Pinned
CS prof @Penn & founder @Rabdos_AI: creating data for the world’s hardest reasoning, from the chalkboard to the wet lab
- Incredibly proud of PhDs #8, #9, and #10 -- congratulations Drs. Jiani Huang (@jiani_huang_ai), Aaditya Naik (@aaditya_naik), and Adam Stein (@adamlsteinl)! I am lucky and grateful for the privilege of working with the three of you!
- Very timely, especially in light of revelation that 1/3rd of problems in FrontierMath are fatally flawed. As expert human validation of frontier math tasks approaches its inevitable limit, LLMs are stepping in to fill the void. But our work below shows that the discovery ofDo we need frontier models to verify math proofs? EpochAI just announced that they found several fatal flaws in their FrontierMath benchmark using GPT-5.5. But isn't verification supposed to be easier than generation, so why were they not spotted earlier? In our recent work, we
- Delighted to announce MathDuels, the first self-play math benchmark! We evaluated 26 frontier models across 780 generated problems from 30 math sub-domains. Check out mathduels.ai for the results, which we plan to update on a regular basis as new models enter theStatic math benchmarks saturate. We built one that doesn't. Announcing MathDuels, the first self-play math benchmark. Every frontier LLM writes problems for the others, and is graded on the ones written for it. As models improve, so does the benchmark.
- What if you could see cardiac arrest coming minutes to hours before it happens? We're building CAMEL, a foundation model for cardiology trained on the ECG signals now captured everywhere from ICUs to wristbands.



