1. X
  2. Mayur Naik
Log inSign up
Mayur Naik
449 posts
Image
user avatar
Mayur Naik
@AI4Code
CS prof @Penn & founder @Rabdos_AI: creating data for the world’s hardest reasoning, from the chalkboard to the wet lab
Philadelphia, PA
cis.upenn.edu/~mhnaik/
Joined January 2019
360
Following
2,439
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Mayur Naik
    @AI4Code
    Apr 14
    Friends, followers, and strangers on X, I recently got excited about mapping the jagged frontier of AI models in math and other STEM areas. With my colleague Prof. Rob Ghrist @prof_g , I founded Rabdos AI rabdos.ai to create original, research-level problems at
    Image
  • user avatar
    Mayur Naik
    @AI4Code
    4h
    Check out Rabdos Math Index! A holistic benchmark for math x AI that draws from our experience supplying math data to frontier AI labs over the last six months. Comprises graduate-level theorems formalized in Lean and checked by semantic rubrics, research-level problems with
    user avatar
    Rabdos_AI
    @Rabdos_AI
    12h
    We are delighted to introduce Rabdos Math Index 📐 One holistic measurement for mathematical reasoning on proof, solution, and perception. Claude Opus 5 leads at 46. GPT-5.6 Sol & Claude Fable 5 follows at 39. No other model scored above 25. Mathematics remains open.
    Image
  • user avatar
    Mayur Naik
    @AI4Code
    May 14
    Incredibly proud of PhDs #8, #9, and #10 -- congratulations Drs. Jiani Huang (@jiani_huang_ai), Aaditya Naik (@aaditya_naik), and Adam Stein (@adamlsteinl)! I am lucky and grateful for the privilege of working with the three of you!
    Image
  • user avatar
    Mayur Naik
    @AI4Code
    May 12
    Very timely, especially in light of revelation that 1/3rd of problems in FrontierMath are fatally flawed. As expert human validation of frontier math tasks approaches its inevitable limit, LLMs are stepping in to fill the void. But our work below shows that the discovery of
    user avatar
    guru
    @guruprerana
    May 12
    Do we need frontier models to verify math proofs? EpochAI just announced that they found several fatal flaws in their FrontierMath benchmark using GPT-5.5. But isn't verification supposed to be easier than generation, so why were they not spotted earlier? In our recent work, we
    Plot comparing open-source models with frontier-models on proof verification.
  • user avatar
    Mayur Naik
    @AI4Code
    Apr 30
    Delighted to announce MathDuels, the first self-play math benchmark! We evaluated 26 frontier models across 780 generated problems from 30 math sub-domains. Check out mathduels.ai for the results, which we plan to update on a regular basis as new models enter the
    user avatar
    Rabdos_AI
    @Rabdos_AI
    Apr 30
    Static math benchmarks saturate. We built one that doesn't. Announcing MathDuels, the first self-play math benchmark. Every frontier LLM writes problems for the others, and is graded on the ones written for it. As models improve, so does the benchmark.
    Image
    00:00

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement