1. X
  2. fig
Log inSign up
fig
67 posts
fig profile banner
user avatar

fig

@figbrains
hyperintelligence for humanity
San Francisco 馃寔
fig.inc
Joined January 2024
2
Following
110
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what鈥檚 happening and join the conversation

Continue with phone
or
Log in with username or email
Terms路Privacy路Cookies路Accessibility路Ads Info路漏 2026 X Corp.
  • Pinned
    user avatar
    fig
    @figbrains
    Oct 31, 2025
    We're excited to announce MultiNet v1.0 - the first cross-domain benchmark for multimodal AI systems. Unlike existing evaluations that test models within single domains, MultiNet reveals what happens when AI systems encounter the full complexity of real-world tasks.
  • user avatar
    fig
    @figbrains
    Aug 26
    Try your hand at beating AI on maze solving in our new evaluation preview!
    user avatar
    harsh鉁岋笍
    @HarshSikka
    Aug 26
    You can solve simple mazes like this in 1 minute. Frontier models can do that too...right? While designing a new frontier evaluation, we were surprised to find complete failure on basic 2D mazes. Excited to share an early research preview, playable demo, and highlights below!
    Image
    00:00
  • user avatar
    fig
    @figbrains
    Jun 16
    GUI-DR applies domain randomization from robotics, varying visual scenes and instructions along controlled axes to expose fragile model behaviors.
    user avatar
    harsh鉁岋笍
    @HarshSikka
    Jun 16
    Computer Control models can score 90%+ on standard benchmarks, but will fail when you set page zoom to 70%. We're built GUI-DR, an OS pipeline that can restyle, reposition, and remove DOM elements on real webpages to reveal model weaknesses that fixed-scene benchmarks miss.
    Image
    00:00
  • user avatar
    fig
    @figbrains
    Jun 5
    Fig wants to directly support researchers working on foundationally new takes on frontier models - targeting hard problems like long horizon multi-environent action. Reach out to contact @ fig . inc if you're working on these or related areas.
    user avatar
    Pranav Guruprasad
    @pranavguru13
    Jun 5
    This week at #CVPR2026 we presented MultiNet v1.0 at the MMFM workshop. It is a benchmark built around a question most evaluations skip: what happens to a multimodal model when you take it out of the one domain it was trained for and ask it to handle everything at once?
  • user avatar
    fig
    @figbrains
    Jun 3
    Come meet the Fig team @ CVPR this week, today through Friday!
    user avatar
    Pranav Guruprasad
    @pranavguru13
    Jun 2
    Headed to #CVPR2026! I'll be there on behalf of @figbrains and @ManifoldRG, presenting our research on next-generation multimodal models and evaluation systems. If you're into multimodal models, VLAs, or how we actually evaluate them, come say hi - I'd love to talk!
Advertisement
Advertisement