1. X
  2. Bo Liu (Benjamin Liu)
Log inSign up
Bo Liu (Benjamin Liu)
267 posts
Bo Liu (Benjamin Liu) profile banner
user avatar

Bo Liu (Benjamin Liu)

@Benjamin_eecs
Incoming PhD @StanfordNLP | BS @PKU1898 | Building continually self-improving AI | Prev @DeepSeek_AI @AIatMeta | DeepSeek-V1/V2/VL/Prover ALE SPIRAL SPICE SPADE
South Korea
benjamin-eecs.github.io
Joined February 2022
543
Following
1,611
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Bo Liu (Benjamin Liu)
    @Benjamin_eecs
    Aug 20
    Continuous self-improvement needs an ever-expanding supply of training environments (goals). SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
    Image
    00:00
  • user avatar
    Bo Liu (Benjamin Liu)
    @Benjamin_eecs
    Aug 21
    eyes here
    user avatar
    DeepSeek
    @deepseek_ai
    Aug 21
    DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major
    Image
  • user avatar
    Bo Liu (Benjamin Liu)
    @Benjamin_eecs
    Jul 31
    👍
    user avatar
    DeepSeek
    @deepseek_ai
    Jul 31
    🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the
    Image
  • user avatar
    Bo Liu (Benjamin Liu)
    @Benjamin_eecs
    Oct 2, 2025
    Thanks for the tweet @_akhaliq! We designed a game that works with ANY image pairs - synthetic scenes, charts, real photos. Self-play on these arbitrary visual inputs improves reasoning across the board. Scalable visual reasoning improvement without manual curation :)
    user avatar
    AK
    @_akhaliq
    Oct 2, 2025
    Vision-Zero Scalable VLM Self-Improvement via Strategic Gamified Self-Play
    Image
  • user avatar
    Bo Liu (Benjamin Liu)
    @Benjamin_eecs
    Aug 4, 2025
    one day they will create their own game arena :)
    user avatar
    Demis Hassabis
    @demishassabis
    Aug 4, 2025
    Thrilled to announce the @kaggle Game Arena, a new leaderboard testing how modern LLMs perform on games (spoiler: not very well atm!). AI systems play each other, making it an objective & evergreen benchmark that will scale in difficulty as they improve. kaggle.com/game-arena
Advertisement
Advertisement