1. X
  2. Tao Yu
Log inSign up
Tao Yu
515 posts
Tao Yu profile banner
user avatar

Tao Yu

@taoyds
@XLangNLP lab, asst. prof. @HKUniversity. author of OpenCUA, OSWorld, Aguvis, Spider, OpenAgents, Text2Reward, Instructor.
Seattle
taoyds.github.io
Joined March 2016
928
Following
6,324
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Tao Yu
    @taoyds
    Jun 26
    From OSWorld 1.0 to 2.0, we went from minutes (~30 steps) to hours (~318), from single apps to real workflows, from high scores (83%) to hard problems (21%). 1+ year, 20+ people, every task rigorously verified. This is what real cua evaluation takes.🙏 👉osworld-v2.xlang.ai
    Image
    Image
    00:22
    user avatar
    XLANG NLP Lab
    @XLangNLP
    Jun 26
    Two years ago, we built OSWorld 1.0 — the benchmark that became the standard for computer-use agents. Agents now score 83.5% on it. Problem solved? Not even close. 🚀Today we introduce OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. What's new:
  • user avatar
    Tao Yu
    @taoyds
    Jul 4
    Excited to have you @StevenyzZhang join HKU!!!🔥🎆🎉
    user avatar
    Yanzhe Zhang
    @StevenyzZhang
    Jul 1
    EXCITED to share that I’ll be joining the University of Hong Kong (HKU) as an Assistant Professor in January 2027! I’m interested in AI agents 🤖, Human-AI Interaction 🤝 and their safety🛡️. Email me if you’d like to chat! I’ll also be at ACL from July 3–6!
  • user avatar
    Tao Yu
    @taoyds
    Jun 24
    Robots can pick&place well now, but ask them to push sth over, drag it closer, or roll it — and they fail! The gap isn't task difficulty; it's that VLA embodied instruction following is brittle and doesn't generalize to how to execute. @Erics_Tong's work tackles exactly this!👇
    user avatar
    xintong hu
    @Erics_Tong
    Jun 24
    Current robot policies overfit specific language templates, handling 'pick and place' but freezing on 'drag it to me ' or 'push it closer to me.' They also lack control over execution: which hand, what approach angle, where to grasp, which path to follow. 🤖 FineVLA make robots
    Image
    00:00
  • user avatar
    Tao Yu
    @taoyds
    May 26
    We've seen nice recent progress on Scaling CUA RL envs/tasks — but 🎯verifiable rewards🎯 for RL training have been largely missing, and that matters a lot for preventing reward hacking. @BowenWangNLP's work tackles exactly this: 32K+RLVR tasks across 110 envs, check it out! 👇
    user avatar
    Bowen Wang
    @BowenWangNLP
    May 26
    RLVR has become the recipe for agentic post-training. But for Computer-Use Agents, the bottleneck is not the algorithm, it is the data. 🐌 🚀 We introduce CUA-Gym: a scalable, lightweight synthesis engine that turns arbitrary task queries into verifiable RLVR data for
    Image
    00:00
  • user avatar
    Tao Yu
    @taoyds
    Apr 22
    Big congrats to @ysu_nlp and the team on the launch of @NeoCognition 🎉!! Looking forward to seeing more great agent work from the team soon.
    user avatar
    Yu Su
    @ysu_nlp
    Apr 21
    Introducing @NeoCognition, the agent lab for specialized intelligence. Everyone needs experts, but human expertise does not scale. Backed by $40M seed funding, we build self-learning agents that specialize across domains to make expertise abundant.
    Image
    00:00

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement