Pinned
From OSWorld 1.0 to 2.0, we went from minutes (~30 steps) to hours (~318), from single apps to real workflows, from high scores (83%) to hard problems (21%).
1+ year, 20+ people, every task rigorously verified. This is what real cua evaluation takes.🙏
👉osworld-v2.xlang.ai
Two years ago, we built OSWorld 1.0 — the benchmark that became the standard for computer-use agents. Agents now score 83.5% on it. Problem solved?
Not even close.
🚀Today we introduce OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks.
What's new:






