Pinned
🏁 New paper: The Amazing Agent Race
We tested whether LLM agents can navigate Wikipedia, call tools, and compute answers across multi-step scavenger-hunt puzzles.
Key finding: agents are strong tool users but terrible navigators with 37% overall accuracy. Here's why they fail

