The NetHack Learning Environment (@NetHack_LE) has been around since 2020. Since 2024, balrogai.com is evaluating frontier LLM performance on it. We are slowly seeing progress on it, but it still seems far from being solved.
The NetHack Learning Environment (@NetHack_LE) has been around since 2020. Since 2024, balrogai.com is evaluating frontier LLM performance on it. We are slowly seeing progress on it, but it still seems far from being solved.
BALROG’s leaderboard has three new entries, courtesy of @creus_roger
Very interesting that GPT 5.6 Sol at max effort is still within error of Gemini 3 and 3.1 pro. GPT Astra 6 on max reasoning effort however reaches new heights, and a @NetHack_LE avg. progression of 13% 🏰