Cool!! building games is a great way to test LLM. It covers reasoning, planning, coding, tool (lib) use, multimdal understanding, and more. But it seems like the community still lacks a standardized benchmark for building games?
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for





