Cool!! building games is a great way to test LLM. It covers reasoning, planning, coding, tool (lib) use, multimdal understanding, and more. But it seems like the community still lacks a standardized benchmark for building games?
This post is from a protected account.




