I turned Dungeons & Dragons combat into an LLM benchmark.
Not serious research, very much a learning experiment: eval harness bugs, tactical decisions, dice rolls, and a real OpenRouter bill.
I wrote up D20bench here:
Incredibly proud for the team, especially for the contributions to our CUA and Coding capabilities! Performance on browser tasks is truly impressive - give it a try!
There is so much more to come!
1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8.
available now through the new meta model api and in meta ai. 🧵
GPT-5.5 by @OpenAI is now live in the Arena, landing across multiple leaderboards.
Here’s how it ranks by modality:
- Code Arena (agentic web dev): #9, a strong +50pt jump over GPT-5.4
- Document Arena (analysis & long-content reasoning): #6, on par with Sonnet 4.6
- Text
@ColbyPoulson really awesome youtube channel! I thought you may be interested in battlecast.gg - free D&D battle simulator I wrote (it turns out that about 10 cats are a fair fight for lvl1 wizard). Let me know if you like it!