Do harness evolution methods really find better harnesses or are they yet another way to improve performance by using more inference time compute? @yikewang_ and @TengX6 have an answer.
- Check out the super cool web interface of DR Tulu!Super excited to share our open interactive demo for DR Tulu-8B! It supports web and literature search with full transparency — you can see the model's thinking traces and tool outputs as it reasons through your query. 🔗 dr-tulu.org 📝 arxiv.org/abs/2511.19399
- I have been trying out Telugu queries on the Indic LLM Arena over the last few days and most of the responses are surprisingly bad, with lots of hallucinations and sometimes even grammatical errors, even from strong (in English) models. Clearly there is a huge gap between EnglishFor AI to be truly inclusive, it must understand more than just grammar—it must understand context. @ai4bharat at @iitmadras had launched the Indic LLM Arena. This isn't just another leaderboard; it’s a public utility for: ✅ Developers: Test your models against real-world
- Here's a really quick follow up to our recent big Olmo 3 release! ✨Olmo 3.1✨: - a longer RL'ed 32B Think, with stronger math and reasoning skills - a 32B Instruct, a larger Instruct model also with function callingOlmo 3.1 is here. We extended our strongest RL run and scaled our instruct recipe to 32B—releasing Olmo 3.1 Think 32B & Olmo 3.1 Instruct 32B, our most capable models yet. 🧵





