This benchmark for Ai scientific capabilities is beautifully thought out - I especially like it's clear enumeration of design principles...
Computer science prof, entrepreneur & leader at Ai2. Excited by AI for science, human-AI interaction, and Web-scale NLP.
- Truly open scientific question answering - that's good!
- I'm so excited by this! Our system is generating some insightful & novel theories (e.g., internally for LM post-training). And it's still getting better!
- Smart analysis analysis of scholar output when authors adopted LLMs as part of their writing: 1) huge 36% boost in # papers published 2) LLMs mitigate skill disparities, eg native language - enough to shift market share of production toward China bit.ly/4qliJGo @yian_yin
- Impressive deep-research performance by a tiny & open model!🔥Thrilled to introduce DR Tulu-8B, an open long-form Deep Research model that matches OpenAI DR 💪Yes, just 8B! 🚀 The secret? We present Reinforcement Learning with Evolving Rubrics (RLER) for long-form non-verifiable DR tasks! Our rubrics: - co-evolve with the policy model -



