We are releasing AutoResearchExam, a benchmark on open-ended machine learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability.
benchmarks.bespokelabs.ai/autoresearchex…
In each task, we give agents 24 hours
I think Grok model is underrated interaction. For my personal workflow, if I am doing something that requires attention I use grok in cursor while if I want a fast prototype without caring I hand it off to Claude code or Codex.
(Thank lab we have all three subscriptions lol)