We're releasing HLE-Diamond: a refined version of Humanity's Last Exam, built with @CAIS.
A year of review and community feedback went into refining this subset to make it more reliable for measuring frontier models.
Top model tested 60.6% overall. We expect HLE-Diamond to
welcome to the lab.
from the researchers at @scale_AI



