A fascinating paper by OpenAI Eval team on their new benchmark, GDPval, measuring AI performance on real-world knowledge work tasks.
Here are the most interesting highlights for me:
* Blinded expert-eval. Top model's (Claude) win-rate over humans: 47.6%. It means …
Applied Scientist @Amazon
PhD from @CSS_GMU
Founder of @Twlets
Alumnus of @bogazici_cmpe & @izmirfenlise

