Evaluations.
I am currently a member of technical staff at Arena, working on evals—specifically large-scale, real-world agentic evaluations. I’m particularly interested in measurement, and in how to turn measurements into future improvements in capability.
Before Arena, I finished my master’s and bachelor’s at UC Berkeley. During this time, I worked on reward model training and evaluation. I’ve been lucky to be mentored by Ion Stoica, Jiantao Jiao, Banghua Zhu, Anastasios Angelopoulos, and Wei-Lin Chiang.
I also worked at Nexusflow, where I trained open-weights LLMs like Athene-70B and Athene-V2-Chat.
Most generally, I want to help build better AI for humans.
Auditing millions of claims made by LLMs in Text and Search Arena battles to construct factuality-aware leaderboards.
Building the methodologies behind the Agent Arena leaderboard.
Evan Frick*, Connor Chen*, Joseph Tennyson*, Tianle Li*, Wei-Lin Chiang*, Anastasios N. Angelopoulos*, and Ion Stoica. (2025).
Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang, Anastasios N. Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. (2024).
Banghua Zhu*, Evan Frick*, Tianhao Wu*, Hanlin Zhu, Karthik Ganesan, Wei-Lin Chiang, Jian Zhang, and Jiantao Jiao. (2024).
Tianle Li*, Wei-Lin Chiang*, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. (2024).
Evan Frick*, Peter Jin*, Tianle Li*, Karthik Ganesan, Jian Zhang, Jiantao Jiao, and Banghua Zhu. (2024).
Evaluations.
Chatbot Arena and reward models.
RLHF and function calling.
RLAIF and synthetic data. Advised by Professor Jiantao Jiao.
Change tracking and analytics system.
ViT robustness in dermatology.
Secure server allocation systems.
Made buttons button, clicks click, and moves move.
M.S. in Electrical Engineering and Computer Science.
B.A. in Computer Science, Minor in Data Science, Highest Distinction in General Scholarship.
Email: evanfrick[at]berkeley[dot]edu
Google Scholar: Evan Frick
LinkedIn: Evan Frick
Twitter: @evan_a_frick