Hi all, I am a 3rd year undergrad at Stanford studying computational physics. I also lead agent evaluations at Browser Use.
I'm starting this X account as a place to voice my thoughts on AI and agentic developments
Grok 4.6 just released! But something is wrong, performance is 9pts worse on browser-use tasks (BU_Bench)
Also the cache read price was raised from 4.5, so it costs more. I am disappointed
I take it all back. We just got access to eval Grok 4.5 and it has landed above GPT-5.6-Sol and just shy of Opus for browser use.
Because cache input is expensive, the overall cost is only 10% cheaper than opus. Its overall a bit faster.
We have another opus-class model in the
I have finally closed the loop on bugfixing our cloud agent. Here is how it works:
> A user's agent tries to send a large email attachment, breaks it up into 20MB at a time as instructed
> But I made a mistake - I said input file size limit is 20MB but its actually a limit on