The adoption of open weights is accelerating, but how do we securely deploy them?
We've assembled a framework for assessing and mitigating risks in open weight model deployments based on our experiences post-training with frontier enterprises.
Gap analysis is how our Applied Researchers, like @_brylee10, surface agent failures across billions of tokens of RL traces for our customers. Using our platform, AC2, and low-cost classifiers like Jev, we can catch 85% of failure modes at a fraction of the cost of LLM judges and
I implemented a system in @appliedcompute’s platform for automated failure mode clustering with Jev to surface errors at an even larger scale than before.
RL training produces billions of tokens in traces. I always manually read many traces to understand model behavior, but
A 35B open-weight model trained to search a precomputed index answers repo search questions at 100x lower cost than a frontier model.
We partnered with @turbopuffer to train Qwen3.6-35B-A3B to find code across ~9,000 repositories. It tops the needle-in-a-haystack task outright
“50% of DoorDash’s agentic restaurant orders are going to places users have never ordered from before.”
@andyfang tells our CEO @ypatil125 what happens when agents become the discovery layer. If models increasingly decide what gets surfaced and bought, companies have a strong