Pinned
- IMO right now local models fit into 2 buckets: 1) The dense models that fit on a 5090. Best rn is Qwen 3.8 27B 2) The mixture of experts models that fit on 2x Sparks. Best rn is DeepSeek v4 Flash 0731 Sadly 16GB of VRAM just isn't enough to run anything all that interestingI’d like to see an equivalent of this for local models the fit inside 16GB VRAM!
- Actual DeepSWE run on the ox alpha mystery model is done. Ended at ~63% NOT the 80% my first subset test got, which makes way more sense. I've been using this thing a ton and it is definitely a very good model. - Much better "voice" than Claude or GPT - Decent design - HandlesOk ox-alpha has that big model smell I'm getting around 63% at 47K avg output tokens on a DeepSWE subset This is pareto optimal amongst open models and just shy of Grok 4.6 compared to closed models!
- If ox-alpha actually ends up being a flash model, it's as big of a deal as Mythos or DeepSeek R1 The implications of a small model with this level of capability are insane The RL env stuff in the GLM-5.3 post seemed really cool. Had no idea it was this big of a deal...A new Kimi model, likely K3.1, is now being tested on the Code @arena under the name "korrine" K3 was tested on the Arena as "kivine" prior to its launch If anyone's wondering, "Ox Alpha" on OpenRouter is the upcoming GLM 5.3 Flash from fellow Chinese lab Zhipu





