This is why local models must win.
Cloud models refused to reverse-engineer Apple's RDMA protocol.
DeepSeek and Nemotron said, "hold my beer".
TBF, Nemotron needed KV Cache injection to comply lol.
The next-gen architecture powering Qwen4 is now here! ✨
Get ready for the open release of Qwen3.8-Flash-Next 🚀
The countdown starts now! ⏳🔥modelscope.cn/models/Qwen/Qw…
A fun compare:
- M5 Ultra Mac Studio with 256GB memory, 8TB storage is $14,299
- NVIDIA RTX PRO 6000 with 96GB is street price $14,999. Sill need a $3K to $5K system to run it.
The NVIDIA card supports full CUDA and will be 2-4x faster at prefill, but for many people the Mac