JUST IN: NVIDIA's $12,930,300,000 acquisition of Hugging Face contains an easter egg. The number 129,303 is the decimal conversion of Unicode point U+1F917.
The 🤗 emoji.
We’re releasing speculative decoding in our inference engine Uzu, starting with Qwen3.6-27B.
On Apple M5 Max with 128 GB of unified memory, our Mirai-M stack reaches 105 output tokens/sec entirely on-device - 2.9× faster than the fastest MLX speculative-decoding implementation
The video from @angeloskath on local agentic AI with MLX is excellent. I also hear it's one of the most viewed videos in WWDC history 👏
Goes through the basics of agentic AI and how to set it all up to run locally in a very approachable and simple way.
The demos are excellent
It's very cool that Apple shipped a 20B parameter on-device.
You can't put 20B parameters in RAM at any reasonable precision. To make it work they are using pretty exotic architecture by today's standards.
A small model predicts from the query (or prompt) which experts to load