Your Mac can run a 26B diffusion LLM now📷
We just added DiffusionGemma 26B-A4B to Fast-dLLM-mlx, our open-source MLX inference framework for diffusion LLMs on Apple Silicon. Inference is ~30% faster than mlx-vlm 0.6.4 on M4 Pro, and the speedup is training-free (dual cache +
The research department at MacPaw.
On-device AI, LLM efficiency, AI memory, HCI.


