Handrolled implementation is now at 320 tok/s for the 0.6B-4bit variant, while plain Axon rewrites can get the bumblebee model to 100 tok/s on the non-quantized variant!
New Hex version shipping tomorrow.
In the past few days I shipped a proper EMLX compiler that leverages MLX tracing compiler, with support to all Nx constructs.
This was inspired by the other MLX-based backend. PR below has Qwen3 perf comparison. We can now reach 170+ tok/s!
github.com/elixir-nx/emlx…

