📣 We've just published our latest research on state-of-the-art speculative decoding in interactivity.
Introducing a new approach that achieves a 4.37× speedup over autoregressive decoding in small batch setting and outperforms the strongest tuned DFlash baseline by 24.7%.
Frontier on-device AI lab. Models, runtime & infrastructure to make
on-device AI interactive, ambient & continuous.
- Gemma 4 architecture analysis thread Just as Gemma3n, this thing has a galaxybrained architecture, very much not a standard transformer
- If you've implemented speculative decoding, you've run into this: your speculator predicts the distribution correctly and still gets rejected. Let’s take an example: "The random number from 1 to 10 inclusive is" Ideal LLM: uniform 1/10 across all numbers. Good speculator: same
- Replying to @liquidaiDay 0 support across the stack: > Hardware: @AMD, @Intel, @Qualcomm > On-device: @lmstudio , @Cactuscompute, @RunAnywhereAI , @zeticai_ , @trymirai > Customization: @distil_labs
- LFM2.5-350M is now available on Mirai. @liquidai smallest model outperforms Qwen3.5-0.8B on reasoning and agentic tool use. Running on Mirai in full precision, it exceeds 70 tokens/second on iPhone.



