really interesting to see array runtime inference with mlx and mlx-lm for oss local search. turning your own battery power into intelligence. 🙂
Replying to @perplexity_ai
Hybrid compute splits work between cloud models and a local model on the Mac. Local inference must keep pace with the rest of the task.
MLX-LM is a general-purpose framework, while Lily is purpose-built for this inference workload.




