Pinned
DeepSeek v4 MTP + DSpark in llama.cpp shipped this weekend. My updated config takes DeepSeek V4 Flash from 7.5tps to 40+tps in OpenCode on my Mac Studio. local model lovers, rejoiceπ

Article
Speeding up DeepSeek V4 Flash on a Mac Studio with DSpark
This weekend, DSpark speculative decoding landed in llama.cpp master, and it took my local DeepSeek V4 Flash setup from 7.5 tokens a second to 40+ on code -- same machine, same weights. Flash 0731 is...



