Real world numbers for DSv4-Flash-0731 on Freetoken running 1x3090 + 256GB DDR4 2400 TR 3945wx, hits 10.5t/s on tg 👀 Compared to of 4x 3090's, 1x 4090 and 1x 5060ti in Llama.cpp hitting 25t/s tg on the DSv4F-0731. 🧵
this is a critically important moment for cyber defense with AI; there is not much time to act.
we are happy if you want to work with us or any of our competitors or partners, but please take this moment seriously.
only an urgent and intense collective response will work.
for computer things or for connecting sensors and such? one of these with a 7th gen and newer cpu can run your whole homelab. rpi is good if you need it to be actually battery powered/mobile or the gpio.
New blog! 🚀 MTP, EAGLE-3, DFlash or DSpark, which speculative decoding method should you actually use?
There’s no universal winner. The best choice changes with the model, workload, and speculation depth.
We break down how 5 methods work, how to enable and tune them in vLLM,