Deepseek v4 Flash 4.2bpw has been running stable for 2 days on our LLM box thanks to the cheap ADT-Link PCIE > NVME adapter. And it's frickin' amazing.
I decided to try to find the limit of it yesterday because i needed to do a multi hour stability test of the OC. So i went to
PHP & simple coding apologist, UI Designer, Fractional CTO, Programmer, Server admin, Bootstrapper, Low latency Enjoyer.
Creator of Zerolith, zl.css, y más
Utah, USA
Joined November 2022
- Whoo-wee! Got one of those cheap NVME to PCIE adapters ( ADT-Link ) and plugged a 4070 into my LLM box for 140gb of vram total. I ran a 4.25 bit quant of Deepseek V4 Flash, and the penalty is going from peak 77 tok/sec to 67 tok/sec, not bad! Despite having 15% less bits than
- I kinda want to make a user interface based around this color scheme, what you think?
- 3TB/sec? damn, that's faster than the ram on my RTX PRO 6000 🤩
- i'm officially eating my hat. I used to talk crap about Zed for being much too basic, but in recent versions, more creature comforts are there & i can see the potential of this leapfrogging VS Code at the very least. The agent is good ( significantly better than the new ACP


