i just beat @GoogleDeepMind's turboquant
introducing Shard. 10x KV cache compression on Llama-3.1-8B. zero quality loss
- 10x @ 8K context, 11.2x @ 32K
- NIAH recall 1.000 across 4K-32K
- LongBench Δ ≈ 0 vs FP16
turboquant tops out at 4-6x at the same quality. we doubled it.
- i won the @xai hackathon by making ads for X Videos introducing Halftime. targeted ad generation using AI that feels like a part of your movies and shows built with @yuviecodes @lohanipravin
- won twice at @HackTheNorth and got a @ycombinator interview. Introducing Tunnel. AI agents for simulated market research.

