We've been pushing search to uncomfortable scales @turbopuffer: 100B vector indexes
To get there, we had to rethink how search algos map onto hardware. The key insight was balancing bandwidth demands across the memory hierarchy
I wrote about the approach and the math behind it
tpuf ANN v3, for when you need to index the entire web
100B+ vectors @ 50ms p50 / 200ms p99 latency
blog: tpuf.link/ann-v3



