A 70B model at FP16 needs about 140GB of VRAM before you've loaded a single token of context. Quantization cuts that 2 to 4x.
H200 reaches 31,712 tokens/s on Llama 2 70B, about 42% above H100's 22,290. Both numbers matter less than what's driving them: LLM inference is
One platform, every GPU cloud. Deploy H100s to B300s across multiple providers. On-demand, no lock-in, from a single dashboard.

