The MLOps community is an open and transparent community where all are welcome to participate. It is a place where MLOps practitioners can collaborate and share
"As a company we are spending $2 per million token. What does it mean? It doesn't mean anything." Kuntal Patel and Abhinav Lad of Palo Alto Networks on why per-token cost is a weak KPI, and what they measure once agents start looping on their own.
Ambud Sharma of @Pinterest maps AI efficiency to five layers: hardware, capacity, inference engine, model and quantization, and routing and governance. The first two are immutable once bought, so no tuning higher up recovers a wrong order.
Drew Breunig on labs training their harnesses into the models: system prompts shrink each release because last cycle's hotfixes get post-trained in. Reliability up, diversity down. Anything that doesn't look like Claude Code fights the weights.
Billing alerts run a day or two behind. Agentic loops and retries can burn a budget faster than that. Brent Eubanks, FinOps architect at @Wayfair, on tracking cost per step from observability logs and alerting when one step moves 10 to 15 percent.