why did everyone suddenly decide specialized inference engines are a good idea?
i've seen at least 5 of these released in the last month.
what's wrong with vLLM/sgLang?
sure you can generate a slop inference engine quickly now but why fragment the ecosystem? it's harmful.