At LiveKit we benchmarked time to first token across the LLMs people actually put in voice agents. Gemma 4 31B on LiveKit Inference came back at 192ms. GPT-4.1 came back at 1,006ms.
Here's where that speed comes from, and how to switch.
Still super impressed by the newspapers from @aiDotEngineer world’s fair. Best thing I’ve seen at an event in a long time. Great work @MLHacks and @ThePracticalDev.
Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits.
Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.
Demand for Fable has been challenging to