Qwen3.8-27B input just dropped 25% to $0.30/M on Tiyuvta Inference. Cached input $0.10 and output $2.03 unchanged. Same 180 tok/s, same full tool calling, same OpenAI-compatible API. Agents burn budgets on input tokens. Yours just got cheaper. inference.tiyuvta.ai/pricing
Systems SWE @ AWS ElastiCache | Valkey GLIDE maintainer | Low-level systems · GPU kernels | Applied ML research → avifenesh.ai | OSS all the way
Joined February 2013
- I could also write code by hand and do CPU optimization myself with the same method, but the point is to have a helpful and easy-to-use model. Opus is too verbose, blasts changes, rewrites what isn't asked for, and acts like your moral teacher. Everything is possible;Replying to @scaling01Agree. People are sleeping on using Opus to hill climb. We use it for optimizing CPU and memory, optimizing CI times, improving frame rates, reducing latency, any other kind of problem in the shape of “iterate on X with a profiler and dataset until it hits Y”
- huggingface.co/Avifenesh/Orni… trained the mtp head to fit the model, instead of the base mode, big performance improvements.
- Said I'm still tuning, and it's up #qwen 3.8 27b live with 260 tok/s with a much lower price than open router providers. Input 0.40, cached 0.10, output 2.03 (cache hit 77.9%), $5 credit on signup, no card needed. inference.tiyuvta.ai


