1. X
  2. Avi Fenesh
Log inSign up
Avi Fenesh
308 posts
Avi Fenesh profile banner
user avatar

Avi Fenesh

@avi_fenesh
Systems SWE @ AWS ElastiCache | Valkey GLIDE maintainer | Low-level systems · GPU kernels | Applied ML research → avifenesh.ai | OSS all the way
avifenesh.ai
Joined February 2013
91
Following
75
Followers
RepliesRepliesMediaMedia
  • user avatar
    Avi Fenesh
    @avi_fenesh
    18h
    Qwen3.8-27B input just dropped 25% to $0.30/M on Tiyuvta Inference. Cached input $0.10 and output $2.03 unchanged. Same 180 tok/s, same full tool calling, same OpenAI-compatible API. Agents burn budgets on input tokens. Yours just got cheaper. inference.tiyuvta.ai/pricing
  • user avatar
    Avi Fenesh
    @avi_fenesh
    Aug 23
    I could also write code by hand and do CPU optimization myself with the same method, but the point is to have a helpful and easy-to-use model. Opus is too verbose, blasts changes, rewrites what isn't asked for, and acts like your moral teacher. Everything is possible;
    user avatar
    Boris Cherny
    @bcherny
    Aug 22
    Replying to @scaling01
    Agree. People are sleeping on using Opus to hill climb. We use it for optimizing CPU and memory, optimizing CI times, improving frame rates, reducing latency, any other kind of problem in the shape of “iterate on X with a profiler and dataset until it hits Y”
  • user avatar
    Avi Fenesh
    @avi_fenesh
    Aug 22
    huggingface.co/Avifenesh/Orni… trained the mtp head to fit the model, instead of the base mode, big performance improvements.
    Image
    Avifenesh/Ornith-1.5-35B-A3B-NVFP4-MTP-GGUF · Hugging Face
    From huggingface.co
  • user avatar
    Avi Fenesh
    @avi_fenesh
    Aug 22
    @ornith_ Ornith-1.5 35B-A3B $0.25/$0.09/$1.20 350 tok/s inference.tiyuvta.ai
  • user avatar
    Avi Fenesh
    @avi_fenesh
    Aug 21
    Said I'm still tuning, and it's up #qwen 3.8 27b live with 260 tok/s with a much lower price than open router providers. Input 0.40, cached 0.10, output 2.03 (cache hit 77.9%), $5 credit on signup, no card needed. inference.tiyuvta.ai

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement