1. X
  2. DeepInfra
Log inSign up
DeepInfra
721 posts
DeepInfra profile banner
user avatar

DeepInfra

@DeepInfra
Fast ML inference. Run top AI models using a simple API.
Palo Alto
deepinfra.com
Joined February 2023
68
Following
25.8K
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    DeepInfra
    @DeepInfra
    Aug 15
    DeepSeek-V4-Pro-0813 is live on DeepInfra. Congrats to the @deepseek_ai team! 1.6T MoE, 49B active, 1M context. Running in US on @nvidia Blackwell: → priority + flex service tiers → prompt cache retention $1.30 in / $2.60 out / $0.10 cached input
    Image
  • user avatar
    DeepInfra
    @DeepInfra
    Aug 14
    Congrats to the @z_ai team on GLM-5.3! 🚀 We’re proud to host their open-source models at DeepInfra and excited to see them continue pushing the frontier of coding, agentic AI, and cybersecurity.
    user avatar
    Z.ai
    @Zai_org
    Aug 14
    Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: z.ai/blog/glm-5.3
    Image
  • user avatar
    DeepInfra
    @DeepInfra
    Aug 13
    New on DeepInfra: Qwen3.8-2.4T-A95B 🚀 @Alibaba_Qwen's latest sparse MoE — 2.4T total params, 95B active, 512 experts. Built for coding, agentic workflows, and complex reasoning, with native 262K context. Live now at $2.00/M in · $6.00/M out · $0.20/M cached @Alibaba_Qwen
    Image
  • user avatar
    DeepInfra
    @DeepInfra
    Aug 11
    New: Prompt Cache Retention on DeepInfra 🔥 Pin your prompt's KV cache for 5 min or 1 hour. Reuse skips prefill — faster TTFT, billed at the cache-read rate. Built for agents, multi-turn chat & doc Q&A/RAG. Live on Nemotron-3-Ultra & Kimi-K2.7-Code 👇
    Image
  • user avatar
    DeepInfra
    @DeepInfra
    Aug 11
    Day 0: @nvidia Nemotron 3.5 Lightning is live on DeepInfra. The fastest open model in its class — 30B MoE, 3B active params, 1M context, up to 4x higher throughput for always-on agents. $0.05/M in · $0.20/M out, OpenAI-compatible. More details in technical blog:

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement