Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    Models

    Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices

    Optimized by NVIDIALaunch from Hugging FaceBeta

    Filters

    Use Case
    Inference Providers
    Publisher
    NIM Container GPUs
    124 models

    models

    • NVIDIA
      DownloadableFree Endpoint

      nemotron-3.5-lightning-30b-a3b

      Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
      • Customization
      • Text-to-Text
      • Long-running agents
      • Open
      Last updated on August 11, 2026
    Items per page
    of 6 pages
  • Meta
    DownloadableFree Endpoint

    muse-glimmer-30b

    Muse Glimmer 30B is a multimodal reasoning model accepting text and images, served on vLLM with native Onyx tool-calling and reasoning parsers.
    • Multimodal
    • Image-to-Text
    • Reasoning
    • Chat
    • Text-to-Text
    • Large Language Models
    Last updated on August 10, 2026
  • NVIDIA
    Free Endpoint

    riva-translate-4b-instruct-v2

    Translation model in 37 languages with few-shots example prompts capability.
    • nvidia nim
    • neural machine translation
    • Text Translation
    Last updated on July 27, 2026
  • NVIDIA
    Free Endpoint

    ising-calibration-1.5-31b

    NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.
    • Quantum Computing
    • Calibration
    • NVIDIA NIM
    • Vision Language Model
    Last updated on July 23, 2026
  • NVIDIA
    Downloadable

    Video Super Resolution NIM

    Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.
    • broadcast
    • video upscaling
    • streaming
    • nvidia ai for media
    • video super resolution
    Last updated on July 22, 2026
  • NVIDIA
    Free Endpoint

    nemotron-3-embed-1b

    1B embedding model for semantic search, retrieval, and RAG applications.
    • Nemotron Retriever
    • Agentic Retrieval
    • Code Retrieval
    • Text-to-Embedding
    • Retrieval Augmented Generation
    Last updated on July 16, 2026
  • Thinkingmachines
    DownloadableFree Endpoint

    inkling

    Inkling is a multimodal (text + image) reasoning model from Thinking Machines — a Mamba-hybrid, 256-expert Mixture-of-Experts architecture with tool use and switchable reasoning.
    • text-to-text
    • reasoning
    • image-to-text
    • multimodal
    Last updated on July 16, 2026
  • Poolside
    Free Endpoint

    laguna-xs-2.1

    Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
    • Agentic AI
    • Coding
    • Reasoning
    • Tool Use
    Last updated on July 15, 2026
  • Z.ai
    DownloadableFree Endpoint

    glm-5.2

    GLM-5.2 is a flagship LLM for agentic workflows, coding, and long-horizon reasoning tasks.
    • Agentic AI
    • Coding
    • Reasoning
    • Tool Use
    8M API calls in the last 30 days
    Last updated on July 3, 2026
  • NVIDIA
    Downloadable

    qwen-image-edit-nvpcb-ovsl2sl

    An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stations
    • Synthetic Data Generation
    • Image Generation
    • Physical AI
    Last updated on July 3, 2026
  • NVIDIA
    Downloadable

    nemotron-ocr-v2

    Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.
    • Table Extraction
    • nemo retriever
    • data ingestion
    • extraction
    • Optical Character Recognition
    338K API calls in the last 30 days
    Last updated on June 24, 2026
  • Minimaxai
    Free Endpoint

    minimax-m3

    MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
    • coding
    • text-to-text
    • reasoning
    10M API calls in the last 30 days
    Last updated on June 12, 2026
  • Google
    DownloadableFree Endpoint

    diffusiongemma-26b-a4b-it

    Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
    • diffusion-llm
    • text-to-text
    • reasoning
    4M API calls in the last 30 days
    Last updated on June 10, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-ultra-550b-a55b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    • Agent
    • MoE
    • Frontier
    • Reasoning
    • Long Context
    52M API calls in the last 30 days
    Last updated on June 4, 2026
  • Resemble.AI
    Downloadable

    chatterbox-multilingual-tts

    Natural and expressive voices in 23 languages. For voice agents and brand ambassadors.
    • TTS
    • Chatterbox
    • Speech Generation
    • multilingual
    • Text-to-Speech
    22K API calls in the last 30 days
    Last updated on June 3, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3.5-content-safety

    Multilingual, multimodal model for detecting unsafe and toxic content.
    • llm safety
    • safety and moderation
    • multilingual content safety
    • ai safety nemo guardrails
    2M API calls in the last 30 days
    Last updated on June 2, 2026
  • NVIDIA
    DownloadableFree Endpoint

    cosmos3-nano

    Generates physics-aware videos from text prompts or an image prompt for physical AI development.
    • autonomous vehicles
    • Physical AI
    • robotics
    • text-to-world
    • image-to-world
    • Synthetic Data Generation
    2K API calls in the last 30 days
    Last updated on June 1, 2026
  • NVIDIA
    DownloadableFree Endpoint

    cosmos3-nano-reasoner

    Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
    • video understanding
    • autonomous vehicles
    • industrial
    • Physical AI
    • vision language model
    • reasoning
    • robotics
    • smart cities
    • Synthetic Data Generation
    2K API calls in the last 30 days
    Last updated on June 1, 2026
  • Stepfun-ai
    DownloadableFree Endpoint

    step-3.7-flash

    A sparse MoE multimodal reasoning model good for enterprise, agentic and coding tasks.
    • Coding
    • Vision
    • Agents
    7M API calls in the last 30 days
    Last updated on May 29, 2026
  • Qwen
    Downloadable

    qwen-image

    Qwen-Image is a text-to-image foundation model with advanced multilingual text rendering.
    • Text-to-Image
    • Image Generation
    Last updated on May 1, 2026
  • Qwen
    Downloadable

    qwen-image-edit

    Qwen-Image-Edit is an image editing model with multilingual text editing and strong subject consistency.
    • Text-to-Image
    • Image Generation
    Last updated on May 1, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-nano-omni-30b-a3b-reasoning

    Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
    • Image-to-Text
    • VLM
    • Video
    • Omni
    • OCR
    8M API calls in the last 30 days
    Last updated on April 28, 2026
  • NVIDIA
    Downloadable

    Relighting

    Re-illuminate people in video to match target lighting from a 360 HDRI environment map.
    • HDRI
    • remote contribution
    • lighting
    • nvidia ai for media
    242 API calls in the last 30 days
    Last updated on April 17, 2026
  • NVIDIA
    DownloadableFree Endpoint

    synthetic-video-detector

    NVIDIA Synthetic Video Detector is an AI-powered micro-service for detecting AI‑generated (synthetic) videos.
    • broadcast
    • media2
    • forensics
    • nvidia ai for media
    • diffusion models
    309K API calls in the last 30 days
    Last updated on April 16, 2026
  • Advertisement
    Advertisement