GPUStack v2.2 — Now Available

From GPU to
Token Factory
in Minutes

Deploy, govern, and scale LLMs and GPU compute on any hardware — on-premise, cloud, or hybrid. One platform for MaaS and GPUaaS.

GPU HARDWARENVIDIAAMDAscendT-headHygonMetaXMoore ThreadsMore accelerators...GPUStackENTERPRISE AI PLATFORMModel ManagementGPU InstancesInference EngineGPU SchedulingSecurity & RBACQuota & Rate LimitObservabilityBilling & MeteringINTEGRATIONSOpenAI CompatibleAnthropic CompatibleLangChain / n8nDify / RAGFlowClaude / OpenClawDocker / KubernetesPrometheus / GrafanaSSH / Jupyter
Any GPU
NVIDIA · AMD · Ascend · More
Any Model
LLM · Multimodal · Embedding · More
Any Inference Engine
vLLM · SGLang · llama.cpp · More
Any API
OpenAI · Anthropic · Custom

Run AI on Any GPU

From NVIDIA data-center GPUs to heterogeneous AI accelerators — unified management, one platform.

NVIDIA
H200 / H100 / A100 / more
AMD
MI300X / Radeon / Ryzen / more
Ascend
910B / 910C / 310P / more
T-head PPU
ZW810E / ZW910E / more
Hygon DCU
BW1000 / K100_AI / more
MetaX
C500 / C550 / C600 / more
Moore Threads
MTT S5000 / S4000 / more
Cambricon
MLU370 / MLU590 / more
Iluvatar
BI-V150 / MR-V100 / more
More AI Accelerators
Custom
x86_64 / ARM64
Ubuntu / RedHat / OpenEuler / Kylin / More Linux
Docker / Podman / Kubernetes

From GPU to Token in Minutes

A unified workflow that abstracts complexity across the entire AI inference stack.

01

Connect Your Model Sources

Browse and download from Hugging Face, ModelScope, or bring your own local model files. Compatibility checks are automated.

Hugging FaceModelScopeLocal FilesMulti-LoRA
02

Auto-Select Inference Engine

GPUStack automatically maps your hardware to the best-matching inference engine version. No manual configuration needed.

vLLMSGLangllama.cppTensorRT-LLMMindIE
03

Scale with Distributed Inference

Run large models across multiple nodes and GPUs. Tensor parallel, pipeline parallel, vLLM Ray cluster, MP distributed — all auto-orchestrated.

GPU PartitioningAuto Ray ClusterAuto MP distributed
04

Serve with Standard APIs

Expose models through OpenAI-compatible and Anthropic-compatible endpoints. Drop-in replacement for any AI framework or application.

OpenAI-compatibleAnthropic-compatibleCustom API

Enterprise MaaS + GPUaaS

MaaS for AI inference and GPUaaS for compute — unified under one control plane.

MaaS

Model as a Service

Full lifecycle management for AI models — deployment, traffic routing, performance tuning, and observability.

Multi-model deployment with load balancing & failover
Traffic weight routing, retry & fallback policies
Throughput / Latency / Standard / Custom performance modes
KV Cache optimization & token decode acceleration
Virtual model routing for zero-downtime upgrades
Public model provider integration (OpenAI, Claude, DeepSeek and more)
GPUaaS

GPU as a Service

Provision and manage GPU instances with persistent storage, flexible access, and fine-grained resource allocation.

GPU instance full lifecycle: create, start, stop, destroy
GPU partitioning with flexible slicing and overcommit strategies
SSH key auto-injection & Jupyter Notebook access
Persistent storage: S3 & NFS, multi-region mount
Custom instance templates with configurable environments
Transparent multi-cluster GPU provisioning

Support for Any Model Type

Deploy state-of-the-art open-source and private models with day-one support for new releases.

DE
DeepSeek
QW
Qwen
GL
GLM
KI
Kimi
MI
MiniMax
MI
Mistral
GE
Gemma
PH
Phi
LLMVLMMultimodalEmbeddingRerankerImageSpeechOCR
Search models, tasks, providers...⌘KCATEGORIESAll ModelsLLMMultimodalEmbeddingRerankerImageText-to-SpeechSpeech-to-TextQWQwen 3.6Multimodal · 27B● ReadyDeployGLGLM 5.1Language · 754B● ReadyDeployDEDeepSeek V4 ProLanguage · 1.6T● ReadyDeployKIKimi 2.6Multimodal · 1.1T● ReadyDeployMIMiniMax M2.7Language · 229B● ReadyDeployGEGemma 4Multimodal · 31B● ReadyDeploy

Deploy Anywhere

On-premise servers, Kubernetes clusters, or dynamic public cloud GPU instances — unified under one control plane.

PRIVATE DATA CENTERNode-018× NVIDIA H10085%Node-028× NVIDIA A10062%Node-034× NVIDIA H20091%

On-Premise

Full control over your existing GPU servers. Air-gapped deployments supported.

KUBERNETES CLUSTERControl PlaneAPI · etcd · SchedulerWorker-1PodPodPod4× GPU · RunningWorker-2PodPodPod8× GPU · RunningWorker-3PodPod---2× GPU · Pending

Kubernetes

Native Kubernetes integration for seamless GPU workload orchestration and elastic scaling.

MULTI-CLOUD GPUGPUStackControl PlaneAWSEC2 · P4d · P5AzureND · NC A100GCPA3 · H100 · TPUAlibabaGN · ebmgn · GPU● Auto-Scale

Multi-Cloud

Dynamically provision GPU resources on AWS, Azure, GCP, Alibaba Cloud, and more.

Inference Performance Modes

Optimize for your workload with preset or fully custom configurations.

Throughput

Max requests/s under high concurrency

Latency

Minimal TTFT for real-time applications

Standard

Full precision, maximum compatibility

Custom

Full parameter control for experts

Security & Compliance Built In

Production-grade security controls for enterprise AI deployments — governance from day one.

RBAC & Multi-Tenancy

Role-based access with full organization resource isolation.

SSO Integration

OIDC, SAML, and AD/LDAP protocol support for enterprise identity providers.

API Key Management

Scoped keys with expiry, rate limits, and per-model permission control.

IP Allowlisting

Restrict inference API access to trusted networks and CIDR ranges.

Token Quotas

Per-user and per-key token quota and rate limiting enforcement.

Usage Analytics

Model call counts, token usage by user, key, and model dimension.

High Availability

Multi-node redundancy with automatic failover and health monitoring.

Metering & Billing

Token, request, and GPU time-based usage tracking for cost visibility and billing.

Monitor, Measure, Optimize

Real-time metrics, historical trends, and resource topology — everything to keep models running at peak performance.

TOKENS/S · LIVE120k90k60k30kavg84.2kp95112kerr0.02%

Real-time Metrics

Inference latency, token rate, queue depth

REQUESTS · 7DMonTueWedThuFriSatToday↑ +23% vs last week · 2.4M total requests

Historical Trends

Performance analytics over time

GPU CLUSTER · LIVE89%GPU Util78%VRAM54%CPU67%RAMNode-01 · 8×H100 · 640 GB VRAM

Resource Usage

GPU/VRAM/CPU/RAM utilization across nodes

RESOURCE TOPOLOGYNode-01 · 4×H100Node-02 · 4×A100G0G1G2G3DeepSeek-V4-FlashTP4 · 89% utilG4G5G6G7Qwen-3-235BTP4 · 72% util8 GPU devices · 2 model instances · cluster view

Resource Topology

Visual map of nodes, GPUs, and model instances

Integrates with Your Entire Stack

Standard APIs and protocols mean GPUStack works with every major AI framework, Cloud Native tool, and DevOps platform out of the box.

AI Frameworks

Application frameworks & AI tools

OpenWebUI
LangChain
n8n
Dify
RAGFlow
Claude Code
OpenClaw
More ...

Cloud Native & DevOps

Deploy and operate at scale

Docker
Podman
Kubernetes
Helm
Higress
Prometheus
Grafana
More ...

Everything Under One Interface

A unified web UI to manage models, GPU clusters, users, API keys, and billing — no command-line required.

Manage Models
Deploy, update, and monitor models visually
GPU Resource Management
Allocate and schedule GPUs, workers, and clusters
Team Collaboration
User groups, model permissions, and API key scoping
Custom Branding
White-label with your logo, colors, and login page
GPUStack/ ModelsModelsClustersWorkersGPUsUsersAPI KeysUsageModels+ Deploy ModelNAMEENGINESTATUSGPUDeepSeek V4 ProvLLM● Running8×H200Qwen3.6 27BvLLM● Running4×4090GLM 5.1SGLang● Running16×H100Gemma 4llama.cpp● Starting1×5090Kimi 2.6vLLM● Stopped—

Frequently Asked Questions

What is GPUStack?

GPUStack is an open-source GPU cluster manager for deploying, governing, and scaling AI models on any hardware — on-premise, cloud, or hybrid. It supports inference engines like vLLM, SGLang, and TensorRT-LLM with OpenAI-compatible and Anthropic-compatible APIs.

Which GPU hardware does GPUStack support?

GPUStack supports NVIDIA, AMD, Ascend NPU, Hygon DCU, Moore Threads, MetaX, Cambricon MLU, Iluvatar, and T-Head PPU — making it hardware-agnostic across all major accelerator vendors.

Is GPUStack open source?

Yes. GPUStack is fully open source under the Apache 2.0 license. The source code is available on GitHub at github.com/gpustack/gpustack.

How does GPUStack compare to running vLLM directly?

vLLM handles single-node inference serving. GPUStack adds a full management layer on top: multi-cluster orchestration, load balancing, high availability, RBAC, monitoring, billing, and a unified API gateway across multiple inference engines.

Does GPUStack support enterprise features?

Yes. GPUStack Enterprise Edition adds high availability across management and business planes, full multi-tenancy with resource isolation, fine-grained RBAC, white-label branding, and dedicated support.

How do I get started with GPUStack?

You can install GPUStack with a single Docker command. After the server starts, add GPU workers via the UI, deploy a model from the catalog, and access it through the OpenAI-compatible API. Full instructions are at docs.gpustack.ai.

Start Your Enterprise
AI Infrastructure Journey

Join enterprises worldwide building their token factory with GPUStack.