Deploy, govern, and scale LLMs and GPU compute on any hardware — on-premise, cloud, or hybrid. One platform for MaaS and GPUaaS.
Heterogeneous GPU
From NVIDIA data-center GPUs to heterogeneous AI accelerators — unified management, one platform.
Platform Workflow
A unified workflow that abstracts complexity across the entire AI inference stack.
Browse and download from Hugging Face, ModelScope, or bring your own local model files. Compatibility checks are automated.
GPUStack automatically maps your hardware to the best-matching inference engine version. No manual configuration needed.
Run large models across multiple nodes and GPUs. Tensor parallel, pipeline parallel, vLLM Ray cluster, MP distributed — all auto-orchestrated.
Expose models through OpenAI-compatible and Anthropic-compatible endpoints. Drop-in replacement for any AI framework or application.
Performance Lab
GPUStack's tuned deployment achieves significant improvements over unoptimized baselines.
GLM-4.6 Throughput
Throughput improvement over unoptimized vLLM baseline on 8× NVIDIA H200.
Qwen3-8B Latency
Latency reduction on Qwen3-8B inference, ideal for real-time applications.
Qwen3-235B-A22B
Throughput improvement on MoE model deployment across multi-node H100 clusters.
Two Platforms, One Stack
MaaS for AI inference and GPUaaS for compute — unified under one control plane.
Full lifecycle management for AI models — deployment, traffic routing, performance tuning, and observability.
Provision and manage GPU instances with persistent storage, flexible access, and fine-grained resource allocation.
Model Ecosystem
Deploy state-of-the-art open-source and private models with day-one support for new releases.
Infrastructure
On-premise servers, Kubernetes clusters, or dynamic public cloud GPU instances — unified under one control plane.
Full control over your existing GPU servers. Air-gapped deployments supported.
Native Kubernetes integration for seamless GPU workload orchestration and elastic scaling.
Dynamically provision GPU resources on AWS, Azure, GCP, Alibaba Cloud, and more.
Optimize for your workload with preset or fully custom configurations.
Max requests/s under high concurrency
Minimal TTFT for real-time applications
Full precision, maximum compatibility
Full parameter control for experts
Enterprise Ready
Production-grade security controls for enterprise AI deployments — governance from day one.
Role-based access with full organization resource isolation.
OIDC, SAML, and AD/LDAP protocol support for enterprise identity providers.
Scoped keys with expiry, rate limits, and per-model permission control.
Restrict inference API access to trusted networks and CIDR ranges.
Per-user and per-key token quota and rate limiting enforcement.
Model call counts, token usage by user, key, and model dimension.
Multi-node redundancy with automatic failover and health monitoring.
Token, request, and GPU time-based usage tracking for cost visibility and billing.
Observability
Real-time metrics, historical trends, and resource topology — everything to keep models running at peak performance.
Inference latency, token rate, queue depth
Performance analytics over time
GPU/VRAM/CPU/RAM utilization across nodes
Visual map of nodes, GPUs, and model instances
Open Ecosystem
Standard APIs and protocols mean GPUStack works with every major AI framework, Cloud Native tool, and DevOps platform out of the box.
Application frameworks & AI tools
Deploy and operate at scale
Management
A unified web UI to manage models, GPU clusters, users, API keys, and billing — no command-line required.
FAQ
GPUStack is an open-source GPU cluster manager for deploying, governing, and scaling AI models on any hardware — on-premise, cloud, or hybrid. It supports inference engines like vLLM, SGLang, and TensorRT-LLM with OpenAI-compatible and Anthropic-compatible APIs.
GPUStack supports NVIDIA, AMD, Ascend NPU, Hygon DCU, Moore Threads, MetaX, Cambricon MLU, Iluvatar, and T-Head PPU — making it hardware-agnostic across all major accelerator vendors.
Yes. GPUStack is fully open source under the Apache 2.0 license. The source code is available on GitHub at github.com/gpustack/gpustack.
vLLM handles single-node inference serving. GPUStack adds a full management layer on top: multi-cluster orchestration, load balancing, high availability, RBAC, monitoring, billing, and a unified API gateway across multiple inference engines.
Yes. GPUStack Enterprise Edition adds high availability across management and business planes, full multi-tenancy with resource isolation, fine-grained RBAC, white-label branding, and dedicated support.
You can install GPUStack with a single Docker command. After the server starts, add GPU workers via the UI, deploy a model from the catalog, and access it through the OpenAI-compatible API. Full instructions are at docs.gpustack.ai.
Join enterprises worldwide building their token factory with GPUStack.