Skip to content

Pinned Loading

  1. vllm vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 93.4k 23.1k

  2. vllm-omni vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 7.1k 1.9k

  3. recipes recipes Public

    Common recipes to run vLLM

    JavaScript 1k 445

  4. llm-compressor llm-compressor Public

    State-of-the-art LLM compression, built for production inference with vLLM

    Python 3.9k 684

  5. speculators speculators Public

    A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

    Python 865 247

  6. semantic-router semantic-router Public

    An open, programmable decision layer for models and compute.

    Go 6.1k 1k

Repositories

Showing 10 of 50 repositories
  • semantic-router Public

    An open, programmable decision layer for models and compute.

    vllm-project/semantic-router's past year of commit activity
    Go 6,054 Apache-2.0 1,017 411 (1 issue needs help) 193 Updated Oct 8, 2026
  • vllm-gaudi Public

    Community maintained hardware plugin for vLLM on Intel Gaudi

    vllm-project/vllm-gaudi's past year of commit activity
    Python 61 Apache-2.0 157 17 63 Updated Oct 8, 2026
  • vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    vllm-project/vllm's past year of commit activity
    Python 93,378 Apache-2.0 23,123 2,580 (30 issues need help) 5,000+ Updated Oct 8, 2026
  • vllm-ascend Public

    Community maintained hardware plugin for vLLM on Huawei Ascend

    vllm-project/vllm-ascend's past year of commit activity
    Python 2,930 Apache-2.0 2,393 1,572 (77 issues need help) 2,222 Updated Oct 8, 2026
  • vllm-omni Public

    A framework for efficient model inference with omni-modality models

    vllm-project/vllm-omni's past year of commit activity
    Python 7,076 Apache-2.0 1,868 869 (151 issues need help) 1,048 Updated Oct 8, 2026
  • ci-infra Public

    This repo hosts code for vLLM CI & Performance Benchmark infrastructure.

    vllm-project/ci-infra's past year of commit activity
    Python 59 Apache-2.0 91 0 83 Updated Oct 8, 2026
  • tpu-inference Public

    TPU inference for vLLM, with unified JAX and PyTorch support.

    vllm-project/tpu-inference's past year of commit activity
    Python 455 Apache-2.0 332 109 (2 issues need help) 420 Updated Oct 8, 2026
  • afd-plugin Public

    vLLM plugin for attention-ffn disaggregation support

    vllm-project/afd-plugin's past year of commit activity
    Python 237 Apache-2.0 54 24 (4 issues need help) 14 Updated Oct 8, 2026
  • humming Public

    Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

    vllm-project/humming's past year of commit activity
    Python 249 Apache-2.0 47 6 11 Updated Oct 8, 2026
  • recipes Public

    Common recipes to run vLLM

    vllm-project/recipes's past year of commit activity
    JavaScript 1,042 Apache-2.0 445 56 136 Updated Oct 8, 2026