Skip to content

Pinned Loading

  1. vllm vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 93.6k 23.2k

  2. vllm-omni vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 7.1k 1.9k

  3. recipes recipes Public

    Common recipes to run vLLM

    JavaScript 1k 447

  4. llm-compressor llm-compressor Public

    State-of-the-art LLM compression, built for production inference with vLLM

    Python 3.9k 685

  5. speculators speculators Public

    A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

    Python 871 250

  6. semantic-router semantic-router Public

    An open, programmable decision layer for models and compute.

    Go 6.1k 1k

Repositories

Showing 10 of 50 repositories
  • vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    vllm-project/vllm's past year of commit activity
    Python 93,562 Apache-2.0 23,244 2,581 (30 issues need help) 5,000+ Updated Oct 11, 2026
  • semantic-router Public

    An open, programmable decision layer for models and compute.

    vllm-project/semantic-router's past year of commit activity
    Go 6,081 Apache-2.0 1,024 406 (1 issue needs help) 171 Updated Oct 11, 2026
  • vllm-omni Public

    A framework for efficient model inference with omni-modality models

    vllm-project/vllm-omni's past year of commit activity
    Python 7,122 Apache-2.0 1,885 882 (151 issues need help) 1,060 Updated Oct 11, 2026
  • humming Public

    Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

    vllm-project/humming's past year of commit activity
    Python 250 Apache-2.0 47 6 11 Updated Oct 11, 2026
  • vllm-ascend Public

    Community maintained hardware plugin for vLLM on Huawei Ascend

    vllm-project/vllm-ascend's past year of commit activity
    Python 2,949 Apache-2.0 2,421 1,599 (77 issues need help) 2,331 Updated Oct 11, 2026
  • vllm-metal Public

    Community maintained hardware plugin for vLLM on Apple Silicon

    vllm-project/vllm-metal's past year of commit activity
    Python 1,835 Apache-2.0 297 13 23 Updated Oct 11, 2026
  • tpu-inference Public

    TPU inference for vLLM, with unified JAX and PyTorch support.

    vllm-project/tpu-inference's past year of commit activity
    Python 457 Apache-2.0 336 108 (2 issues need help) 427 Updated Oct 11, 2026
  • MSA Public Forked from MiniMax-AI/MSA
    vllm-project/MSA's past year of commit activity
    Python 1 MIT 68 0 1 Updated Oct 11, 2026
  • afd-plugin Public

    vLLM plugin for attention-ffn disaggregation support

    vllm-project/afd-plugin's past year of commit activity
    Python 238 Apache-2.0 54 20 (4 issues need help) 16 Updated Oct 11, 2026
  • guidellm Public

    Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs

    vllm-project/guidellm's past year of commit activity
    Python 1,683 Apache-2.0 263 56 55 Updated Oct 10, 2026