Skip to content

Pinned Loading

  1. vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 87.6k 20k

  2. vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 5.7k 1.4k

  3. recipes Public

    Common recipes to run vLLM

    JavaScript 939 349

  4. llm-compressor Public

    Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

    Python 3.6k 592

  5. speculators Public

    A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

    Python 669 172

  6. semantic-router Public

    Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference

    Go 5.1k 785

Repositories

Showing 10 of 46 repositories