-
-
Notifications
You must be signed in to change notification settings - Fork 20k
Projects
Repository projects
PRs and issues related to NVIDIA hardware
#31 updatednowJul 30, 2026 - #16 updated
24 minutes agoJul 30, 2026 Work on the Transformers modeling backend: running Transformers model implementations inside vLLM.
#28 updated2 hours agoJul 30, 2026 Tracking failures that are occurring in CI.
#20 updated4 hours agoJul 29, 2026 Track CPU related issues & tasks
#42 updated8 hours agoJul 29, 2026 Maintainer's tracking board for Prometheus metrics related PRs and issues
#44 updated15 hours agoJul 29, 2026 torch.compile integration related
#12 updated15 hours agoJul 29, 2026 Tracks Ray issues and pull requests in vLLM
#7 updatedyesterdayJul 28, 2026 2025-02-25: DeepSeek V3/R1 is supported with optimized block FP8 kernels, MLA, MTP spec decode, multi-node PP, EP, and W4A16 quantization
#5 updated2 days agoJul 28, 2026 Main tasks for the multi-modality workstream (#4194)
#8 updated2 days agoJul 27, 2026 Optimization and bugfixes for Qwen3.5 model series.
#50 updated3 days agoJul 26, 2026 - #47 updated
4 days agoJul 26, 2026 Community requests for multi-modal models
#10 updated4 days agoJul 25, 2026 - #57 updated
5 days agoJul 25, 2026 - #24 updated
last weekJul 24, 2026 - #51 updated
last weekJul 24, 2026 - #46 updated
2 weeks agoJul 16, 2026 - #29 updated
2 weeks agoJul 14, 2026 - #33 updated
on Jun 22Jun 22, 2026 A list of onboarding tasks for first-time contributors to get started with vLLM.
#6 updatedon May 31May 31, 2026 - #45 updated
on May 17May 17, 2026 Backlog for CI feature requests
#35 updatedon May 8May 8, 2026 - #25 updated
on Mar 7Mar 7, 2026 Tracker of known issues and bugs for serving Llama on vLLM
#14 updatedon Feb 6Feb 6, 2026 Enhancement to Llama herd of models. See also https://github.com/vllm-project/vllm/issues/16114
#13 updatedon Nov 20, 2025Nov 20, 2025 [Testing] Optimize V1 PP efficiency.
#1 updatedon Oct 6, 2025Oct 6, 2025 - #2 updated
on Aug 15, 2025Aug 15, 2025