-
-
Notifications
You must be signed in to change notification settings - Fork 20k
Insights: vllm-project/vllm
July 22, 2026 – July 29, 2026
Overview
Summary
Excluding merges, 211 authors have pushed 297 commits to main and 369 commits to all branches.
On main, 1034 files have changed and there have been 84,940 additions and 11,478 deletions
1 release published by 1 person
- v0.26.0 v0.26.0
3 days agoJul 27, 2026
297 pull requests merged by 154 people
- [docs] Add documentation for pynvvideocodec video decoding backend
#49066 merged
24 minutes agoJul 30, 2026 - fix(step3p5-mtp): honor exclude_modules for the MTP head via prefix
#48883 merged
49 minutes agoJul 30, 2026 - [KV Connector] Fix NIXL mamba state pairing for multi-slot block tables
#50153 merged
1 hour agoJul 30, 2026 - [ROCm][CI] Fix Kimi K3 KDA on ROCm
#50262 merged
2 hours agoJul 30, 2026 - [Spec Decode][Perf] Replicate DSpark Markov head across TP ranks
#49731 merged
2 hours agoJul 30, 2026 - [BugFix] Fix `num_output_placeholders` preemption underflow
#48245 merged
4 hours agoJul 29, 2026 - [ModelRunner V2] Enable sequence pooling for embedding and classification models
#48791 merged
4 hours agoJul 29, 2026 - [Bugfix][Frontend] Return transcription and translation verbose as float
#49073 merged
4 hours agoJul 29, 2026 - [Rust Frontend] Send multimodal tensors in auxiliary frames
#49341 merged
5 hours agoJul 29, 2026 - [CI][Test] Fix pooling truncation test after VLLMError hierarchy change
#50241 merged
5 hours agoJul 29, 2026 - [XPU] Route weightless RMSNorm to _C dispatch
#47121 merged
6 hours agoJul 29, 2026 - [Kernel][Mamba] Fused-kernel support for align-mode DS-conv state migration with num_accepted_tokens > 1
#49291 merged
6 hours agoJul 29, 2026 - [Bugfix][Multimodal] Include media IO config in MM cache hash
#49975 merged
7 hours agoJul 29, 2026 - [KV Offload] Move CPUOffloadingSpec onto SharedOffloadRegion
#50094 merged
7 hours agoJul 29, 2026 - [CI/Perf] Fix malformed serving benchmark config
#43538 merged
7 hours agoJul 29, 2026 - [Rust Frontend] Add --limit-mm-per-prompt support
#49604 merged
8 hours agoJul 29, 2026 - [CI] Fix MXFP8 MOE backend selection tests on gfx942
#50222 merged
8 hours agoJul 29, 2026 - [EC Connector] Add has_pending_push_work
#49582 merged
8 hours agoJul 29, 2026 - [XPU] upgrade to torch 2.13
#48677 merged
9 hours agoJul 29, 2026 - [Misc][Minimax-M3]add default video_processor
#50092 merged
9 hours agoJul 29, 2026 - [BugFix] eagle draft max position embeddings
#49343 merged
9 hours agoJul 29, 2026 - [Model] Support Qwen3.5 text-only dense and MoE models
#50210 merged
9 hours agoJul 29, 2026 - [Model] Add Kimi K3 support: Python frontend [2/2]
#50093 merged
9 hours agoJul 29, 2026 - [Frontend] Reuse prefill token ids on the decode chat path for disaggregated serving
#48145 merged
9 hours agoJul 29, 2026 - [CPU] Fix FP8 attention scratchpad sizing
#50194 merged
10 hours agoJul 29, 2026 - [CI] Allow PR comment acknowledgements
#50211 merged
10 hours agoJul 29, 2026 - [XPU] [UT] [CI] add xpu config to run gpt-oss accuracy in ut and ci
#48703 merged
11 hours agoJul 29, 2026 - [Model] Add Kimi K3 support: model files and kernels [1/N]
#50089 merged
11 hours agoJul 29, 2026 - fused_moe: add VLLM_TRITON_USE_TD tensor-descriptor path
#42436 merged
12 hours agoJul 29, 2026 - [Test] dynamic_shapes_compilation
#49974 merged
12 hours agoJul 29, 2026 - [CPU] Fix s390x builds and update torch version in dockerfile
#50144 merged
13 hours agoJul 29, 2026 - Add CachePolicyFactory for pluggable/external eviction policies
#49114 merged
13 hours agoJul 29, 2026 - [CI][ROCm] Stabilize Qwen2-VL LoRA test
#50161 merged
13 hours agoJul 29, 2026 - [Frontend][Core] Standardize request error handling with VLLMError hierarchy
#49665 merged
13 hours agoJul 29, 2026 - [KV Connector] Support NIXL P/D for hybrid MLA+SSM models
#49762 merged
14 hours agoJul 29, 2026 - Integrate CuTeDSL MoE for ReLU2 NVFP4
#49580 merged
15 hours agoJul 29, 2026 - [CompressedTensors] FP4 Qutlass Integration
#43229 merged
15 hours agoJul 29, 2026 - [ROCm][CI] Stabilize ngram and suffix correctness test
#50190 merged
15 hours agoJul 29, 2026 - [Model] Add Inkling compressed-tensors dynamic FP8 support
#48876 merged
15 hours agoJul 29, 2026 - [CI] Allow comment-triggered builds past pipeline filters
#50197 merged
15 hours agoJul 29, 2026 - [CI] Add comment-based Buildkite triggers
#50132 merged
16 hours agoJul 29, 2026 - [BugFix] Stop dummy runs from writing mamba state through stale block-table rows
#49757 merged
16 hours agoJul 29, 2026 - [Bugfix] Fix /wake_up crash on hybrid models (Mamba/DeltaNet)
#41602 merged
17 hours agoJul 29, 2026 - [CI][NIXL] Fix flaky DP+EP test port conflict
#50171 merged
17 hours agoJul 29, 2026 - [ROCm][CI] Stabilize ROCm audio streaming test
#50163 merged
17 hours agoJul 29, 2026 - [MXFP8][ROCm] Fix MXFP8 MoE backend selection
#49747 merged
17 hours agoJul 29, 2026 - [Test] Make EPD correctness tests configurable for XPU
#50110 merged
18 hours agoJul 29, 2026 - [Perf] Zero-copy torch.Tensor pickling in shm_broadcast MessageQueue
#48442 merged
18 hours agoJul 29, 2026 - perf: dispatch non-grouped bias-less topk routing methods to fused path
#49618 merged
19 hours agoJul 29, 2026 - [Model] Add Kimi K3 support: Rust frontend [1/2]
#50104 merged
20 hours agoJul 29, 2026 - [Bugfix][Spec Decode] Size DFlash query buffers for cudagraph-padded batches
#50065 merged
yesterdayJul 29, 2026 - [Rust Frontend] Extract shared tracing setup logic into `vllm-tracing`
#50129 merged
yesterdayJul 29, 2026 - [BugFix] Fix clang spinloop mwaitx include
#45532 merged
yesterdayJul 29, 2026 - [ROCm][Bugfix] Sanitize AITER paged-MQA logits before sparse top-k for DeepSeek-V4
#49714 merged
yesterdayJul 29, 2026 - [Bugfix][Multimodal] Fix video temporal padding estimates
#49030 merged
yesterdayJul 29, 2026 - [compressed-tensors] update `find_matched_target` order to prioritize fused name matches over class match
#49483 merged
yesterdayJul 29, 2026 - [ROCm] Fix and optimize GPT-J-style MRoPE
#49906 merged
yesterdayJul 29, 2026 - [ROCm] Cache fp32 upcast of static e8m0 weight scale in AITER scaled_mm
#47773 merged
yesterdayJul 29, 2026 - [Rust][Benchmark] Make `vllm bench serve` Rust delegation opt-in
#50081 merged
yesterdayJul 29, 2026 - [Attention] Skip sparse indexer scoring for dense short prefills
#48407 merged
yesterdayJul 29, 2026 - [Bugfix] Don't reuse engine core payload buffer while zmq is sending it
#50053 merged
yesterdayJul 29, 2026 - [Docs] Expand llm-d integration page
#45432 merged
yesterdayJul 29, 2026 - [Bugfix][KV Offload][P2P] Scope serve state to fetch rounds
#49877 merged
yesterdayJul 29, 2026 - [Core] Warm up runner-owned Triton kernels before the first request
#49903 merged
yesterdayJul 28, 2026 - [Bugfix][KV Offload] Keep Mamba block span unscaled under DCP
#49964 merged
yesterdayJul 28, 2026 - [Bugfix][Kernel] Fix batch invariance in RMSNorm kernels by pinning block size
#48391 merged
yesterdayJul 28, 2026 - [CI] Wire untethered test files into CI jobs
#49340 merged
yesterdayJul 28, 2026 - [Bugfix] Add missing `vllm/models/kimi_k3/__init__.py`
#50131 merged
yesterdayJul 28, 2026 - [KV Connector] Support NIXL heterogeneous P/D block sizes for hybrid models
#49612 merged
yesterdayJul 28, 2026 - add epilogue hook to flex attention
#45841 merged
yesterdayJul 28, 2026 - [PD][NixlPush] Skip extra `add_remote_agent` step in D->P handshake
#49345 merged
yesterdayJul 28, 2026 - [Bugfix] Enhance extra_config handling for layer name suffix matching
#48589 merged
yesterdayJul 28, 2026 - [Elastic EP] Async preparation
#47288 merged
yesterdayJul 28, 2026 - [Build] Fix DeepEP CUDA driver stub linking
#50103 merged
yesterdayJul 28, 2026 - [Rust Frontend] Align sampling validation with Python
#47494 merged
yesterdayJul 28, 2026 - [Rust Frontend] Fix finish reason for named tool choices
#49496 merged
yesterdayJul 28, 2026 - [CI] Increase Qwen3.5 MTP GSM8K generation length
#49881 merged
yesterdayJul 28, 2026 - [CI] Add PyTorch stable ABI audit check
#48164 merged
yesterdayJul 28, 2026 - [Test] Skip ROCm AITER MLA prefill tests on non-ROCm platforms
#49945 merged
yesterdayJul 28, 2026 - [Rust Frontend] Add ordinary-text tokenizer encoding
#49992 merged
yesterdayJul 28, 2026 - [Kimi-K3] Add AttnRes kernels
#50090 merged
yesterdayJul 28, 2026 - [CI][ROCm] Soft fail LoRA mirror
#50086 merged
yesterdayJul 28, 2026 - [Bugfix][Spec Decode] Preserve draft buffers across level-2 sleep
#49774 merged
yesterdayJul 28, 2026 - [Bugfix] Respect cgroup memory limits on all platforms
#49966 merged
yesterdayJul 28, 2026 - [Core][Frontend] Add weight version tagging for RL rollouts
#49040 merged
yesterdayJul 28, 2026 - [Bugfix] Fix multi-modal support on CPU MRV2
#50073 merged
yesterdayJul 28, 2026 - Remove triton per group quant [ROCm] [Bugfix]
#49621 merged
yesterdayJul 28, 2026 - [Build] Fix CUDA arch detection producing kernel-less builds on SM121
#49904 merged
2 days agoJul 28, 2026 - [Bugfix] Only pad transformers backend `value` when it is narrower
#50060 merged
2 days agoJul 28, 2026 - [Bugfix][KV Offload][OBJ] Preserve job completion during cleanup
#49947 merged
2 days agoJul 28, 2026 - [KV Offload] Make compact secondary identity TP-independent
#49858 merged
2 days agoJul 28, 2026 - [Bugfix] Fix DeepseekV4FP8 Quark MXFP4 crash on list-valued weight
#49634 merged
2 days agoJul 28, 2026 - [AMD] Revert `Mxfp4MoeBackend.TRITON_UNFUSED` fallback
#46491 merged
2 days agoJul 28, 2026 - [KV-offload][FS] : Batch store/load_block in C
#49152 merged
2 days agoJul 28, 2026 - [Feature] Add VidCom2 video token pruning
#47750 merged
2 days agoJul 28, 2026 - [ROCm][Quark][6/N] Use MXFP4 linear kernel abstraction for `aiter` backend
#49348 merged
2 days agoJul 28, 2026 - [XPU] Add online fp8 quantization test
#44513 merged
2 days agoJul 28, 2026 - [CI][ROCm] Soft-fail Python-only installation mirror
#50041 merged
2 days agoJul 28, 2026 - [ROCm][DSv3.2] Eliminate per-decode FillFunctor launches in sparse-MLA hot loop
#44527 merged
2 days agoJul 28, 2026 - [ROCm][KVConnector][MoRI-IO] Fix WRITE-mode remote-TP rank collapse (#46332 follow-up)
#47764 merged
2 days agoJul 28, 2026 - [Docs] Remove experimental warning for EP
#50057 merged
2 days agoJul 28, 2026 - [Core][PCP] Select MRV2 when PCP is enabled
#50034 merged
2 days agoJul 28, 2026 - [Tests][Spec Decode] Add gemma4 MTP acceptance rates test
#47920 merged
2 days agoJul 28, 2026 - Fix Humming non-gated MoE
#49096 merged
2 days agoJul 28, 2026 - Fix MQA with tensor parallelism on transformers modeling backend
#49987 merged
2 days agoJul 28, 2026 - [ROCm] [BugFix] Fix Quark GLM-5.2 Checkpoint inference: indexer wk per-channel FP8 dequant + missing sparse-MLA metadata fields
#48886 merged
2 days agoJul 28, 2026 - [CI][ROCm] Fix `test_ocp_mx_wikitext_correctness` reference value
#49690 merged
2 days agoJul 28, 2026 - [CI] Initialize DeepEP FP8 test weights
#49912 merged
2 days agoJul 28, 2026 - [ROCm][Quantization][5/N] Refactor quark_moe w8a8-int8 w/ oracle
#46765 merged
2 days agoJul 28, 2026 - [Refactor] Remove dead code in multiple files
#49745 merged
2 days agoJul 28, 2026 - [DSv4 Perf] Adaptive topk width, 1.0% E2E throughput improvement
#50004 merged
2 days agoJul 28, 2026 - [Core] Fail fast when /dev/shm is too small for the shm ring buffer
#48879 merged
2 days agoJul 28, 2026 - [Test] Regression test for hybrid-Mamba eagle cache-peek in Mooncake connector (#43559)
#48361 merged
2 days agoJul 28, 2026 - [Bugfix][CPU] Fall back to torch for unaligned swigluoai on NEON/vec MoE
#49985 merged
2 days agoJul 28, 2026 - Fix MLA padding and grouped topk routing in the Transformers modelling backend
#49982 merged
2 days agoJul 28, 2026 - [Attention] Integrate FlashAttention 4 SM100 headdim 256 support
#42669 merged
2 days agoJul 28, 2026 - [MRV2] Always build attn metadata at capture time (#49364)
#49995 merged
2 days agoJul 28, 2026 - [Bugfix] Changed speech to text chunk timestamp to cumulative approach
#41131 merged
2 days agoJul 28, 2026 - [Rust Frontend][gRPC] Add server and model discovery
#49491 merged
2 days agoJul 28, 2026 - [Misc][PD] Nixl cleanup `get_backend_aware_kv_block_len` and `virtually_split_kv_in_blocks`
#49988 merged
2 days agoJul 28, 2026 - [Bugfix] Fix VLLM_ENFORCE_STRICT_TOOL_CALLING mutation in tests
#49846 merged
2 days agoJul 28, 2026 - [Perf] Hash videos by source bytes
#49607 merged
2 days agoJul 28, 2026 - [MRV2][Performance] Skip no-op FP32 logits materialization
#47711 merged
2 days agoJul 28, 2026 - [Bugfix] Restore truncate_prompt_tokens for Jina rerank/score online
#49963 merged
2 days agoJul 27, 2026 - [Bugfix] Reject pipeline parallelism for DiffusionGemma
#45828 merged
2 days agoJul 27, 2026 - [Perf] Tune LL BF16 Router GEMM
#48774 merged
2 days agoJul 27, 2026 - [Core] Fix internal LB load-balancing
#49204 merged
2 days agoJul 27, 2026 - [Model] Enable EVS for Qwen3.5
#48912 merged
2 days agoJul 27, 2026 - [Perf] Make merge attention context count a runtime argument
#48739 merged
2 days agoJul 27, 2026 - [Perf] Skip ll_bf16 router GEMM warmup for non-MoE models
#49659 merged
2 days agoJul 27, 2026 - [Bugfix]Reject invalid FlashInfer MNNVL workspaces
#49043 merged
2 days agoJul 27, 2026 - Improve Transformers modelling backend `fx` tracer
#49957 merged
2 days agoJul 27, 2026 - [KV Offloading] Per-request tier filtering with TierFilter/TierMatcher
#48123 merged
2 days agoJul 27, 2026 - [ROCm] Use backend-default dot precision for ReplaySSM
#49909 merged
2 days agoJul 27, 2026 - [ROCm] [Model] Enable TML inkling
#48841 merged
2 days agoJul 27, 2026 - [ROCm][CI] Skip three torchao tests of gfx950 until `torchao==0.18` is released
#49732 merged
2 days agoJul 27, 2026 - [CI][ROCm] Make hf-xet reconstruction safe on shared NFS
#49837 merged
2 days agoJul 27, 2026 - [Bugfix][ROCm] Use batch DMA for CPU KV cache loads
#49843 merged
2 days agoJul 27, 2026 - [communication] [bugfix] fix quickreduce acc error in cudagraph mode
#46913 merged
2 days agoJul 27, 2026 - [Bugfix][CPU] Zero-pad MoE intermediate size for grouped-gemm TP alignment
#49591 merged
2 days agoJul 27, 2026 - [Tokenizer] Use HF config for HF tokenizers
#49907 merged
2 days agoJul 27, 2026 - [XPU][CI] Use platform device in InputBatch V2 test
#49939 merged
2 days agoJul 27, 2026 - [Bugfix] Normalize sparse MLA warmup compression ratios
#49392 merged
2 days agoJul 27, 2026 - [Rust Frontend] Keep `--max-model-len` engine-owned
#49944 merged
2 days agoJul 27, 2026 - [CI][ROCm] Reduce kernel test runtime
#49915 merged
2 days agoJul 27, 2026 - [Quantization][INC]Add MXFP8 Linear Support
#47514 merged
2 days agoJul 27, 2026 - [3/N][Core][KV Connector] Support reliable partial-tail KV offload for sub-block prompts
#49502 merged
2 days agoJul 27, 2026 - [CPU][Spec Decode] Optimize GDN conv path for speculative decoding
#48577 merged
2 days agoJul 27, 2026 - [CPU][Perf] INT8 Fused MoE Kernel for Arm CPUs
#48637 merged
2 days agoJul 27, 2026 - [Hardware][Power] Add FAST_EXP for Power
#49571 merged
3 days agoJul 27, 2026 - [CI] Explicitly tear down speculative decode runners
#49910 merged
3 days agoJul 27, 2026 - [Bugfix][KV Offload][P2P] Fix EngineCore crash reconnecting to a reaped peer
#49823 merged
3 days agoJul 27, 2026 - [ROCm] Make vllm_c RMSNorm output contiguous
#49913 merged
3 days agoJul 27, 2026 - [Bugfix] Fix mHC block-M prenorm GEMM cross-row reduction carry-over
#49429 merged
3 days agoJul 27, 2026 - [XPU][CI] Add more test cases in Intel GPU CI
#49422 merged
3 days agoJul 27, 2026 - [CI] Add kimi and k3 auto-labeling rules
#49895 merged
3 days agoJul 27, 2026 - [Docs] Document NVFP4 GEMM kernel selection and Marlin weight-only fallback
#49376 merged
3 days agoJul 27, 2026 - [ROCm][CI] Keep native datasets cache off shared NFS
#49516 merged
3 days agoJul 27, 2026 - [Frontend] expose stream_interval as req sampling param
#49754 merged
3 days agoJul 27, 2026 - [XPU] Enable QK Norm + RoPE fusion pass on XPU
#49394 merged
3 days agoJul 27, 2026 - [CI][ROCm] Keep global GPU memory cleanup opt-in
#49911 merged
3 days agoJul 27, 2026 - [CI][ROCm] Reduce V1 attention test runtime
#49916 merged
3 days agoJul 27, 2026 - [Core] Fix gpu<->cpu syncs in MRV2 mamba_hybrid.py
#49736 merged
3 days agoJul 27, 2026 - [Core][KV-transfer] MoRIIO: heterogeneous TP<->DP prefill/decode read routing
#46116 merged
3 days agoJul 27, 2026 - [BugFix][MRV2] Don't create dummy requests longer than `max_model_len`
#49751 merged
3 days agoJul 27, 2026 - [Bugfix] Prevent NaN poisoning in xpu_mla_sparse for fully-masked index chunks
#48366 merged
3 days agoJul 27, 2026 - [ModelRunner V2] Support encoder-only attention
#49331 merged
3 days agoJul 27, 2026 - [CI/Build] Refresh tags before building macOS wheel
#49901 merged
3 days agoJul 27, 2026 - [Bugfix][CuMem] Make KV-cache wake cleanup tag-safe
#49857 merged
3 days agoJul 27, 2026 - [Core][Distributed] Add process-checkpoint lifecycle hooks for communicators (starting with Flashinfer)
#46877 merged
3 days agoJul 27, 2026 - [Bugfix][KV Offload] Namespace auto cache dtype by effective dtype
#49438 merged
3 days agoJul 27, 2026 - [Bugfix] Fix handling 5D KV cache in kv_postprocess_layout_on_receive
#47791 merged
3 days agoJul 26, 2026 - [KV Offload] Fix num_tokens_after_batch for different termination types
#49285 merged
3 days agoJul 26, 2026 - [Bugfix] Reject contradictory custom-op directives
#49134 merged
3 days agoJul 26, 2026 - [Bugfix][KV Offload] Bound unaligned SWA loads by physical GPU blocks
#49052 merged
3 days agoJul 26, 2026 - [KVOffload][P2P] Generic P2P secondary tier: peer lookup and serving via ParentManager
#48021 merged
3 days agoJul 26, 2026 - [Model] Support top_k and top_p sampling for DiffusionGemma
#45429 merged
3 days agoJul 26, 2026 - [Perf] Isolate MM preprocessing on its own executor
#49524 merged
3 days agoJul 26, 2026 - [CI] Fix speech correctness check rejecting improved WER
#49853 merged
3 days agoJul 26, 2026 - [Perf] DeepSeek-OCR-2 TTFT Optimize
#49531 merged
3 days agoJul 26, 2026 - [Doc] Add compile cache volume example to the Docker deployment page
#49782 merged
4 days agoJul 26, 2026 - [KV Offload] Deduplicate replicated MLA KV in the shared CPU region
#48906 merged
4 days agoJul 26, 2026 - [Bugfix][KV Offload] Namespace persistent cache by model runner
#49440 merged
4 days agoJul 26, 2026 - [Bugfix] Respect declared attention contract for ColQwen3.5 retrievers
#49372 merged
4 days agoJul 26, 2026 - [Bugfix][KVConnector] Disable cross-layer KV blocks for per-token-head quant
#49226 merged
4 days agoJul 26, 2026 - [XPU][CI] add heterogeneous TP UT
#49651 merged
4 days agoJul 26, 2026 - [CI] Compute speech WER directly with jiwer
#49773 merged
4 days agoJul 26, 2026 - [Kernel] TD operand loads for batched MoE GEMM (moe_mmk) on XPU
#46340 merged
4 days agoJul 26, 2026 - [Bugfix][KV Offloading] Defer request finalization until final store
#49671 merged
4 days agoJul 26, 2026 - [UX] DCP Topology Validation
#49777 merged
4 days agoJul 26, 2026 - [MM][CG] Support ViT CUDA Graph for Gemma-4
#46837 merged
4 days agoJul 26, 2026 - [Build] Fix for DeepEP manylinux pidfd sycall usage
#49814 merged
4 days agoJul 26, 2026 - [CI] Stabilize Pooling Rerank Equivalence Test
#49822 merged
4 days agoJul 26, 2026 - [Bugfix] Wait for the linear bias before layerwise online processing
#49805 merged
4 days agoJul 26, 2026 - [Model] Add VaultGemma via Transformers modeling backend
#49803 merged
4 days agoJul 26, 2026 - [Perf] Fix moe `reduce_scatter` perf regression by removing additional comm, 5% E2E throughput gain back.
#48763 merged
4 days agoJul 26, 2026 - Mergify message not on cancelled
#45117 merged
4 days agoJul 26, 2026 - [Feature] Add fault tolerance framework (simplified) for DP+EP external LB deployments
#44428 merged
4 days agoJul 26, 2026 - [BugFix] Increase the max supported duration for MOSS-TD
#49403 merged
4 days agoJul 25, 2026 - [Docs] Use `gen-files` for generated docs content
#49587 merged
4 days agoJul 25, 2026 - [Bugfix][MiniMax-M3] Fix token-major top-k buffer handling in Triton …
#49149 merged
4 days agoJul 25, 2026 - [Bugfix][Tool Parser] Fix dropped streaming arguments in Jamba and InternLM2 parsers
#48852 merged
4 days agoJul 25, 2026 - [CI] Stop flaky test from downloading model every time
#49800 merged
4 days agoJul 25, 2026 - feat[vLLM × v5]: Add audio support for the Transformers backend
#39330 merged
4 days agoJul 25, 2026 - Make bare `hugging_face` imports forbidden
#49726 merged
4 days agoJul 25, 2026 - [Bugfix][CI] Fix stale Mooncake lookup expectation broken by a merge race
#49802 merged
4 days agoJul 25, 2026 - [Core] Keep attention backends eligible for text-only serving of prefix-LM models
#48796 merged
4 days agoJul 25, 2026 - [Model] Remove Ouro
#49786 merged
4 days agoJul 25, 2026 - [Perf][V1] Skip LRU hash-split in free_blocks when prefix caching is off
#48017 merged
4 days agoJul 25, 2026 - [Docs] Fix confusing docstring indentation in nemotron_h.py
#49781 merged
4 days agoJul 25, 2026 - [multimodal] Make PyNvVideoCodec decoder concurrency configurable
#49753 merged
4 days agoJul 25, 2026 - Stabilize GPU memory teardown between ROCm CI tests
#49242 merged
5 days agoJul 25, 2026 - [CI] fix compile test | refactor VLLM_DISABLE_COMPILE_CACHE for tests
#49770 merged
5 days agoJul 25, 2026 - [ROCm][CI] Force native compile caches onto local disk
#49763 merged
5 days agoJul 25, 2026 - [Bugfix][KV Connector][Mooncake] Keep TP-sharded Mamba state out of the KV-head dedup
#49499 merged
5 days agoJul 25, 2026 - [XPU] add warning for xpu graph limitations
#49419 merged
5 days agoJul 25, 2026 - [BUGFIX] Fix log capture in KV test
#49655 merged
5 days agoJul 25, 2026 - Revert "[Perf][GLM-5.2] Blackwell decode optimizations"
#49768 merged
5 days agoJul 25, 2026 - [ROCm][CI] Wait for ROCm VRAM to settle between compiled and eager LL…
#49739 merged
5 days agoJul 25, 2026 - [ROCm][Docker] Drop MORI_GPU_ARCHS so MoRI autodetects the device arch
#49737 merged
5 days agoJul 25, 2026 - [ROCm][CI] Fix XPASS(strict) on mixed audio embeds test
#49733 merged
5 days agoJul 25, 2026 - [KV Offload][CI] Fall back to buffered I/O without O_DIRECT; fix flaky api-server test
#49734 merged
5 days agoJul 25, 2026 - [Model] Remove Plamo2
#49729 merged
5 days agoJul 25, 2026 - [CI] Stabilize memory-sensitive compile and structured output tests
#49749 merged
5 days agoJul 25, 2026 - [Bugfix] Support non-uniform page sizes in KVBlockZeroer
#49704 merged
5 days agoJul 25, 2026 - [AMD][Bugfix][EPLB] Fix elastic EP scaling accuracy on ROCm
#47206 merged
5 days agoJul 25, 2026 - [WideEP] Update NCCL to 2.30.7 to enable DeepEPv2 in the vllm/vllm-openai image
#45321 merged
5 days agoJul 25, 2026 - [Bugfix] Register axk1 config to fix A.X-K1 init
#49727 merged
5 days agoJul 25, 2026 - Fix GLM-4.1V video placeholder token ID handling.
#49484 merged
5 days agoJul 25, 2026 - [CI] Avoid unnecessary Hugging Face metadata requests
#49508 merged
5 days agoJul 25, 2026 - [CI] Reuse loaded config for cached tokenizer
#49509 merged
5 days agoJul 25, 2026 - Add `sm_107` for Rubin
#49387 merged
5 days agoJul 25, 2026 - [CI][AMD] Deprecate DinD for MI355 tests
#49257 merged
5 days agoJul 25, 2026 - [Kernel] ReplaySSM: cache SSM inputs for faster Mamba2 standard decode
#48018 merged
5 days agoJul 25, 2026 - [Bugfix] Accept RFC 2397 parameters in base64 data URLs
#48973 merged
5 days agoJul 25, 2026 - [UX] Improve data-parallel launch validation
#49124 merged
5 days agoJul 24, 2026 - [UX] Reject incompatible nested runtime overrides
#49247 merged
5 days agoJul 24, 2026 - [ROCM] Fix AITER Fused AllReduce RMSNorm for Transformers Backend
#49673 merged
5 days agoJul 24, 2026 - [Bugfix][Benchmarks] Restore --skip-tokenizer-init with custom dataset
#49180 merged
5 days agoJul 24, 2026 - Encoder cache extension hooks
#48218 merged
5 days agoJul 24, 2026 - [Bugfix] Skip linear bias in layerwise reload to avoid corruption
#49586 merged
5 days agoJul 24, 2026 - [PD][NixlPush][Bugfix] Fix blocking handshake call on writer thread
#49221 merged
5 days agoJul 24, 2026 - Remove Quantization test parallelism
#49693 merged
5 days agoJul 24, 2026 - [Perf][GLM-5.2] Blackwell decode optimizations
#48597 merged
last weekJul 24, 2026 - Bump Transformers version to 5.14.1
#49223 merged
last weekJul 24, 2026 - [CI] Use explicit devices in IR tests
#49513 merged
last weekJul 24, 2026 - [Bugfix] Fix humming kernel crash when layer.has_bias is None
#48769 merged
last weekJul 24, 2026 - [CompressedTensors] DeepSeek4 CT Quantization Support
#41276 merged
last weekJul 24, 2026 - [Bugfix] Detect mixed precision in packed KV cache specs
#49623 merged
last weekJul 24, 2026 - [Model] Support llm-compressor Inkling NVFP4 weights
#49258 merged
last weekJul 24, 2026 - [ROCm][Quantization] Add Quark W4A8 (INT4-FP8) MoE CI coverage
#48050 merged
last weekJul 24, 2026 - [Docs] Fix broken anchor links in serving/pooling/MoE docs
#49654 merged
last weekJul 24, 2026 - [ROCm][CI] Prepare AMD mirrors for regating
#49270 merged
last weekJul 24, 2026 - [BugFix][LoRA] Skip marlin-backend gpt-oss LoRA tests on XPU
#49385 merged
last weekJul 24, 2026 - [CI][PD] Add hybrid SSM P_TP>D_TP accuracy sweep entry
#49593 merged
last weekJul 24, 2026 - [CI] Disable reasoning in Responses smoke test
#49511 merged
last weekJul 24, 2026 - [ROCm][CI] Language Models tests tiny-mixtral with aiter fix
#49551 merged
last weekJul 24, 2026 - [Bugfix][KV cache] Support sparse-MLA targets with SWA drafts
#48776 merged
last weekJul 24, 2026 - [Bug] Fix batch invariance rms norm comparison
#49603 merged
last weekJul 24, 2026 - [Bugfix] Restore structured output logger initialization
#49626 merged
last weekJul 24, 2026 - [Core] Update PyTorch to 2.13.0, torchvision to 0.28.0, triton to 3.7.1
#48155 merged
last weekJul 24, 2026 - [CI] Bump PyTorch Compilation Unit Tests timeout to 150 min
#49606 merged
last weekJul 24, 2026 - [CI][Bugfix] Fix test isolation in block_int8/ptpc_fp8 MoE kernel tests
#49609 merged
last weekJul 24, 2026 - [CI] Isolate cudagraph tests in child processes
#49510 merged
last weekJul 24, 2026 - [DSv4 Perf] Skip topk and router when not needed, 3.4% E2E TTFT improvement for Decode case
#49486 merged
last weekJul 24, 2026 - [CI] Use explicit devices in quantization tests
#49512 merged
last weekJul 24, 2026 - Revert "[MRV2] Always build attn metadata at capture time" (#49364)
#49451 merged
last weekJul 24, 2026 - [Bugfix][Structured Output][Spec Decode] Advance grammar across reasoning boundary
#44993 merged
last weekJul 24, 2026 - Fix GPTQ quantized Qwen3.5 MTP weight loading with spec decode
#48816 merged
last weekJul 23, 2026 - [Perf] Defer MM embeds loading off the event loop
#49477 merged
last weekJul 23, 2026 - [Bugfix][CI/Build] Fix Plamo2 HF runner crash on transformers v5 (_tied_weights_keys list→dict)
#44239 merged
last weekJul 23, 2026 - [Bugfix][Spec Decode] Select earliest-completing stop string in check_stop_strings
#49391 merged
last weekJul 23, 2026 - [Bugfix] handle grammar compilation failures to avoid engine crash
#47312 merged
last weekJul 23, 2026 - [Bugfix][Core] shm_broadcast: bound idle reader waits and release read slots
#45224 merged
last weekJul 23, 2026 - [MRV2][Spec Decode] Avoid rejection sampler OOM by chunking
#48630 merged
last weekJul 23, 2026 - [Core] Simplify KVBlockZeroer index tensor handling
#48399 merged
last weekJul 23, 2026 - [MooncakeStore] Re-derive full external hits on stored boundaries
#49481 merged
last weekJul 23, 2026 - [Perf][KVConnector][Mooncake] Vectorize prepare_value on the KV load path
#48531 merged
last weekJul 23, 2026 - [CPU][Docs] Update docs and dockerfile for s390x
#49523 merged
last weekJul 23, 2026 - [Misc] Use VLLMValidationError in chat_utils content-part validation
#49217 merged
last weekJul 23, 2026 - [Bugfix] Fix DeepSeek-V4 DSpark draft shared-expert padding for TP > 8
#49415 merged
last weekJul 23, 2026 - [Performance][Model] Avoid transient Inkling result allocations (performance, and OOM prevention on smaller memory configurations)
#49487 merged
last weekJul 23, 2026 - [Bugfix][Model] Remove SciPy dependency from Inkling scale planning
#49485 merged
last weekJul 23, 2026 - [Docs] Re-add Reo.dev analytics beacon
#49474 merged
last weekJul 23, 2026 - [Bugfix] Make shared NVFP4 MoE scales writable
#49489 merged
last weekJul 23, 2026 - [ROCm] Fused Shared Expert Support for AMD Quark DeepSeek-V4 Model Checkpoints
#48044 merged
last weekJul 23, 2026 - [BugFix] Handle per-group prefix-hit divergence for hybrid models with KV connector
#48425 merged
last weekJul 23, 2026 - Add quantization label automation
#49492 merged
last weekJul 23, 2026 - [Core][DSV4] Compact MXFP4 indexer KV cache and packed group overlays
#48993 merged
last weekJul 23, 2026 - [Bugfix] Exclude location-derived path vars from torch.compile cache factors
#47573 merged
last weekJul 23, 2026 - [Bugfix] Fix DeepGEMM warmup when using `FlashInferFp8DeepGEMMDynamicBlockScaledKernel`
#49467 merged
last weekJul 23, 2026 - [Bugfix][CI] Fix `topk_softplus_sqrt` no-op on non-XPU platforms
#49452 merged
last weekJul 23, 2026 - [Bugfix] Retry config read to survive concurrent HF cache refresh
#49001 merged
last weekJul 23, 2026 - [Bugfix] Restore `gather_and_maybe_dequant_cache` OOB guard
#49427 merged
last weekJul 23, 2026 - [CI] Fix stale/fragile untethered kernels-root tests
#49423 merged
last weekJul 23, 2026 - Bump Flashinfer version to 0.6.15
#48914 merged
last weekJul 23, 2026 - [Bugfix][Parser] Fix special tokens (EOS/BOS) leaking into reasoning content
#48748 merged
last weekJul 23, 2026
428 pull requests opened by 277 people
- Fix GLM-4.1V vision TP fallback for non-divisible attention heads
#49479 opened
last weekJul 23, 2026 - [Bugfix][Spec Decode] Enable Step-3.7-Flash (Step3p7) MTP on Model Runner V2
#49490 opened
last weekJul 23, 2026 - [Frontend] Add cache_salt support to Anthropic Messages API
#49498 opened
last weekJul 23, 2026 - [Performance] Skip redundant full-fit admission pass
#49500 opened
last weekJul 23, 2026 - Fix Humming W4 packing for dense INT4 weights
#49501 opened
last weekJul 23, 2026 - [Model Runner V2][Spec Decode] Optimize block verification kernels
#49503 opened
last weekJul 23, 2026 - perf(structured-output): cache Guidance grammar serialization
#49504 opened
last weekJul 23, 2026 - [Bugfix] Avoid repeated layerwise reload warning scans
#49505 opened
last weekJul 23, 2026 - [KV Offload] Carry chunk index in OffloadKey
#49506 opened
last weekJul 23, 2026 - [ROCm][CI] Use the same-build wheel in Python-only CI
#49514 opened
last weekJul 23, 2026 - [ROCm][CI] Select CPU platform for native no-GPU jobs
#49515 opened
last weekJul 23, 2026 - [Kernel][Core][Shared-PCP][KV Update] Replace PCP update collectives with direct peer stores
#49517 opened
last weekJul 23, 2026 - [CI][AMD] Retry Hugging Face tests online
#49518 opened
last weekJul 23, 2026 - [Bugfix][Model Loader] Defer post-load attention weight processing
#49519 opened
last weekJul 23, 2026 - [Bugfix] Honor per-request StructuredOutputsParams.disable_any_whitespace
#49520 opened
last weekJul 23, 2026 - [DRAFT][Bugfix][Structured Output] Let terminal grammars stop under min_tokens
#49521 opened
last weekJul 23, 2026 - [Refactor][KV Offloading] Unify fence field and move fence check after _build_store_jobs
#49522 opened
last weekJul 23, 2026 - [Bugfix] Shared MM text-LoRA mapper fallback for language_model wrappers
#49525 opened
last weekJul 23, 2026 - [Bugfix][Frontend] Prevent server_load double-decrement on client disconnect
#49526 opened
last weekJul 23, 2026 - [XPU] Support EC connector KV Offloading on XPU
#49532 opened
last weekJul 23, 2026 - [CI] Prototype for Rust coverage
#49534 opened
last weekJul 23, 2026 - [Bugfix][Tool Parser] Fix hunyuan_a13b/xlam streaming: argument named "name" spawns a phantom tool call
#49535 opened
last weekJul 23, 2026 - [CI][Intel] Add GSM8K DeepSeek-V2-Lite-Instruct-FP8 XPU eval job
#49536 opened
last weekJul 23, 2026 - [Bug]: Fused_moe dimension mismatch for Qwen mxfp4 model on ROCM
#49539 opened
last weekJul 23, 2026 - [Bug]: vLLM serving Qwen3.6-27B with multimodal inputs causes InternalServerError due to CUDA OOM despite sufficient GPU memory
#49540 opened
last weekJul 23, 2026 - [Bug]: Beam Search is slower by a factor of 1000
#49541 opened
last weekJul 23, 2026 - [Bugfix] Carry positions through ubatch metadata slicing for DBO
#49542 opened
last weekJul 23, 2026 - [XPU] Support test_cp_generation and add to Intel CI
#49543 opened
last weekJul 23, 2026 - [ROCm][Perf] gfx942: use FlyDSL fp8 MQA logits kernel (ROCm/aiter#3913)
#49544 opened
last weekJul 23, 2026 - [Core] Improve compile_sizes cudagraph padding error message
#49549 opened
last weekJul 23, 2026 - [Doc][Core] Clarify FlashInfer spec-decode PIECEWISE cudagraph downgrade warning
#49550 opened
last weekJul 23, 2026 - [Models][Quantization] Qwen3.5 MTP: honor checkpoint per-layer quantization config for draft layers
#49553 opened
last weekJul 23, 2026 - [Feature] implement return indexer topk of sparse attention
#49555 opened
last weekJul 23, 2026 - [Bugfix][NIXL] Mark transfer as failed when check_xfer_state() raises
#49557 opened
last weekJul 23, 2026 - [Bugfix][MoE] Filter packed expert weights during EP loading
#49558 opened
last weekJul 23, 2026 - [Model] Add Eagle3 support for Gemma3
#49560 opened
last weekJul 23, 2026 - [Bugfix] Remove failed transfer from inflight on NIXL poll exception (#49528)
#49561 opened
last weekJul 23, 2026 - [Feature] Make placement group wait timeout configurable
#49563 opened
last weekJul 23, 2026 - [MRv2] FlashAttention PCP support for GQA on MRv2
#49564 opened
last weekJul 23, 2026 - [Core] Make Ray placement group strategy configurable via VLLM_RAY_PG_STRATEGY
#49565 opened
last weekJul 23, 2026 - [Bugfix] Propagate retrieve failures in in-tree LMCache fallback adapter
#49566 opened
last weekJul 23, 2026 - [Bugfix][Tool Parser] Preserve whitespace in Step3 parameter values
#49567 opened
last weekJul 23, 2026 - [Doc] Make /dev/shm shared memory mount configurable in the Helm char
#49568 opened
last weekJul 23, 2026 - [MyPy][1/N] Fix mypy errors in some tests/ directories and enforce follow-imports=silent
#49570 opened
last weekJul 23, 2026 - [EPLB] Support custom initial expert maps
#49572 opened
last weekJul 23, 2026 - [EPLB] Add expert backup region and descriptor primitives
#49573 opened
last weekJul 23, 2026 - [Exp][Hybrid KV Cache] Evaluate semantic checkpoints for recurrent-state prefix caching
#49574 opened
last weekJul 23, 2026 - [Bugfix] Preserve explicit None when swapping dict values
#49575 opened
last weekJul 23, 2026 - [Rust Frontend] Extract shared chat types into `vllm-chat-types`
#49576 opened
last weekJul 23, 2026 - [Feature] Mask Replay
#49577 opened
last weekJul 23, 2026 - [Rust Frontend] Add MiMo reasoning parser
#49578 opened
last weekJul 23, 2026 - [EC Connector] Call to EC Connector update_connector_output from scheduler
#49579 opened
last weekJul 23, 2026 - [Bugfix][Frontend] Support namespace tools for harmony/gpt-oss models in Responses API
#49581 opened
last weekJul 23, 2026 - [CPU] Bump triton-cpu pin to include autotune cache-key fix (#275)
#49583 opened
last weekJul 23, 2026 - [EC Connector] Added Build Connector Worker Meta for EC Connector
#49585 opened
last weekJul 23, 2026 - [Kernel] Skip out-of-window KV blocks in paged decode sliding-window path
#49588 opened
last weekJul 23, 2026 - [Bugfix][ROCm][CI] Fix qk_norm_rope_kvcache fusion test for packed KV cache layout
#49590 opened
last weekJul 23, 2026 - [XPU] Enable XPU blockfp8 for DSv3
#49596 opened
last weekJul 23, 2026 - [XPU] Add Kimi Delta Attn(KDA) Support
#49598 opened
last weekJul 23, 2026 - [Do Not Merge!] Update vllm to point to flash-attention commit that builds FA3 with torch stable API. (Retry)
#49599 opened
last weekJul 23, 2026 - [CI] Build mamba-ssm with C++20 for torch 2.14 nightly compatibility
#49600 opened
last weekJul 23, 2026 - [Weight processing] Copy over `new_data` attributes in `replace_parameter`
#49601 opened
last weekJul 23, 2026 - [Bugfix] Hoist $defs/definitions in Cohere parser tool schema composition
#49602 opened
last weekJul 24, 2026 - fix: include zero temperature in trace attributes
#49605 opened
last weekJul 24, 2026 - [Core] Offload raw-prompt preprocessing to renderer thread pool in AsyncLLM
#49608 opened
last weekJul 24, 2026 - [Refactor] refactor humming linear and moe backends to use explicit layer configs
#49610 opened
last weekJul 24, 2026 - [Benchmark] Add probe requests to vllm bench serve
#49611 opened
last weekJul 24, 2026 - [Bugfix][Sampling] Clear empty side on thinking-budget asymmetric SWAP
#49613 opened
last weekJul 24, 2026 - Fix speculators dspark attribute loading
#49617 opened
last weekJul 24, 2026 - [Bugfix] Recompute block hashes after streaming session updates
#49619 opened
last weekJul 24, 2026 - fix(v1): prevent IMA after speculative preemption
#49620 opened
last weekJul 24, 2026 - [Bugfix][V1] Validate prompt_logprobs + kv_sharing_fast_prefill incompatibility at request admission
#49622 opened
last weekJul 24, 2026 - [Bugfix] Bump CUDA TileLang to 0.1.10
#49624 opened
last weekJul 24, 2026 - [Bugfix] Reject mixed structured output backends during validation
#49625 opened
last weekJul 24, 2026 - [Feat] Add warmup doc and migrate DSv4 to shared warmup contract
#49627 opened
last weekJul 24, 2026 - [CI][Debugging] Dump diagnostics for slow engine execution stages
#49628 opened
last weekJul 24, 2026 - HiSparse: host-resident sparse-MLA decode hot-buffering + GLM-5.2 indexCache opts
#49629 opened
last weekJul 24, 2026 - [Bugfix] NIXL: mark transfer failed when check_xfer_state raises in poll()
#49632 opened
last weekJul 24, 2026 - [Bugfix] Harmony: support namespace tool type in get_developer_message
#49633 opened
last weekJul 24, 2026 - [Model][MoE] DeepSeek-V4: add opt-in FlashInfer moe_ep expert backend
#49636 opened
last weekJul 24, 2026 - fix(kernel): restore scalar_t RMSNorm intermediate rounding boundary (#49616)
#49639 opened
last weekJul 24, 2026 - [Bugfix] Fix HF benchmark datasets crashing when --hf-split is omitted
#49641 opened
last weekJul 24, 2026 - [Model][Bugfix] step3.7-flash: bounds-check per-layer config lookups for MTP layers
#49642 opened
last weekJul 24, 2026 - [Feat][Core] Add disk offloading support to SimpleCPUOffloadConnector
#49644 opened
last weekJul 24, 2026 - [Bugfix][Core] Propagate DBO padding through grouped top-k routing
#49645 opened
last weekJul 24, 2026 - [Bugfix][Spec Decode] Route dflash on Qwen3DSparkModel arch to dspark
#49646 opened
last weekJul 24, 2026 - [Rubin] Enable NVLink all-reduce paths on SM107
#49647 opened
last weekJul 24, 2026 - [Bugfix][Tool Parser] Migrate Granite to the streaming Parser Engine
#49648 opened
last weekJul 24, 2026 - [CPU][Spec Decode] Centralize CPU causal-conv dispatch
#49650 opened
last weekJul 24, 2026 - [Bugfix][Spec Decode] Fix autoregressive draft decode capture with dynamic SD
#49652 opened
last weekJul 24, 2026 - [Bugfix][CPU] Fix LoRA no-op row aliasing in PunicaWrapperCPU
#49656 opened
last weekJul 24, 2026 - [Bugfix] Fix Qwen3-ASR realtime infinite repetition
#49658 opened
last weekJul 24, 2026 - [Bugfix][Kernel] Fix integer overflow in libtorch_stable/activation_kernels.cu
#49660 opened
last weekJul 24, 2026 - [Docs] Document ExampleConnector migration for disaggregated prefill
#49663 opened
last weekJul 24, 2026 - [XPU] [Linear] add torch as xpu linear backend
#49664 opened
last weekJul 24, 2026 - [Render][2/n] Paged shared memory storage for mm tensor ipc.
#49667 opened
last weekJul 24, 2026 - [graceful shutdown] fix http server start firstly before app signal handler register
#49668 opened
last weekJul 24, 2026 - [Bugfix] Avoid complex64 advanced indexing in pixtral vision encoder for backend portability
#49669 opened
last weekJul 24, 2026 - [Perf] Fuse mul_+moe_sum in TopKWeightAndReduceContiguous for large reductions
#49670 opened
last weekJul 24, 2026 - [Bugfix][Core] Stop zero-progress preemption cascades for deferred KV frees
#49675 opened
last weekJul 24, 2026 - [Bugfix] Fix missing f-string in Keye smart_resize aspect-ratio error message
#49676 opened
last weekJul 24, 2026 - [Perf][GLM-5.2] Store sparse physical indices in attention metadata
#49678 opened
last weekJul 24, 2026 - [Bugfix][CI/Build] Fix InternViT load crash on transformers v5 (missing all_tied_weights_keys)
#49679 opened
last weekJul 24, 2026 - [Observability] Add multimodal media load metrics
#49682 opened
5 days agoJul 24, 2026 - [Bugfix][V1] Free _re_block_ids on aborted requests with enable_return_routed_experts
#49684 opened
5 days agoJul 24, 2026 - [CI][Intel] Run lm-eval-harness large fp8 model on XPU
#49685 opened
5 days agoJul 24, 2026 - [Multimodal] Expose mm hash algothrim selection to cli args
#49686 opened
5 days agoJul 24, 2026 - add total start up time to the log
#49687 opened
5 days agoJul 24, 2026 - [Bugfix][CPU] Use C++ causal_conv1d kernels for GDN attention on non-AMX AVX-512BF16 CPUs
#49688 opened
5 days agoJul 24, 2026 - Fix KV event aggregation across workers
#49689 opened
5 days agoJul 24, 2026 - Fix duplicate HunyuanVL image boundary tokens
#49691 opened
5 days agoJul 24, 2026 - [Bugfix][MoE] Preserve TRTLLM MoE runtime state across sleep reload
#49695 opened
5 days agoJul 24, 2026 - p800 mrv2
#49697 opened
5 days agoJul 24, 2026 - [Model] Support MiniCPM-RobotTrack
#49698 opened
5 days agoJul 24, 2026 - [EPLB] Add Platform Backend interface for out-of-tree devices
#49700 opened
5 days agoJul 24, 2026 - [Rust Frontend] Extract chat renderer to a separate crate
#49703 opened
5 days agoJul 24, 2026 - Added kv cache budget in metrics endpoint
#49706 opened
5 days agoJul 24, 2026 - Fix NVIDIA DeepSeek V4 MHC warmup coverage
#49707 opened
5 days agoJul 24, 2026 - [Bugfix] Track token usage for ParsableContext
#49709 opened
5 days agoJul 24, 2026 - [Docs] Validate SmolLM2 batch invariance
#49710 opened
5 days agoJul 24, 2026 - [Rust Frontend] Add poolside_v1 reasoning parser
#49713 opened
5 days agoJul 24, 2026 - [Core][Kernel] Support dynamic activation scaling for ModelOpt NVFP4
#49715 opened
5 days agoJul 24, 2026 - [Attention] Add FlashInfer XQA decode support
#49718 opened
5 days agoJul 24, 2026 - [Attention] Add experimental ZoomKV backend
#49719 opened
5 days agoJul 24, 2026 - [Fix] Strip whitespace for VLLM_USE_MODELSCOPE env check
#49720 opened
5 days agoJul 24, 2026 - [Bugfix] Terminate background stream consumers when producer exits early
#49721 opened
5 days agoJul 24, 2026 - [Bugfix] Finalize streaming requests at terminal sentinel
#49722 opened
5 days agoJul 24, 2026 - fix(reasoning): force-open poolside_v1 on bare </think>
#49728 opened
5 days agoJul 24, 2026 - [Bugfix][Structured Output][Spec Decode] Fix async grammar bitmask alignment after draft trimming
#49738 opened
5 days agoJul 25, 2026 - [Kernel][Core][Shared-PCP][Capacity] Owner-shard KV history for up to 4x capacity
#49741 opened
5 days agoJul 25, 2026 - [Bugfix] Count reasoning tokens when the start token is absent
#49743 opened
5 days agoJul 25, 2026 - [Kernel] Fused QK-Norm + mRoPE for Qwen3-VL-class models
#49744 opened
5 days agoJul 25, 2026 - [tracing] Use explicit is not None checks for optional span parameter…
#49748 opened
5 days agoJul 25, 2026 - [Perf] RMSNorm uncontiguous support, 1.2~3.1x kernel performance improvement
#49750 opened
5 days agoJul 25, 2026 - [ROCm][MLA] Guard sparse-MLA persistent-only path for gqa_ratio=64 fp8 (#49649)
#49755 opened
5 days agoJul 25, 2026 - [Core][Shared-PCP][Output] Restore only sampled final rows after PCP prefill
#49756 opened
5 days agoJul 25, 2026 - [ROCm][MoE] Fix expert_map vs AITER expert_mask for non-AITER experts under EP
#49758 opened
5 days agoJul 25, 2026 - [Bugfix][Worker] Respect gpu_memory_utilization on integrated (unified-memory) GPUs; fail clean instead of wedging the host
#49760 opened
5 days agoJul 25, 2026 - Add Gemma4 B200 FP8 optimized path
#49761 opened
5 days agoJul 25, 2026 - [Quantization] Share online weight scales across TP
#49764 opened
5 days agoJul 25, 2026 - [Kernel] FlashInfer CuTe-DSL NVFP4 Quantization
#49775 opened
5 days agoJul 25, 2026 - [AMD][Bugfix] Disable MXFP4 Triton backend on gfx950
#49779 opened
5 days agoJul 25, 2026 - [Bugfix][Structured Output] Stage V2 grammar bitmask copies through preallocated pinned buffers
#49780 opened
5 days agoJul 25, 2026 - [Bugfix] Give DeepGEMM power-of-two FP32 scale factors for UE8M0 packing
#49784 opened
4 days agoJul 25, 2026 - [Bugfix] Reject unsupported forced tool choice in Responses API
#49785 opened
4 days agoJul 25, 2026 - [Model] Enable LoRA support for tower and connector in LlavaNextForConditionalGeneration
#49788 opened
4 days agoJul 25, 2026 - [Draft][Reload] Validate weight updates as transactions
#49789 opened
4 days agoJul 25, 2026 - [Model][NVIDIA] Route DSA models to the SM100 implementation
#49790 opened
4 days agoJul 25, 2026 - [Kernel][CUDA] Optimize small-batch decode GEMMs
#49791 opened
4 days agoJul 25, 2026 - [Kernel][SM100] Add a CuTeDSL fused query kernel
#49792 opened
4 days agoJul 25, 2026 - [Spec Decode][Perf] Optimize MTP draft decoding
#49793 opened
4 days agoJul 25, 2026 - [Feature] Add /v1/responses/input_tokens endpoint for preflight token counting
#49794 opened
4 days agoJul 25, 2026 - fix(mooncake): reject duplicate transfer_id to prevent KV-cache leak
#49796 opened
4 days agoJul 25, 2026 - Fix Gemma 4 for upcoming Transformers version
#49797 opened
4 days agoJul 25, 2026 - fix: add TQFullAttentionSpec guard in v1 _reshape_kv_cache_tensors
#49798 opened
4 days agoJul 25, 2026 - feat(parser): add xLAM tool parser to Rust frontend
#49799 opened
4 days agoJul 25, 2026 - docs: fix typo in cli.md
#49806 opened
4 days agoJul 25, 2026 - docs: fix typo in features README
#49807 opened
4 days agoJul 25, 2026 - docs: fix grammar in custom_op.md
#49808 opened
4 days agoJul 25, 2026 - [Feature][Model Runner V2] Support extract_hidden_states speculation
#49811 opened
4 days agoJul 26, 2026 - [CI] Shut down EngineCore explicitly in dynamic shapes compilation tests
#49812 opened
4 days agoJul 26, 2026 - [XPU] Prefer MXFP8 MoE XPU backend; forward softcap/ALIBI
#49813 opened
4 days agoJul 26, 2026 - [Bugfix][MiMo] Apply vision attention sinks in the window attention path
#49815 opened
4 days agoJul 26, 2026 - [Attention] Enable NVFP4 KV cache on SM120 (consumer Blackwell)
#49818 opened
4 days agoJul 26, 2026 - [Model] Add Cohere2MoE Eagle3 auxiliary hidden states
#49819 opened
4 days agoJul 26, 2026 - fix(cli): include inherited field docstrings in get_attr_docs
#49821 opened
4 days agoJul 26, 2026 - [Bugfix][Responses] FunctionTool to ChatCompletionToolsParam Rendering Parity
#49824 opened
4 days agoJul 26, 2026 - benchmarks: add speculative decoding analysis script
#49825 opened
4 days agoJul 26, 2026 - [Model] Enable batch-invariant mixed decode and prefill for Qwen GDN
#49827 opened
4 days agoJul 26, 2026 - kernels: fused silu_and_mul + dynamic per-token FP8 quantization
#49828 opened
4 days agoJul 26, 2026 - [Test] Add a Hugging Face cache artifact verifier
#49838 opened
4 days agoJul 26, 2026 - [Test][ROCm] Account for gfx950 FP8 RMSNorm rounding
#49839 opened
4 days agoJul 26, 2026 - [Bugfix] Shut down private Tensorizer engines
#49840 opened
4 days agoJul 26, 2026 - [Bugfix] Clean distributed state after worker initialization failure
#49841 opened
4 days agoJul 26, 2026 - [Model] Support native Transformers ERNIE 4.5 VL
#49842 opened
4 days agoJul 26, 2026 - [Bugfix] Pick a KV block size supported by every attention backend
#49845 opened
4 days agoJul 26, 2026 - [Kernel] ReplaySSM: cache SSM inputs for faster Mamba2 speculative decode
#49847 opened
4 days agoJul 26, 2026 - [CI] Add trusted /ci policy and catalog metadata
#49849 opened
4 days agoJul 26, 2026 - [Bugfix][KV Offload] Bound the HIT_PENDING wait so a stalled write cannot defer requests indefinitely
#49850 opened
4 days agoJul 26, 2026 - [MRV2][Multimodal] Enable encoder cuda graph for model runner v2
#49852 opened
4 days agoJul 26, 2026 - [Bugfix] Plumb SwiGLU-OAI gemm1 alpha/beta for ModelOpt NVFP4 MoE
#49854 opened
4 days agoJul 26, 2026 - [Rust Frontend] Support --allowed-local-media-path and --allowed-media-domains
#49855 opened
4 days agoJul 26, 2026 - [Bugfix][AsyncScheduler] Reset num_output_placeholders after discarding stale async frames
#49860 opened
3 days agoJul 26, 2026 - [Kernel] Guard A8 Marlin repack K alignment
#49862 opened
3 days agoJul 26, 2026 - [Bugfix][ROCm] Honor VLLM_WSL2_ENABLE_PIN_MEMORY on the ROCm platform
#49864 opened
3 days agoJul 26, 2026 - [Doc] Fix stale script references in CI contributor guides
#49866 opened
3 days agoJul 26, 2026 - Rename flashinfer_utils.py → flashinfer_moe.py for clarity
#49867 opened
3 days agoJul 26, 2026 - [Docs] Add SharedStorageConnector->ExampleConnector migration note to disagg prefill docs
#49868 opened
3 days agoJul 26, 2026 - [Model] Fix GLM-OCR MTP weight loading
#49869 opened
3 days agoJul 26, 2026 - [ROCm][Perf] Speed up single-group MoE routing
#49870 opened
3 days agoJul 26, 2026 - [ROCm][Perf] Add MI250X selective state update tuning
#49871 opened
3 days agoJul 26, 2026 - [ROCm][Perf] Speed up large MoE ReLU-squared on gfx90a
#49872 opened
3 days agoJul 26, 2026 - [Bugfix][MoE] Use backend expert count for workspace sizing
#49873 opened
3 days agoJul 26, 2026 - fix: handle Inkling non-reasoning responses
#49876 opened
3 days agoJul 26, 2026 - [FEAT] Support fast engine recovery through weight cache
#49879 opened
3 days agoJul 26, 2026 - [Bugfix] Reject custom tools on non-Harmony Responses routes
#49882 opened
3 days agoJul 26, 2026 - [XPU] WA topk_sigmoid argument mismatch for XPU platform
#49884 opened
3 days agoJul 26, 2026 - [Feature][Frontend] Make strict tool calling an explicit override
#49885 opened
3 days agoJul 27, 2026 - [Kernel] ReplaySSM: cache SSM inputs for faster Gated DeltaNet speculative decode
#49887 opened
3 days agoJul 27, 2026 - [ROCm] Support AITER paged attention on gfx90a
#49888 opened
3 days agoJul 27, 2026 - SM120 NVFP4 KV cache support + MTP cudagraph fix + KV offload crash fix
#49891 opened
3 days agoJul 27, 2026 - [Misc] Add unit test for moe_fused_mul_sum Triton kernel
#49894 opened
3 days agoJul 27, 2026 - [Bugfix] Route DSv4 sparse-indexer prefill top-k around NaN-broken kernel path on SM12x
#49897 opened
3 days agoJul 27, 2026 - [Bugfix] Reject whitespace-only bad_words in SamplingParams
#49898 opened
3 days agoJul 27, 2026 - [Bugfix] Handle non-text frames in realtime speech WebSocket
#49899 opened
3 days agoJul 27, 2026 - [Bugfix][Quantization] Match draft model quant targets under its root prefix
#49900 opened
3 days agoJul 27, 2026 - [CI] Retry Hugging Face processor loading
#49908 opened
3 days agoJul 27, 2026 - [Frontend] Lazily initialize chat media connectors
#49914 opened
3 days agoJul 27, 2026 - [Core] Explicitly manage torch CPU threads in workers
#49919 opened
3 days agoJul 27, 2026 - [Bugfix] Fix llama3_json parser dropping parallel and prose-adjacent tool calls
#49923 opened
3 days agoJul 27, 2026 - [WIP] Switch to the Rock, Keep Python 3.12 and Ubuntu 22.04
#49925 opened
3 days agoJul 27, 2026 - [Bugfix][Kernel] Keep DeepGEMM FP8 quant inside opaque op for TMA scales
#49928 opened
3 days agoJul 27, 2026 - [Model Loader] Introduce custom weight copy
#49929 opened
3 days agoJul 27, 2026 - fix: Add fp8 einsum warmup
#49930 opened
3 days agoJul 27, 2026 - [Linear] [Kernel] add block-wise scaled_mm
#49932 opened
3 days agoJul 27, 2026 - [Bugfix] Gate MiniMax M3 warmup imports by model and platform
#49933 opened
3 days agoJul 27, 2026 - [1/N] Unify multiple-path encoder cuda graph support
#49934 opened
3 days agoJul 27, 2026 - [Doc] Document FP8 GEMM kernel selection and Blackwell support
#49936 opened
3 days agoJul 27, 2026 - [ROCm] Add AITER FP8 ViT encoder attention
#49937 opened
3 days agoJul 27, 2026 - Fix MTP position-zero embeddings
#49938 opened
3 days agoJul 27, 2026 - [Bugfix] Forward SwiGLU alpha/beta to the Marlin NVFP4 MoE quant config
#49941 opened
3 days agoJul 27, 2026 - [CPU] Add CPU FP8 W8A8 linear/MoE support
#49942 opened
3 days agoJul 27, 2026 - [XPU] Support mxfp4 Sequence Parallelism on XPU
#49943 opened
3 days agoJul 27, 2026 - Fix DoS via sample-rate forgery bypassing audio decode duration guard
#49948 opened
2 days agoJul 27, 2026 - [Bugfix] cumem_allocator: return None via Py_IncRef, not Py_RETURN_NONE
#49950 opened
2 days agoJul 27, 2026 - Fix ModelOpt NVFP4 TP shard alignment
#49951 opened
2 days agoJul 27, 2026 - [ROCm][AITER] Add GDN long-prefill split-QKV fast path
#49953 opened
2 days agoJul 27, 2026 - [Metrics] Export EPLB rebalancing state and events
#49956 opened
2 days agoJul 27, 2026 - Support Voxtral audio generation in the Transformers backend
#49958 opened
2 days agoJul 27, 2026 - [CPU] Fix torch.compile crash from torch.accelerator.synchronize on CPU-only hosts
#49960 opened
2 days agoJul 27, 2026 - [Bugfix][KV Offload] Fix hybrid DCP block geometry
#49962 opened
2 days agoJul 27, 2026 - [Bugfix][KV Connector] Clip DCP block lists by effective span
#49965 opened
2 days agoJul 27, 2026 - fix(responses): reject code_interpreter tool call missing "code" field
#49967 opened
2 days agoJul 27, 2026 - [Bugfix][KV Connector] Fix MoRIIO DCP block span
#49968 opened
2 days agoJul 27, 2026 - [Spec Decode] Add top-k DSpark Markov projection
#49969 opened
2 days agoJul 27, 2026 - [Bugfix][KV Connector] Fix HF3FS DCP block span
#49970 opened
2 days agoJul 27, 2026 - [Bugfix][KV Events] Report DCP cache hits at effective span
#49971 opened
2 days agoJul 27, 2026 - [Bugfix] Drain all response_mqs on FT failure to prevent residual messages
#49976 opened
2 days agoJul 27, 2026 - [Perf][Multimodal] Speed up special mm token check
#49978 opened
2 days agoJul 27, 2026 - [Kernel] Add SM90 W4AFP8 grouped MoE CUTLASS operator
#49979 opened
2 days agoJul 27, 2026 - [Feat] Add request-level preemption count histogram metric
#49984 opened
2 days agoJul 27, 2026 - bugfix: fp8 MLA with Marlin kernels
#49989 opened
2 days agoJul 27, 2026 - Resolve revision to commit_hash once per model load (when loading from HF Hub)
#49990 opened
2 days agoJul 27, 2026 - [ROCm] Add VLLM_ROCM_CLONE_MMAP_WEIGHTS to avoid slow H2D copies from mapped weights
#49991 opened
2 days agoJul 27, 2026 - [EC Connector] EC Offloading Connector use events instead of StepTracker
#49994 opened
2 days agoJul 27, 2026 - fix: reject string schemas that mix pattern/format with length bounds
#49996 opened
2 days agoJul 28, 2026 - 3a80b-kda-attnres-dsrouting
#49997 opened
2 days agoJul 28, 2026 - [Rust Frontend] Support EngineCore reattachment across frontend restarts
#49998 opened
2 days agoJul 28, 2026 - [New model] Kimi K3
#50000 opened
2 days agoJul 28, 2026 - [Core] Zero KV cache when NaN logits detected
#50002 opened
2 days agoJul 28, 2026 - [Feat][ROCm] Implement get_all_gpu_pci_bus_ids for ROCm platform
#50003 opened
2 days agoJul 28, 2026 - [Bugfix][DCP] Fix NVIDIA DeepSeek-V3.2 / GLM-5.2 fused attention
#50005 opened
2 days agoJul 28, 2026 - [ROCm] Add tuned selective_state_update float16 config for AMD Instinct MI325X
#50006 opened
2 days agoJul 28, 2026 - [ROCm] Add tuned selective_state_update float32 config for AMD Instinct MI325X
#50007 opened
2 days agoJul 28, 2026 - [rocm] perf: drop redundant -inf prefill of decode paged MQA-logits buffer
#50008 opened
2 days agoJul 28, 2026 - [DCP][Performance] Add Shared-DCP peer-addressable decode data paths
#50009 opened
2 days agoJul 28, 2026 - [DCP][Kernel] Add Shared-DCP consumer-direct Top-K
#50010 opened
2 days agoJul 28, 2026 - [Bugfix][KV Offload] Protect unread promotions from eviction by other speculative promotions
#50014 opened
2 days agoJul 28, 2026 - [Bugfix][Parser] gemma4: seed non-streaming parsing from the prompt
#50015 opened
2 days agoJul 28, 2026 - [HelionLinearBackend][4/N] Add HelionINT8ScaledMMLinearKernel
#50016 opened
2 days agoJul 28, 2026 - Chunked prefill paged decode masked load perf [ROCm] [bugfix]
#50017 opened
2 days agoJul 28, 2026 - Enable ModelOpt FP8 emulation on SM80
#50019 opened
2 days agoJul 28, 2026 - [Bugfix][MRV2] Support encoder timing stats in model runner V2
#50020 opened
2 days agoJul 28, 2026 - [Bugfix] Bound two num_accepted-derived indices in the GDN spec-decode path (MTP + hybrid GDN CUDA illegal access)
#50021 opened
2 days agoJul 28, 2026 - [Bugfix][Attention] Size FlashInfer paged-KV buffers for local attention
#50022 opened
2 days agoJul 28, 2026 - [Bugfix] Prevent get_open_port() from spinning forever when VLLM_PORT is in the DP reserved range
#50025 opened
2 days agoJul 28, 2026 - [Bugfix] Reject stream and tools in batch chat completions instead of returning empty output
#50027 opened
2 days agoJul 28, 2026 - [Quantization] Fix online MoE weight preparation
#50029 opened
2 days agoJul 28, 2026 - [Quantization] Add per-token NVFP4 CuTe-DSL MoE backend
#50030 opened
2 days agoJul 28, 2026 - [Attention][MiniMax-M3] Add MSA speculative decode verification
#50032 opened
2 days agoJul 28, 2026 - feat(grpc): add KV event source discovery
#50033 opened
2 days agoJul 28, 2026 - [Bugfix] Coordinate scheduler with process checkpoint hooks
#50036 opened
2 days agoJul 28, 2026 - Fix/granitemoe fp8 split experts rebased
#50037 opened
2 days agoJul 28, 2026 - [XPU] Gate FlashAttention-in-graph on oneAPI 2026.0+ runtime support
#50038 opened
2 days agoJul 28, 2026 - Add timeout enforcement to wait_for_engine_startup
#50039 opened
2 days agoJul 28, 2026 - [Perf] Use Triton moe backend for tensor fp8 quant scheme on Hopper
#50040 opened
2 days agoJul 28, 2026 - fix(entrypoints): reject stream=True and tools in BatchChatCompletionRequest
#50042 opened
2 days agoJul 28, 2026 - fix(utils): advance next_port when candidate_port is in reserved DP range in get_open_port
#50043 opened
2 days agoJul 28, 2026 - fix(metrics): add graceful fallback for MultiProcessCollector initialization failure
#50044 opened
2 days agoJul 28, 2026 - Backpressure
#50045 opened
2 days agoJul 28, 2026 - [Bugfix][NIXL] Release a dead peer's NIXL state without waiting for TTL
#50047 opened
2 days agoJul 28, 2026 - [XPU][INC] Add configurable backend selection for W4A16 on XPU
#50048 opened
2 days agoJul 28, 2026 - fix(dp): avoid get_open_port() infinite loop when VLLM_PORT is DP-reserved
#50049 opened
2 days agoJul 28, 2026 - [Bugfix] Fix indexer expanded_block_table_buffer size mismatch
#50050 opened
2 days agoJul 28, 2026 - docs: clarify post-generation checks vs logits processors
#50051 opened
2 days agoJul 28, 2026 - perf(rocm): fuse Kimi-K3 KDA decode gate
#50054 opened
2 days agoJul 28, 2026 - [Kimi-K3] DCP support
#50055 opened
2 days agoJul 28, 2026 - [Rust Frontend] Prevent invalid token IDs in random benchmarks
#50058 opened
2 days agoJul 28, 2026 - [Model Runner V2][Spec Decode] Add KV cache support for multi-layer MTP
#50062 opened
2 days agoJul 28, 2026 - [Kimi-K3] MoRIIO KV transfer of hybrid mamba/KDA (conv+ssm) recurrent state (READ + WRITE)
#50063 opened
2 days agoJul 28, 2026 - [Refactor][PCP] Make PCPManager construction extensible
#50066 opened
2 days agoJul 28, 2026 - [Model] Enable Qwen3.8 for AMD Rocm
#50068 opened
2 days agoJul 28, 2026 - [CUDA] Select the canonical runtime for host-registration fallback
#50070 opened
2 days agoJul 28, 2026 - [Misc][Docs] Add comprehensive Responses API documentation and unit tests
#50071 opened
2 days agoJul 28, 2026 - warmup: surface async CUDA errors on GB10 (sm_121) by adding explicit synchronize points
#50072 opened
2 days agoJul 28, 2026 - [Bugfix][Quantization] Reuse online NVFP4 MoE kernel across reloads
#50074 opened
2 days agoJul 28, 2026 - [Core] Decide layerwise reload completion by manifest, not element count
#50075 opened
2 days agoJul 28, 2026 - [Kimi K3] Add EPLB support
#50076 opened
2 days agoJul 28, 2026 - [Model] Add native RWKV7 serving, fused execution, and quantization
#50077 opened
2 days agoJul 28, 2026 - [Bugfix][Frontend] Handle malformed InternLM2 tool output
#50078 opened
2 days agoJul 28, 2026 - [Bugfix][Spec Decode] Scope MTP completeness checks for Kimi K3
#50080 opened
2 days agoJul 28, 2026 - [Bugfix] Add Kimi K3 MoE support to benchmark_moe.py
#50082 opened
2 days agoJul 28, 2026 - [KV Offload] Per-request store strategy hook (covers #42050 + admission seam)
#50087 opened
2 days agoJul 28, 2026 - [MoE] Latent-MoE runner, fused MoE tail, SituGLU activation and single-group top-k routing
#50088 opened
2 days agoJul 28, 2026 - [Refactor] reuse GateLinear eligibility for ll_bf16 router warmup
#50091 opened
yesterdayJul 28, 2026 - [Kernel] Verify s4 ldmatrix layout
#50096 opened
yesterdayJul 28, 2026 - [XPU] fix_test_worker_memory_snapshot_for_xpu
#50097 opened
yesterdayJul 28, 2026 - [Rust frontend] treat stream_interval as no-op instead of unsupported
#50101 opened
yesterdayJul 28, 2026 - [ROCm][Perf] DeepSeek-V4 skip ragged layout packing for topk indices in sparse attn prefill
#50106 opened
yesterdayJul 28, 2026 - Support mm_processor_cache in the Transformers multimodal backend
#50107 opened
yesterdayJul 28, 2026 - [Kernel][CI] `--jit-monitor-mode error` e2e tests for kernel warmup infra
#50109 opened
yesterdayJul 28, 2026 - [MoE] Consolidate expert placement strategy resolution
#50113 opened
yesterdayJul 28, 2026 - [Core][Perf] Skip no-op block-hasher calls between full-block boundaries
#50114 opened
yesterdayJul 28, 2026 - fix(security): Validate ZMQ frame count in MooncakeConnector sender listener
#50115 opened
yesterdayJul 28, 2026 - [chore] clean-up weight prepack for INT8 MoE
#50116 opened
yesterdayJul 28, 2026 - [Docs] Fix ECExampleConnector references in disagg_encoder docs
#50117 opened
yesterdayJul 28, 2026 - fix(security): use effective block size in DCP full KV cache report
#50118 opened
yesterdayJul 28, 2026 - [ROCm] Enable pinned memory on supported WSL2 kernels
#50126 opened
yesterdayJul 28, 2026 - [CPU] Migrate unquantized MoE to the modular-kernel experts structure
#50133 opened
yesterdayJul 28, 2026 - [Core] Add KV event state snapshots
#50134 opened
yesterdayJul 28, 2026 - [Bugfix] Don't transpose fused MoE quantization scales in `RoutedExperts.load_weights`
#50137 opened
yesterdayJul 28, 2026 - Codex/kimi k3 dspark pp
#50138 opened
yesterdayJul 28, 2026 - WIP: FlashInfer ReplaySSM kernel
#50140 opened
yesterdayJul 28, 2026 - [Misc] Clarify mono audio requirement
#50141 opened
yesterdayJul 28, 2026 - [ROCm][CI] Relax DeepEP fp8 dispatch tolerance on gfx950
#50145 opened
yesterdayJul 28, 2026 - [Core] Support classical hybrid draft models (short_conv + attention) for speculative decoding
#50146 opened
yesterdayJul 28, 2026 - [Attention]: Use KVCacheSpec for AttentionMetadataBuilder type hints
#50148 opened
yesterdayJul 28, 2026 - [Bugfix] Fix sequence-parallel detection for small decode batches
#50155 opened
yesterdayJul 29, 2026 - [Cohere][Spec Decode] Add CohereEagleProposer to support multi layer eagle drafts
#50156 opened
yesterdayJul 29, 2026 - [Kernel] Add support for Flashinfer Mamba SSU algorithm selection
#50157 opened
yesterdayJul 29, 2026 - [Feature] Support pre-computed mm processor outputs
#50160 opened
yesterdayJul 29, 2026 - [ROCm][Kimi-K3] Model enablement - aiter moe environment variable cleanup
#50162 opened
yesterdayJul 29, 2026 - [Frontend][Rust] Support mm_processor_kwargs in chat completions
#50164 opened
yesterdayJul 29, 2026 - docs: Add Apple Silicon/Metal execution requirements to quickstart
#50167 opened
yesterdayJul 29, 2026 - LUT-B quantization accuracy prototype
#50168 opened
yesterdayJul 29, 2026 - [Core][Spec Decode] Fix KV cache allocation for sliding-window drafters and local-attention pool sizing
#50169 opened
yesterdayJul 29, 2026 - [Bugfix][Spec Decode] Warn when Eagle3 aux layers fall back to a default
#50170 opened
yesterdayJul 29, 2026 - [Feature] Qwen3-Next (GDN): mamba_cache_mode="all" prefix caching with speculative decoding (MTP) on V1
#50172 opened
yesterdayJul 29, 2026 - [Bugfix][Core][Spec Decode] Skip spec-decode block reservation on the external-KV-load step (P/D)
#50173 opened
yesterdayJul 29, 2026 - [3/N][Feat][Perf] Add new warmup infrastructure for JITs. Add provider registry and orchestration for JIT warmup
#50174 opened
yesterdayJul 29, 2026 - [1/N][warmup][DSv4] Migrate shared NVIDIA and MLA kernels
#50175 opened
yesterdayJul 29, 2026 - [2/N][warmup][DSv4] Migrate attention kernels
#50176 opened
yesterdayJul 29, 2026 - [3/N][warmup][DSv4] Migrate MoE kernels
#50177 opened
yesterdayJul 29, 2026 - [4/N][warmup][DSv4] Migrate MHC TileLang kernels
#50178 opened
yesterdayJul 29, 2026 - [Bugfix] Skip generic required/named tool grammar for parsers that set supports_required_and_named
#50180 opened
yesterdayJul 29, 2026 - [Bugfix][MLA] Fix fp8 KV cache prefill query quantization selection for Kimi-K3
#50181 opened
yesterdayJul 29, 2026 - [Bugfix] Only require symm-mem multicast when multimem is the selected algorithm
#50182 opened
20 hours agoJul 29, 2026 - [Bugfix][Spec Decode] Fix NaN handling in rejection sampler tl.argmax
#50183 opened
20 hours agoJul 29, 2026 - attn_res kernel latency improvements
#50185 opened
19 hours agoJul 29, 2026 - Bump the minor-update group across 1 directory with 173 updates
#50186 opened
19 hours agoJul 29, 2026 - [Bugfix][Spec Decode] Fix misspelled EAGLE3 aux-hidden-state layer method overrides
#50187 opened
19 hours agoJul 29, 2026 - [Frontend] Use VLLMValidationError for batch request URL validation
#50191 opened
18 hours agoJul 29, 2026 - [Bugfix] Stabilize V1 GPU model runner initialization
#50192 opened
18 hours agoJul 29, 2026 - [Helion] Improve fused QK config selection
#50193 opened
17 hours agoJul 29, 2026 - [Frontend] Add stateless /v1/responses/render endpoint
#50195 opened
16 hours agoJul 29, 2026 - [Bugfix] Fix Marlin MoE W13 scale permutation under tensor parallelism
#50196 opened
16 hours agoJul 29, 2026 - [Rust Frontend] Add MiniCPM5 XML tool parser
#50198 opened
15 hours agoJul 29, 2026 - [XPU][CI]Adjust Samplers test ENV for Intel GPU
#50199 opened
15 hours agoJul 29, 2026 - [Bugfix][Rust Frontend] Select earliest-completing stop string
#50200 opened
15 hours agoJul 29, 2026 - [Kernel] Harden top_k_per_row against NaN and under-filled output
#50201 opened
15 hours agoJul 29, 2026 - trtllm fp8 moe sm100 compatibility
#50205 opened
14 hours agoJul 29, 2026 - [XPU][CI]Add back tests/v1/e2e/general/test_correctness_sliding_window.py::test_sliding_window_retrieval[True-1-5-google/gemma-3-1b-it]
#50207 opened
14 hours agoJul 29, 2026 - [Bugfix][KV Connector][Mooncake] Preserve HMA region strides
#50208 opened
13 hours agoJul 29, 2026 - [XPU] Fix inc int4 model
#50209 opened
13 hours agoJul 29, 2026 - [ROCm][Perf][Model] Fuse Qwen3-VL attention prologue into single AITER kernel
#50212 opened
11 hours agoJul 29, 2026 - [Doc]: Update Quickstart documentation with Apple Silicon / Metal execution
#50213 opened
11 hours agoJul 29, 2026 - [Bugfix][Quantization] Fix fused AutoGPTQ overrides and mixed-group MoE loading
#50214 opened
11 hours agoJul 29, 2026 - [Bugfix][CI/Build] Honor VLLM_PRECOMPILED_WHEEL_COMMIT=nightly
#50216 opened
11 hours agoJul 29, 2026 - [Cleanup] Remove orphaned output_text_buffer_length and stale spec-decode warning after V0 removal- #50225
#50218 opened
10 hours agoJul 29, 2026 - [CPU][s390x] Optimize inference perf and add oneDNN INT8 GEMM for s390x
#50219 opened
10 hours agoJul 29, 2026 - Fix MoE fused sum row offsets
#50220 opened
10 hours agoJul 29, 2026 - fix(security): enforce audio decode duration limit in NanoNemotronVL
#50221 opened
10 hours agoJul 29, 2026 - [Docs] Add SharedStorageConnector→ExampleConnector migration note and fix connector inventory
#50223 opened
9 hours agoJul 29, 2026 - [Docs] Update Quickstart with Apple Silicon/Metal compatibility note and safetensors-compatible default model
#50224 opened
9 hours agoJul 29, 2026 - [Cleanup] Remove orphaned output_text_buffer_length and stale spec-decode startup warning after V0 removal
#50225 opened
9 hours agoJul 29, 2026 - [Kernel] Optimize SM103 GDN causal-conv prefills
#50226 opened
9 hours agoJul 29, 2026 - [Bugfix] Avoid non-contiguous CPU tensors in Llama-4 ModelOpt weight …
#50227 opened
9 hours agoJul 29, 2026 - [Bugfix][Kimi K3] Fix message-level tools, response_format passthroug…
#50228 opened
9 hours agoJul 29, 2026 - [Parser] Migrate Kimi K3 to Parser Engine
#50229 opened
9 hours agoJul 29, 2026 - [Perf][CUDA] Programmatic dependent launch for the DSA decode kernels
#50230 opened
9 hours agoJul 29, 2026 - [Docs] Add the GLM-5.2 low-latency B300 TP8 MTP=5 serving recipe
#50231 opened
9 hours agoJul 29, 2026 - [Model][Perf] Overlap mixed GDN decode and prefill recurrent kernels
#50233 opened
9 hours agoJul 29, 2026 - [PD][PushConnector] Record last activity of remotes to allow clean up of stale ones
#50234 opened
8 hours agoJul 29, 2026 - [XPU] update warning of XPU Graph
#50236 opened
8 hours agoJul 29, 2026 - [TEST ONLY] Validate /ci retry on a failed Buildkite job
#50238 opened
8 hours agoJul 29, 2026 - [Bugfix] Count prompt-opened Poolside reasoning tokens
#50240 opened
7 hours agoJul 29, 2026 - K3 DSpark AR fusion
#50242 opened
7 hours agoJul 29, 2026 - [Build] Fix CUDA release wheel builds
#50243 opened
7 hours agoJul 29, 2026 - [torch.compile] Compile `CustomOp.forward_native` for ReLU^2 to avoid raw torch ops inside opaque custom ops
#50244 opened
7 hours agoJul 29, 2026 - [Feature]: Add DFly speculative decoding and D-Cut support(Related to issue 50245)
#50246 opened
7 hours agoJul 29, 2026 - [Bugfix][ROCm] Reject incompatible mixed KV-cache layouts
#50247 opened
7 hours agoJul 29, 2026 - [Bugfix] Fix TurboQuant cache dtype propagation and FP8 store on Ampere
#50248 opened
7 hours agoJul 29, 2026 - [Bugfix] Flatten >2D multimodal embeddings, not just 3D
#50250 opened
6 hours agoJul 29, 2026 - [Draft][Reload] Add manifest-driven load receipts and update scopes
#50251 opened
6 hours agoJul 29, 2026 - [Bugfix][Gemma4] Map stacked expert LoRA tensors onto the MoE parent module
#50252 opened
6 hours agoJul 29, 2026 - fix(entrypoints): complete VLLMValidationError migration in chat_utils.py
#50254 opened
6 hours agoJul 29, 2026 - fix(entrypoints): migrate Responses API validation errors to VLLMValidationError
#50257 opened
6 hours agoJul 29, 2026 - Docs/output token control sampling params
#50260 opened
5 hours agoJul 29, 2026 - Fix UVA fallback for environments without pinned memory support (e.g. WSL2)
#50261 opened
5 hours agoJul 29, 2026 - [Bugfix] Strip Gemma4 trailing sentinels regardless of order
#50263 opened
5 hours agoJul 29, 2026 - [Feature] add the option to continue serving unfinished requests after the recovery
#50265 opened
5 hours agoJul 29, 2026 - [CI] KimiLinear PD in nightlies
#50266 opened
4 hours agoJul 29, 2026 - [Hardware][AMD] Enable fused bf16→fp32 router GEMM on ROCm
#50268 opened
4 hours agoJul 29, 2026 - [Bugfix] Fix speculative decoding for short_conv (LFM2) models
#50272 opened
4 hours agoJul 29, 2026 - [Quantization] Honor `--linear-backend` for ModelOpt W4A16
#50273 opened
4 hours agoJul 29, 2026 - Fix Qwen Omni timing for presampled videos
#50274 opened
3 hours agoJul 29, 2026 - [Bugfix][EC Connector] Don't stop an encoder-instance request before its images are encoded
#50275 opened
3 hours agoJul 29, 2026 - [Bugfix] Fix packed KV block zeroing stride
#50276 opened
3 hours agoJul 29, 2026 - [Bugfix] Make Rust HF dataset endpoint configurable
#50277 opened
3 hours agoJul 29, 2026 - [Bugfix][LoRA] Make Punica shrink deterministic with split_k=1
#50278 opened
3 hours agoJul 29, 2026 - [Kernel] Fuse strided GDN RMSNorm gating on SM103
#50279 opened
3 hours agoJul 29, 2026 - [Bugfix][Spec Decode] Make EAGLE weight sharing TP-consistent
#50280 opened
3 hours agoJul 29, 2026 - [Core] Add --enable-nan-fault-tolerance for NaN detection and request abort
#50283 opened
3 hours agoJul 29, 2026 - [CI] Stabilize speculator memory teardown
#50284 opened
2 hours agoJul 30, 2026 - [Refactor] Remove multiple dead codes
#50285 opened
2 hours agoJul 30, 2026 - [Bugfix][Model Runner V2] Preserve Mamba block table capacity under DCP
#50287 opened
2 hours agoJul 30, 2026 - [SM120] Add NVFP4 KV cache support for consumer Blackwell (RTX 5090)
#50288 opened
2 hours agoJul 30, 2026 - [Rust Frontend] Add standalone Rust renderer
#50289 opened
1 hour agoJul 30, 2026 - [Model] Step3: guard tie_word_embeddings access for pipeline parallelism
#50290 opened
1 hour agoJul 30, 2026 - [Bugfix] Preserve truncated Kimi K3 reasoning
#50291 opened
1 hour agoJul 30, 2026 - [Model Runner V2] Enable encoder token classification
#50293 opened
1 hour agoJul 30, 2026 - [Kernel][Model] Optimize FA4 mm_prefix range lookup
#50294 opened
1 hour agoJul 30, 2026 - [Test] Add per-parser tool-call format conformance tests
#50296 opened
1 hour agoJul 30, 2026 - [BugFix] Fix P/D preemption race condition
#50297 opened
52 minutes agoJul 30, 2026 - [DSv4 Perf] Remove redundant full kernel for dsv4, 1.88x kernel performance improvement
#50298 opened
47 minutes agoJul 30, 2026 - [Core] Add TP-invariant tree kernels across TP sizes
#50299 opened
36 minutes agoJul 30, 2026 - [Security] Fix chat template resource-exhaustion DoS (GHSA-4hhp-h66f-…
#50300 opened
36 minutes agoJul 30, 2026 - [KV Offload] Enable single-copy MLA layout for CPUOffloadingSpec
#50301 opened
27 minutes agoJul 30, 2026
155 issues closed by 43 people
- [Bug]: EngineCore dies silently during warmup on GB10/sm_121 (dense model, FLASH_ATTN); avoided by CUDA_LAUNCH_BLOCKING=1
#50067 closed
3 hours agoJul 29, 2026 - [Bug]: verbose_json returns duration as a string instead of a number
#49068 closed
4 hours agoJul 29, 2026 - [CI/Perf] Invalid JSON in serving benchmark config
#43537 closed
7 hours agoJul 29, 2026 - [Bug]: Changing VLLM_CPU_KVCACHE_SPACE drops Qwen 3.5 accuracy on AMD EPYC CPU
#46347 closed
8 hours agoJul 29, 2026 - [Bug] Qwen3_5ForCausalLM / Qwen3_5MoeForCausalLM defined but not registered — text-only checkpoints fail to load
#48060 closed
9 hours agoJul 29, 2026 - [Bug]: GLM-5.1-NVFP4 RuntimeError: The size of tensor a (3072) must match the size of tensor b (6144) at non-singleton dimension 1
#41108 closed
10 hours agoJul 29, 2026 - [Bug]: EP support for Fused MoE LoRA is not implemented yet.
#38005 closed
15 hours agoJul 29, 2026 - [Bug]: Engine V1 crash with WorkerProc leaked shared_memory on multi-GPU setup
#37900 closed
15 hours agoJul 29, 2026 - [Bug]: V1 Engine: EngineDeadError (AssertionError) on max_model_len overflow during realtime audio streaming
#38428 closed
15 hours agoJul 29, 2026 - [Bug]: `limit_mm_per_prompt` is ineffective for Qwen3-VL
#38459 closed
15 hours agoJul 29, 2026 - feat: DNS-AID SVCB endpoint capability advertisement
#38511 closed
15 hours agoJul 29, 2026 - [Feature]: Sharded model loader doesn't support GCS
#38568 closed
15 hours agoJul 29, 2026 - Bug: ValueError: too many values to unpack in dispatch_cpu_unquantized_gemm when loading Qwen3.5-4B
#38591 closed
15 hours agoJul 29, 2026 - [Bug]: ROCm ROCMAiterMLASparseMetadata missing num_decode_tokens, sparse MLA (DSA) models fail to start on MI300X on main
#50064 closed
17 hours agoJul 29, 2026 - [Bug]: DSpark launch failed with FP4 target model on GLM 5.2
#47934 closed
yesterdayJul 29, 2026 - [Build]: csrc/spinloop.cpp fails to compile under clang — direct #include <mwaitxintrin.h> rejected
#45515 closed
yesterdayJul 29, 2026 - [Bug][Multi-modal]: Video frames should be paded right by temporal_patch_size
#47866 closed
yesterdayJul 29, 2026 - [Bug][KV Offload][P2P] Symmetric fetch can remain non-terminal after confirmed lookup hits; consumer stays deferred for 30s
#49820 closed
yesterdayJul 29, 2026 - [Bug]: vLLM+LMCache GLM-5.2 inference causes engine RPCTimeout
#48801 closed
yesterdayJul 28, 2026 - [Bug] _align_hybrid_block_size produces TP-dependent block sizes, currently unsupported when local and remote kernel block size mismatch
#41037 closed
yesterdayJul 28, 2026 - [Bug]: EPD correctness test produces different output for multi-image prompts
#49692 closed
yesterdayJul 28, 2026 - [Feature]: Support Gemma 3 QAT series
#16856 closed
2 days agoJul 28, 2026 - [Performance]: Fuse padding onto GEMM by making the GEMM out-of-place
#24917 closed
2 days agoJul 28, 2026 - [Bug]: tensorizer model loading with tensor parallelism increases memory requirements
#25751 closed
2 days agoJul 28, 2026 - [Bug]: Qwen3-VL {4B,8B} FP8 on vLLM returns only exclamation marks ("!!!!!...") on Jetson Thor
#27364 closed
2 days agoJul 28, 2026 - [Bug]: Engine core proc EngineCore_DP0 died unexpectedly, shutting down client.
#27557 closed
2 days agoJul 28, 2026 - [RFC]: Supporting Multi MTP layers in Speculative Decoding (EagleProposer)
#31204 closed
2 days agoJul 28, 2026 - [Bug][Infrastructure]: Inconsistent Docker Image Versioning and Missing Tags on Docker Hub
#33748 closed
2 days agoJul 28, 2026 - [RFC]:DeepSeek-R1 Moe offload
#33869 closed
2 days agoJul 28, 2026 - [Bug]: the embeddings from vllm does not match hugging face sentence-transformer
#34910 closed
2 days agoJul 28, 2026 - [Bug]: #47327 dense-MHA split breaks FlashMLA sparse: OOB write in top-k index conversion, corrupted fp8_ds_mla context gather
#48611 closed
2 days agoJul 28, 2026 - [CI Failure]: Acceptance length issue in the Speculators Correctness group
#43000 closed
2 days agoJul 28, 2026 - [Bug]: DP request distribution becomes imbalanced under long-context workload on H20 GPU
#48808 closed
2 days agoJul 27, 2026 - [Usage]: Why was the DualChunkFlashAttentionBackend removed in subsequent versions?
#49296 closed
2 days agoJul 27, 2026 - [Feature]: Restore support for 2-bit and 3-bit GPTQ
#45051 closed
2 days agoJul 27, 2026 - [Bug]: MiniMax-M3 Multimodal Model Crashes During Inference
#49940 closed
2 days agoJul 27, 2026 - [Bug][KV Offload][P2P] EngineCore crash reconnecting to peer: stale dead ZmqConnection remains registered
#49809 closed
3 days agoJul 27, 2026 - [Bug]: MiniMax M3 Triton path misinterprets token-major top-k buffer
#48603 closed
3 days agoJul 27, 2026 - Cancelled issue. No longer needed.
#49848 closed
3 days agoJul 27, 2026 - [Installation]: undefined symbol: _ZN3c104cuda29c10_cuda_check_implementationEiPKcS2_jb
#36302 closed
3 days agoJul 27, 2026 - [RFC]: Dynamic Speculation Length (DSL) with Confidence-Threshold Early Exit for vLLM Speculative Decoding
#36657 closed
3 days agoJul 27, 2026 - [Bug]: In DP mode, waiting request stack in a few DP ranks.
#36748 closed
3 days agoJul 27, 2026 - [Bug]: Qwen3-0.6B hangs during `vllm bench throughput` on amd machine
#36989 closed
3 days agoJul 27, 2026 - [Bug]: n_completions + logprobs Causes Significant TTFT Spike for Co-Scheduled Requests on Cold Cache
#37343 closed
3 days agoJul 27, 2026 - [Performance]: Is SamplingParams support set enable_thinking?
#37527 closed
3 days agoJul 27, 2026 - [Feature]: add ParoQuant quantization
#37687 closed
3 days agoJul 27, 2026 - [RFC] Tail-Optimized LRU (T-LRU): Reducing Tail Latency via Conversation-Aware KV Cache Eviction
#37823 closed
3 days agoJul 27, 2026 - [Usage]: Unable to run Qwen3-14B with vLLM (multiple issues)
#37907 closed
3 days agoJul 27, 2026 - [Bug]: ImportError: flash_attn.ops.triton.rotary not found on older versions (< v2.1.2)
#38056 closed
3 days agoJul 27, 2026 - [Bug] scalar_types.int4 weight type not supported in Marlin kernel, making W4A8-INT models undeployable
#38063 closed
3 days agoJul 27, 2026 - [Bug] W4A8-INT compressed_tensors silently runs W4A16 — activations never quantized to int8
#38064 closed
3 days agoJul 27, 2026 - [RFC]: Can we implement n-gram and suffix speculative decoding in model_runner_v2?
#38069 closed
3 days agoJul 27, 2026 - [Bug]: MFU statistics on the Step-3.5-Flash model are inaccurate
#38170 closed
3 days agoJul 27, 2026 - [Bug]: M2.5 tool call result is badcase, deploy 1p1d with nixl connector, P and D use DP8-EP-TP1
#38203 closed
3 days agoJul 27, 2026 - [Feature]: Add Rotorquant support
#38291 closed
3 days agoJul 27, 2026 - [Bug] Step-3.5-Flash MTP Speculative Decoding Has Extremely Low Acceptance Rate (2.4%-4.6%)
#38339 closed
3 days agoJul 27, 2026 - [Bug]: 使用swift rollout启动vllm,推理结果乱码
#38349 closed
3 days agoJul 27, 2026 - [Bug]: DeepSeek V4 Pro crashes in flashmla sparse_prefill_fwd (TMA init failed) on v0.26.0, works on v0.25.x
#49883 closed
3 days agoJul 27, 2026 - [Bug]: xpu_mla_sparse NaN-poisons attention output when a row's leading topk index chunk is fully masked
#48364 closed
3 days agoJul 27, 2026 - CPU offload fatal: valid unaligned SWA external load exceeds aligned pending-block bound
#48959 closed
3 days agoJul 26, 2026 - [CI Failure]: distributed-compile - test_compile_correctness[test_setting3]
#49672 closed
4 days agoJul 26, 2026 - [Bug]: OffloadingConnector corrupts outputs with per-token-head quantized KV cache (cross-layer allocation lacks scale packing)
#48412 closed
4 days agoJul 26, 2026 - [CI Failure]: entrypoints-integration-speech_to_text - WER tests fail through evaluate HfFolder
#49771 closed
4 days agoJul 26, 2026 - [Bug]: Final KV offload store after request finalization crashes EngineCore
#49635 closed
4 days agoJul 26, 2026 - [Bug]: flashinfer cubin version mismatch
#49804 closed
4 days agoJul 26, 2026 - [Bug]: Blank output on google/vaultgemma-1b when context falls outside the last 512 tokens
#49795 closed
4 days agoJul 26, 2026 - [New Model]: microsoft/VibeVoice-ASR support
#32823 closed
4 days agoJul 25, 2026 - [Performance]: Prefix cache hit lower on vLLM than on other inference stacks
#38194 closed
5 days agoJul 25, 2026 - [Bug]: CUDA Illegal Instruction during CUDA Graph capture with Nemotron-3-Nano NVFP4 on sm_121
#38208 closed
5 days agoJul 25, 2026 - [RFC]: Support Dynamic Model Switching and Flexible Collective Communication in External Launcher Mode
#38231 closed
5 days agoJul 25, 2026 - [Bug]: VLLM_CPU_OMP_THREADS_BIND=nobind cannot be used with tp>1 on CPU backends
#38250 closed
5 days agoJul 25, 2026 - [Usage]: How to do offline inference on one rank in a distributed environment?
#38258 closed
5 days agoJul 25, 2026 - [Feature]: PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
#38279 closed
5 days agoJul 25, 2026 - [Bug]: Gemma3n concurrent audio requests crash EngineCore — missing dynamic_dims on audio sequence dimension
#38297 closed
5 days agoJul 25, 2026 - Energy Efficiency: 10 Mathematical Techniques for 60-70% AI Energy Reduction (Phi6Simple, FFT-Mix, Phi MoE)
#38298 closed
5 days agoJul 25, 2026 - [Bug]: ValueError: Gemma4ForConditionalGeneration does not support LoRA yet.
#40693 closed
5 days agoJul 25, 2026 - [Bug]: Gemma4-31B freezes on multiple RTX6000 PRO during loading
#38926 closed
5 days agoJul 25, 2026 - [Bug]: Gemma 4: Engine hang during large prefill caused by Interleaved Attention and p-RoPE implementation
#39914 closed
5 days agoJul 25, 2026 - [Bug]: Gemma 4 31B INT4 on 2×24GB GPUs (TP=2): GPU KV cache size is 25,200 tokens at max_model_len=131072, gpu_memory_utilization=0.96, BF16 KV
#39133 closed
5 days agoJul 25, 2026 - [Bug]: Regression in 0.19.1 - Gemma 4 26B MoE fails to load packed experts (KeyError: down_proj_packed). Worked in dev6.
#40591 closed
5 days agoJul 25, 2026 - [Bug]: Gemma4 multimodal: missing vision-aware bidirectional attention mask for use_bidirectional_attention="vision" models
#40106 closed
5 days agoJul 25, 2026 - sleep-mode leaks HBM for models loaded with --mm-encoder-tp-mode data (blocks multi-container sleep-swap)
#47654 closed
5 days agoJul 25, 2026 - KVBlockZeroer crashes with non-uniform page sizes on GLM-5.2 FP8
#49696 closed
5 days agoJul 25, 2026 - [Bug]: Qwen3MoE 使用 --convert embed 启动时抛出 AttributeError: 'Qwen3MoeForEmbedding' object has no attribute 'logits_processor'
#49657 closed
last weekJul 24, 2026 - [Bug]: IndexError in DeepSeek-V4 MTP module when using TP=16
#47273 closed
last weekJul 24, 2026 - [Bug]: API Connection Error after concurrent API calls
#14365 closed
last weekJul 24, 2026 - [RFC][UX]: debug mode for vLLM-compile
#20394 closed
last weekJul 24, 2026 - [Bug]: triton cache path for standalone_compile seems wrong
#24896 closed
last weekJul 24, 2026 - [Feature]: Ship some basic html UI with vllm for most basic testing (not for prod usage)
#25021 closed
last weekJul 24, 2026 - [Feature]: Will vLLM 0.11.0 provide a pre-compiled wheel for CUDA 11.8?
#26221 closed
last weekJul 24, 2026 - [Bug]: xgrammar cleanup leakage
#26363 closed
last weekJul 24, 2026 - [RFC][ResponsesAPI]: Separating State & Providing Flexibility for serving ResponsesAPI
#26934 closed
last weekJul 24, 2026 - [RFC]: Elastic DP Rank Recovery: Graceful Handling of GPU Hardware Failures Without Full Cluster Restart
#30112 closed
last weekJul 24, 2026 - [Feature]: Models with different quantization for different layers/blocks
#30921 closed
last weekJul 24, 2026 - [Bug]: accuracy issue with VLLM_USE_FLASHINFER_MOE_FP8=1 for Qwen3-Coder-480B-A35B-Instruct-FP8
#31394 closed
last weekJul 24, 2026 - [RFC]: Unified Parser for reasoning, tool calling
#32713 closed
last weekJul 24, 2026 - [RFC]: Support Chunked Pipeline Parallel(CPP) with dynamic chunk size.
#33914 closed
last weekJul 24, 2026 - [Bug]: Enable DBO on Qwen3-VL-235B-A22B raise TypeError: 'NoneType' object is not subscriptable
#34210 closed
last weekJul 24, 2026 - [Performance]: The checkpoint you are trying to load has model type `qwen3_5` but Transformers does not recognize this architecture.
#35847 closed
last weekJul 24, 2026 - [RFC]: Make kernel/op and component tests device-agnostic for OOT plugins
#36602 closed
last weekJul 24, 2026 - [Bug]: POST /wake_up causes vLLM process to crash. 500 Internal Server Error
#36753 closed
last weekJul 24, 2026 - [Bug]: EAGLE3 speculative decoding + multimodal crash under high concurrency
#36906 closed
last weekJul 24, 2026 - [Bug]: sm110: torch.AcceleratorError: CUDA error: an illegal instruction was encountered
#37060 closed
last weekJul 24, 2026 - [RFC]: Opt-in Media URL Cache for `MediaConnector`
#37075 closed
last weekJul 24, 2026 - [Usage]:wen3.5-35B-A3B (FP8) with vLLM 0.17.1 , the first request takes significantly longer than subsequent requests
#37175 closed
last weekJul 24, 2026 - [Bug]: Docker Build Failure for Dockerfile.nightly_pytorch
#37284 closed
last weekJul 24, 2026 - [Bug]: gdn prefill kernel errors
#37365 closed
last weekJul 24, 2026 - [Bug]: key error: ‘Layer.34.mlp.experts.gate_up_proj’
#37644 closed
last weekJul 24, 2026 - Bug: Speculative Decoding (MTP) Causes </think> Detection Failure in Structured Output + Reasoning Mode
#34650 closed
last weekJul 24, 2026 - [Bug]: Qwen3.5 397B GPTQ model outputs all exclamation points on ROCM
#37996 closed
last weekJul 23, 2026 - [Bug]: vLLM fails to start on RDNA 4 (gfx1201) inside containers — amdsmi, circular import, and torch.cuda.device_count() all broken
#40081 closed
last weekJul 23, 2026 - [Bug]: Ubuntu 26.04 with an AMD 9070 and ROCM, an error occurs. cannot open shared object file: No such file or directory
#49468 closed
last weekJul 23, 2026 - [Bug]: Runtime error on ROCm platform serving Deepseek-R1 using VLLM_ROCM_USE_AITER=1
#39485 closed
last weekJul 23, 2026 - [Bug]: Qwen3.5-122B-A10B-FP8 EngineCore crash on concurrent image requests
#37602 closed
last weekJul 23, 2026 - [Bug]: Gemma4MultimodalEmbedder normalization order different from Transformers, causing bad audio inference
#40095 closed
last weekJul 23, 2026 - [Bug]: Fused_moe dimension mismatch for Qwen mxfp4 model on ROCM
#49141 closed
last weekJul 23, 2026 - [Bug]:推理时报错,模型关闭了。部署的Qwen3.5-122B-A10B-FP8模型
#37392 closed
last weekJul 23, 2026 - [Bug]: Mooncake Connector: Decode nodes stuck in WAITING_FOR_REMOTE_KVS after Prefill node restart
#37745 closed
last weekJul 23, 2026 - [Performance]: Deepseek performance regressing with norm fusion enabled
#37832 closed
last weekJul 23, 2026 - [Bug] Potential incorrect tokenizer source path in RunAI object storage pull
#37836 closed
last weekJul 23, 2026 - [Feature Request] Support chat_template in tokenizer_config.json for DeepSeekV32
#37839 closed
last weekJul 23, 2026 - _update_request_as_session does not update max_tokens from StreamingUpdate
#37842 closed
last weekJul 23, 2026 - [Bug]: 0.17.0rc1在A2部署GLM-4.7,开启MTP后工具调用异常
#37846 closed
last weekJul 23, 2026 - [RFC]: Unify the function of getting device count
#37849 closed
last weekJul 23, 2026 - [Bug]: Phi qk_layernorm appears to be unsupported in vLLM
#37852 closed
last weekJul 23, 2026 - [Bug]: does not have the attribute 'FakeTensorMode'
#37858 closed
last weekJul 23, 2026 - [Bug] UVA CPU offload completely broken on WSL with NVFP4 MoE (Qwen3.5-35B-A3B): three distinct crashes across all parameter combinations
#37883 closed
last weekJul 23, 2026 - [Bug]: EBNF grammar not strictly enforced when n > 1 in parallel generation
#37893 closed
last weekJul 23, 2026 - [Bug]: Qwen3.5-35B-A3B compile cache miss 100% on subgraphs.
#37919 closed
last weekJul 23, 2026 - [Bug]: Editing `values.labels` in `chart-helm` breaks Service selector and leaves Endpoints empty
#37942 closed
last weekJul 23, 2026 - [Bug]: Minimax-M2.5 on version 0.17.0 results in an keyerror when the pipeline parallelism (PP) is greater than or equal to 2
#37946 closed
last weekJul 23, 2026 - [Bug]: spec decoding nonparallel 路径 draft/target hidden size 不兼容,建议适配不同hidden size
#37966 closed
last weekJul 23, 2026 - [Bug][Model] Eagle2.5-VL applies ImageNet normalization instead of SigLIP2
#37977 closed
last weekJul 23, 2026 - [Bug]: minimax-m2.5, reasoning token result in negative values
#37988 closed
last weekJul 23, 2026 - [Bug]: Speech-to-Text endpoint may return 501 but not documented in OpenAPI
#38004 closed
last weekJul 23, 2026 - [Bug]: Marlin MoE kernel fails with MXFP4-quantized GPT-OSS 20B - Invalid thread config for non-aligned dimensions (K=2880, N=2880)
#38022 closed
last weekJul 23, 2026 - [Bug]: CPUFusedMOE raises IndexError when running EP across CPU-only nodes
#38033 closed
last weekJul 23, 2026 - [Bug]: DiffusionGemma crashes when a one-token prefill remainder is mistaken for decode
#49495 closed
last weekJul 23, 2026 - [RFC] Persistent, in-place drafting attention metadata for full-CUDA-graph MTP/spec-decode
#49488 closed
last weekJul 23, 2026 - [Bug]: torch-nightly build fails — vllm-flash-attn (_vllm_fa2_C) built with C++17 but ATen now requires C++20
#49260 closed
last weekJul 23, 2026 - [Bug]: qwen3_coder streaming tool parser can emit incomplete JSON arguments with stream_interval > 1
#45256 closed
last weekJul 23, 2026
150 issues opened by 122 people
- [Bug]: [Parser] Kimi-K3 synthesizes response-scoped tool-call IDs (`{name}:{index}`) that repeat across turns, breaking multi-turn tool loops
#50295 opened
1 hour agoJul 30, 2026 - VLLM_ATTENTION_BACKEND env var is silently ignored — attention_backend= LLM()/EngineArgs kwarg is the only mechanism that works
#50292 opened
1 hour agoJul 30, 2026 - [Feature]: Expose startup status and health endpoint before model engine is ready
#50282 opened
3 hours agoJul 29, 2026 - [RFC]: Per-layer online quantization configuration
#50281 opened
3 hours agoJul 29, 2026 - [Bug]: Rust HF benchmark hard-codes the Dataset Viewer endpoint and can stall when it is unreachable
#50271 opened
4 hours agoJul 29, 2026 - [Bug]: Host memory is not reducing after the model is loaded into Intel XPU
#50269 opened
4 hours agoJul 29, 2026 - [Performance][ROCm]: Fused bf16→fp32 router GEMM falls back to a standalone copy kernel on ROCm
#50267 opened
4 hours agoJul 29, 2026 - [Performance]: On RDNA, hybrid-Mamba models fall back to Triton paged attention and decode collapses at long context
#50264 opened
5 hours agoJul 29, 2026 - [Doc]: Add feature documentation for undocumented vLLM-specific SamplingParams fields
#50259 opened
5 hours agoJul 29, 2026 - [Bug]: Kimi K3 parser leaks truncated reasoning into content with speculative decoding
#50258 opened
5 hours agoJul 29, 2026 - [Bug]: Responses API validation errors bypass VLLMValidationError — 4 sites in harmony_utils.py and responses/utils.py
#50256 opened
6 hours agoJul 29, 2026 - [Kernel][Model] Gemma4: optimize FA4 mm_prefix range lookup and CuTe JIT stability
#50255 opened
6 hours agoJul 29, 2026 - [Fix]: 7 user-facing errors in chat_utils.py bypass VLLMValidationError after PR #49665
#50253 opened
6 hours agoJul 29, 2026 - [Bug]: CUDA error: no kernel image is available for execution on the device when running Kimi-K3 on Ampere (A100) with vllm/vllm-openai:kimi-k3
#50249 opened
6 hours agoJul 29, 2026 - [Feature]: Add DFly speculative decoding and D-Cut support
#50245 opened
7 hours agoJul 29, 2026 - [Bug]: UVA is not available on WSL2 with RTX 5050 (sm_120) — crashes on basic model load, no CPU offload
#50239 opened
8 hours agoJul 29, 2026 - [Bug]: ROCm mixed KV-cache layouts abort speculative decoding with AssertionError
#50237 opened
8 hours agoJul 29, 2026 - [Bug]: Kimi-K3 prefix cache miss when prompt length is exactly a 1536-token boundary
#50235 opened
8 hours agoJul 29, 2026 - [Cleanup] Orphaned SamplingParams.output_text_buffer_length and stale spec-decode startup warning after V0 removal
#50217 opened
10 hours agoJul 29, 2026 - [Installation]: VLLM_PRECOMPILED_WHEEL_COMMIT=nightly is rejected
#50215 opened
11 hours agoJul 29, 2026 - [New Model]: Support GLM-5.2-Vision-NVFP4
#50206 opened
14 hours agoJul 29, 2026 - [Bug]: vllm-k3-toolcall repeat and repeat error
#50203 opened
15 hours agoJul 29, 2026 - [Bug]: Xid 31 MMU fault (illegal write) with flashinfer_b12x MoE backend under concurrent chunked prefill (SM120, Qwen3.5-122B-A10B-NVFP4)
#50189 opened
18 hours agoJul 29, 2026 - [Bug]: Explicit --enable-prefix-caching with MTP spec decode + tool calling corrupts tool-call emission on repeated identical prompts
#50188 opened
18 hours agoJul 29, 2026 - [CI Failure]: test_deep_ep_v2_moe FP8 output tolerance failure
#50184 opened
20 hours agoJul 29, 2026 - [Bug]: Multicast check disables the symmetric memory two-shot all-reduce path, which never uses multicast
#50179 opened
yesterdayJul 29, 2026 - [Usage]: Does Kimi-K2.5 support `thinking_token_budget`?
#50166 opened
yesterdayJul 29, 2026 - [Doc]: Update Quickstart documentation with Apple Silicon / Metal execution requirements and compatibility note
#50165 opened
yesterdayJul 29, 2026 - [Bug]: GLM 5.2; Garbled output with sequence-parallel MoE on small decode batches (TP+EP+DP)
#50154 opened
yesterdayJul 29, 2026 - [Feature]: Balanced MoE routing between GPUs
#50151 opened
yesterdayJul 28, 2026 - [Bug]: Host-memory leak on eager `reshape_and_cache_flash` path
#50150 opened
yesterdayJul 28, 2026 - [Feature]: remove vllm user from docker containers
#50149 opened
yesterdayJul 28, 2026 - [Bug]: Kimi-K3 (TP=8, prefix caching): recurring illegal-memory-access crashes under concurrent load
#50147 opened
yesterdayJul 28, 2026 - [Bug/Feature]: w8a8_block_int8_matmul kernel is dormant
#50143 opened
yesterdayJul 28, 2026 - [Bug]: Qwen3-Omni does not propagate effective sampled FPS to audio/video interleaving
#50142 opened
yesterdayJul 28, 2026 - [CI] Transformers backend llava-onevision test needs default_torch_num_threads=1 to avoid hanging
#50130 opened
yesterdayJul 28, 2026 - [Performance] Measure Transformers backend startup time vs native
#50128 opened
yesterdayJul 28, 2026 - [RFC]: Shared Context Parallelism: peer-addressable GPU objects for PCP and DCP
#50127 opened
yesterdayJul 28, 2026 - [Usage]: Does multimodal HF processor overlap with GPU compute in offline `LLM.generate()`?
#50111 opened
yesterdayJul 28, 2026 - [Usage]: It seems that there is no kimi-k3-cu129 docker.
#50102 opened
yesterdayJul 28, 2026 - [Feature]: Kimi K3 DSpark Pipeline Parallelism
#50098 opened
yesterdayJul 28, 2026 - [Bug][DCP] NVIDIA DeepSeek-V3.2 / GLM-5.2 fused attention bypasses DCP handling
#50095 opened
yesterdayJul 28, 2026 - [Bug]: NVFP4 KV writes V block scales in the SM100 trtllm-gen swizzle unconditionally, silently corrupting the SM120 FlashInfer XQA route
#50084 opened
2 days agoJul 28, 2026 - RFC: Kimi K3 RL support roadmap
#50079 opened
2 days agoJul 28, 2026 - [Bug]: kimi-k3 --kvcache-dtype-fp8 error
#50056 opened
2 days agoJul 28, 2026 - [RFC]: Back-pressure detection for KV cache offloading tiers
#50031 opened
2 days agoJul 28, 2026 - [Bug]: /v1/chat/completions/batch accepts stream: true and returns empty output with HTTP 200; tools are silently dropped
#50026 opened
2 days agoJul 28, 2026 - [Bug]: get_open_port() hangs forever when VLLM_PORT falls inside the data parallel reserved port range
#50024 opened
2 days agoJul 28, 2026 - [CI Failure]: LM Eval Qwen3.5 Models Accuracy issues
#50018 opened
2 days agoJul 28, 2026 - [Feature]: AMD Kimi K3 MoRI MI355X disagg
#50013 opened
2 days agoJul 28, 2026 - DSpark: num_speculative_tokens < dspark_block_size guard (from #47419) has no corresponding code path — should be a warning
#50012 opened
2 days agoJul 28, 2026 - [Bug]: Sleep mode wake_up crashes EngineCore natively on DGX Spark (GB10, unified memory) — sleep level 1 works, wake kills the engine
#50011 opened
2 days agoJul 28, 2026 - [Model Support] Kimi K3 Tracking Issue
#50001 opened
2 days agoJul 28, 2026 - [Perf] DSD arms pay a large baseline tax vs no-spec under production defaults; PIECEWISE override identified as one factor
#49986 opened
2 days agoJul 27, 2026 - [Bug]: /metrics returns 500 when PROMETHEUS_MULTIPROC_DIR resolves to a network-backed/mounted volume (TP>1)
#49983 opened
2 days agoJul 27, 2026 - [Bug]: tool_choice: "required" causes xgrammar FSM crash / infinite hang with GLM-5.2-NVFP4 on vLLM 0.26.0
#49981 opened
2 days agoJul 27, 2026 - [Feature]: Add SM90 W4AFP8 grouped MoE support
#49977 opened
2 days agoJul 27, 2026 - [Feature]: Support ubuntu 26.04 runtime container
#49973 opened
2 days agoJul 27, 2026 - [Bug]: FT recovery reads residual responses from unconsumed response_mqs after worker failure
#49972 opened
2 days agoJul 27, 2026 - [Regression] Trailing <turn|> token appearing at the end of generated text in vLLM 0.26.0
#49955 opened
2 days agoJul 27, 2026 - [Bug]: Responses code_interpreter returns HTTP 500 when tool arguments omit required "code" field in experimental ParsableContext
#49954 opened
2 days agoJul 27, 2026 - [Bug]: Nemotron 3 Nano NVFP4 plain TP8 fails on Hopper/Marlin because TP shards split 16-value quantization groups
#49949 opened
2 days agoJul 27, 2026 - [Bug]: EngineDeadError NVFP4 marlin
#49926 opened
3 days agoJul 27, 2026 - [Bug]: [Regression] Assertion res == CUresult::CUDA_SUCCESS failed in FlashMLA (phase1.cuh) for DeepSeek-V4 on v0.26.0 (Works in v0.25.0)
#49922 opened
3 days agoJul 27, 2026 - [Perf] BF16x3 router GEMM gated off family-120 Blackwell (GB10 / DGX Spark, sm_121) — the only barrier for DeepSeek-V4-Flash's fp32 router
#49921 opened
3 days agoJul 27, 2026 - [Feature]: DeepGEMM kernels are never warmed — only 2 of 24 entry points, so fp8_einsum JIT-loads during serving
#49905 opened
3 days agoJul 27, 2026 - [Bug]: Tiered KV offload promotes every waiting request (which fills primary DRAM pool)
#49902 opened
3 days agoJul 27, 2026 - [Bug] DeepSeek-V4 on SM12x: NaN MQA logits drive top_k_per_row_prefill to emit uninitialized smem as indices -> illegal memory access
#49896 opened
3 days agoJul 27, 2026 - [Bug]: SpeculativeConfig method="draft_model" cannot load mixed-precision compressed-tensors checkpoints (config_groups)
#49893 opened
3 days agoJul 27, 2026 - [Bug]: Dramatic KV cache size increase (~40%) for Gemma4 from v0.25.1 to v0.26
#49878 opened
3 days agoJul 26, 2026 - [Bug]: vllm+tilert+glm-5.1 error
#49874 opened
3 days agoJul 26, 2026 - [Bug]: Inkling parser is wrong for non reasoning mode - outputs "<|end_message|>"
#49865 opened
3 days agoJul 26, 2026 - [Bug] Modular MoE reports zero local experts after a backend releases source weights
#49863 opened
3 days agoJul 26, 2026 - [Bug]: VLLM_WSL2_ENABLE_PIN_MEMORY=1 is silently ignored on the ROCm platform
#49861 opened
3 days agoJul 26, 2026 - [Bug]: [MTP] ValueError when loading zai-org/GLM-OCR with MTP speculative decoding (model.layers.16.mtp_block uninitialized)
#49856 opened
3 days agoJul 26, 2026 - [Bug]: Multimodal models fail to load on ROCm/RDNA4 (gfx1201) — `CUDA error: invalid argument` in `vit_torch_sdpa_wrapper` encoder attention
#49851 opened
4 days agoJul 26, 2026 - [Bug]: PP=2 + GlmMoeDsa: inductor compile combined with CUDA-graph capture produces garbage output; either alone is clean (v0.24 & v0.26)
#49844 opened
4 days agoJul 26, 2026 - [Bug][KV Offload] `HIT_PENDING` has no deadline; a stalled write can defer requests until the client timeout
#49829 opened
4 days agoJul 26, 2026 - [Bug]: vllm serve --help=Frontend misses docs for inherited fields from BaseFrontendArgs
#49817 opened
4 days agoJul 26, 2026 - [Feature]: per gpu gpu-memory-utilization
#49816 opened
4 days agoJul 26, 2026 - [Bug]: PCP (#46570) broken for non-compress models (GLM-5.2, compress_ratio=1) — multiple crash paths
#49810 opened
4 days agoJul 26, 2026 - [Bug]: DeepGEMM 2.6.x UE8M0 assert - vLLM passes uninitialized FP32 scale-factor padding to the packing kernel
#49783 opened
4 days agoJul 25, 2026 - [Bug]: gpt-oss-120b fails for ROCm on gfx950 with Triton 3.7.1
#49778 opened
5 days agoJul 25, 2026 - [CI Failure]: PyTorch Compilation Unit Tests - test_dynamic_shapes_compilation
#49772 opened
5 days agoJul 25, 2026 - [RFC]: Native Disaggregated Pull-Based Queue Worker Interface and Heuristic Pull Router
#49765 opened
5 days agoJul 25, 2026 - [RFC]: vLLM Agentic Coding Readiness Survey
#49752 opened
5 days agoJul 25, 2026 - [Tracking Issue]: `sm_107` enablement for Rubin GPUs and the Vera Rubin platform
#49735 opened
5 days agoJul 25, 2026 - [Bug][Qwen 3.5 4B][H100]: Performance of DFlash is lower than expected
#49730 opened
5 days agoJul 24, 2026 - [Bug]: Responses custom tools are coerced to function_call on non-Harmony Qwen routes
#49724 opened
5 days agoJul 24, 2026 - [Performance]: Modelopt quantized model goes slower in fp8 than BF16 on B200 (sm100) using vLLM 0.25.1
#49723 opened
5 days agoJul 24, 2026 - [Bug]: int8_per_token_head KV cache corrupts Gemma-4 (hybrid attention) output under load on Triton
#49716 opened
5 days agoJul 24, 2026 - [Bug]: Responses silently ignores named tool_choice for unsupported XML tool parsers
#49712 opened
5 days agoJul 24, 2026 - [Bug]: poolside_v1 reports zero Responses reasoning_tokens for prompt-opened thinking spans
#49711 opened
5 days agoJul 24, 2026 - [Model Validation] SmolLM2-360M-Instruct batch invariance
#49708 opened
5 days agoJul 24, 2026 - [RFC]: Support router-driven mixtures of multiple LoRA adapters
#49705 opened
5 days agoJul 24, 2026 - [RFC]: Add an EPLB Platform Backend interface for out-of-tree accelerators
#49702 opened
5 days agoJul 24, 2026 - [Performance]: Compile mode 3 degrades triton w4a16 kernel performance in few request scenarios.
#49699 opened
5 days agoJul 24, 2026 - [Installation]: no cu129 nightly wheels for x86_64 platform
#49683 opened
5 days agoJul 24, 2026 - [Bug]: Deferred KV block frees cause zero-progress preemption cascades with async KV consumers
#49674 opened
last weekJul 24, 2026 - [Docs]: MkDocs TOC slugify differs from markdownlint MD051 anchor validation
#49662 opened
last weekJul 24, 2026 - [Feature]: Add a server-side tool strictness level for auto tool choice
#49661 opened
last weekJul 24, 2026 - [ROCm] Sparse-MLA persistent-kernel gate is unsafe for gqa_ratio=64 fp8 (no non-persistent kernel)
#49649 opened
last weekJul 24, 2026 - [RFC]: Enable full CUDA graphs for MoRIIO READ mode via async KV-load gating
#49643 opened
last weekJul 24, 2026 - [Bug]: [CPU] GDN attention falls back to slow torch conv1d on non-AMX AVX-512BF16 CPUs
#49640 opened
last weekJul 24, 2026 - NVFP4 fused-expert scales dropped in Qwen3_5_VL_MoE weight_loader (qwen3_5_vl_moe)
#49638 opened
last weekJul 24, 2026 - [Feature]: Integration of HieraSparse N:M Semi-Structured Sparse KV Cache Attention
#49630 opened
last weekJul 24, 2026 - [Bug]: Suspected RMSNorm precision regression after 225936a causes NGRAM/target token divergence for Qwen3-VL
#49616 opened
last weekJul 24, 2026 - [Bug]: DFlash 推理效率异常
#49614 opened
last weekJul 24, 2026 - [CI][torch nightly] mamba-ssm build fails under torch 2.14 nightly - forces -std=c++17 but ATen now requires C++20
#49597 opened
last weekJul 23, 2026 - [Bug]: deadlock / hang with GLM-5.2-FP8 EP+DP on GB200 - with ablation table
#49594 opened
last weekJul 23, 2026 - [Bug]: if passed temperature=0.0 there is no output of its value in telemetry span
#49589 opened
last weekJul 23, 2026 - [Bug]: vllm0.25版本支持qwen3-vl fp8算子DeepGemmFp8BlockScaledMMKernel 走原始的deepgemm的后端
#49584 opened
last weekJul 23, 2026 - [Feature]: Plan for enabling mypy in tests directory
#49569 opened
last weekJul 23, 2026 - [Feature] Support `extract_hidden_states` on Model Runner V2
#49562 opened
last weekJul 23, 2026 - [Bug]: QWEN 3.6-35Ba3B+DFlash (Train by Speculator) when vllm0.25.1 use V2, the acceptance rate is 0%.
#49559 opened
last weekJul 23, 2026 - [Bug]: In-tree fallback LMCache adapter does not propagate retrieve failures to the engine
#49556 opened
last weekJul 23, 2026 - [Bug]: --moe-backend humming crashes at startup on NVIDIA GB10 / DGX Spark (humming NVML mem-clock query unsupported)
#49554 opened
last weekJul 23, 2026 - [Bug]: MTP draft model ignores checkpoint per-layer quantization config (block_name_to_quantize mapped inconsistently vs target)
#49552 opened
last weekJul 23, 2026 - FlashInfer + spec-decode silently downgrades to PIECEWISE cudagraphs (−16% measured); warning understates the cost
#49547 opened
last weekJul 23, 2026 - [RFC]: Speculative recompute fallback for slow remote KV cache loads
#49545 opened
last weekJul 23, 2026 - [Feature]: Real-time request-queue signal (NUM_REQUESTS_RUNNING/WAITING) for external routers
#49538 opened
last weekJul 23, 2026 - [Feature] Support KV-cache events & KV-transfer for hybrid-attention models
#49537 opened
last weekJul 23, 2026 - [Feature]: Support for Responses API `/v1/responses/input_tokens`
#49533 opened
last weekJul 23, 2026 - [Feature]: Make PG_WAIT_TIMEOUT configurable via environment variable
#49530 opened
last weekJul 23, 2026 - [Perf][Kernel] Adopt PTX 9.4 `ldmatrix.s8.s4` (hardware INT4→INT8 expanding load) in W4A8-INT8 paths
#49529 opened
last weekJul 23, 2026 - [Bug]: NIXL poll leaves transfers inflight when check_xfer_state raises
#49528 opened
last weekJul 23, 2026 - [Feature]: Make Ray placement group strategy configurable via VLLM_RAY_PG_STRATEGY
#49527 opened
last weekJul 23, 2026 - [Bug]: Harmony models reject namespace tool type in Responses API
#49493 opened
last weekJul 23, 2026 - [Bug]: vllm bench serve - huggingface datasets - "TypeError: string indices must be integers, not 'str'"
#49480 opened
last weekJul 23, 2026 - [Bug]: Failed to run RedHatAI/gemma-4-31B-it-speculator.dspark
#49475 opened
last weekJul 23, 2026
1113 Unresolved conversations
Sometimes conversations happen on old items that aren't yet closed. Here is a list of all the Issues and Pull Requests with unresolved conversations.
- [MoE Refactor] Rename FusedMoE to FusedMoEFactory
#44941 commented on
8 hours agoJul 29, 2026 • new comments - Add tiering offloading metrics
#48798 commented on
2 hours agoJul 30, 2026 • new comments - [Do Not Merge][2/N][Attention] Enable masked MHA for sparse MLA prefills
#48770 commented on
46 minutes agoJul 30, 2026 • new comments - [Mypy Fix] Mypy fix for "vllm/model_executor/models/[aA][bB]"
#48977 commented on
yesterdayJul 29, 2026 • new comments - [XPU] Support MXFP8 gemm/bmm weights for INC DeepSeek V4 model
#48476 commented on
yesterdayJul 28, 2026 • new comments - [Perf][Hybrid] 3D-grid tiling of the state-copy Triton kernels
#49436 commented on
2 hours agoJul 30, 2026 • new comments - add torchcomms backend for xpu support
#46210 commented on
2 days agoJul 28, 2026 • new comments - [KV Offload] Fix failed-load livelock by marking the lookup verdict negative
#49328 commented on
yesterdayJul 28, 2026 • new comments - [Online quantization] Add online MXFP4 quantization support
#49347 commented on
yesterdayJul 28, 2026 • new comments - Hardware-agnostic model definition via HF transformer backend (1/N)
#49458 commented on
2 days agoJul 27, 2026 • new comments - Add pushed-based ZMQ metrics logger support
#48969 commented on
2 days agoJul 28, 2026 • new comments - Port DeepGemm MHC Kernel to CuTeDSL
#48619 commented on
6 hours agoJul 29, 2026 • new comments - [Core][KV Connector][Mooncake] Support MRV2 prefill PCP in Mooncake
#49109 commented on
yesterdayJul 28, 2026 • new comments - [PARSER][Mistral] unified engine-based parser for reasoning and tool calls
#48947 commented on
1 hour agoJul 30, 2026 • new comments - [MRV2][Spec] Fuse AR speculator multi-step decodes back into one CUDA graph
#46849 commented on
2 days agoJul 28, 2026 • new comments - [Compilation]Fuse Transformers Residual Add + RMSNorm
#48757 commented on
8 hours agoJul 29, 2026 • new comments - [DSv4 Perf] Optimize workspace reuse for eager break, 3.9% E2E TTFT improvement.
#49236 commented on
3 hours agoJul 29, 2026 • new comments - [KV Connector] Add per-layer canonical KV page mappings for parallelism-agnostic offload
#48408 commented on
33 minutes agoJul 30, 2026 • new comments - [ROCm][Quark][7/N] Use MXFP4 linear kernel abstraction for `emulation` backend
#48949 commented on
1 hour agoJul 30, 2026 • new comments - [AutoRound] Support AutoRound Format Block-Wise FP8 in vLLM
#47434 commented on
yesterdayJul 28, 2026 • new comments - [Perf] Add mm_tensor_ipc=cuda_ipc TP-aware GPU transport
#49462 commented on
yesterdayJul 28, 2026 • new comments - [ROCm] [CI] Support cached K/V (key/value=None) in Triton prefix-prefill
#48257 commented on
1 hour agoJul 30, 2026 • new comments - [WIP] Switch vLLM to The Rock when on AMD hardware
#49470 commented on
2 hours agoJul 30, 2026 • new comments - [Security] Fix DoS via out-of-range stop_token_ids crashing CUDA
#44968 commented on
2 days agoJul 27, 2026 • new comments - [Attention] Batch-level prefill/decode attention backend routing
#48162 commented on
yesterdayJul 29, 2026 • new comments - feat(nixl,dcp): Supports MLA DCP for PD disaggregation with nixl connector
#38433 commented on
9 hours agoJul 29, 2026 • new comments - refactor(envs): migrate vllm/envs.py to pydantic-settings
#42136 commented on
20 hours agoJul 29, 2026 • new comments - Add NVFP4 KV 4-over-6 scale search
#45187 commented on
1 hour agoJul 30, 2026 • new comments - [Model] Honor fp32 head_dtype in Inkling muP logits (target + MTP draft)
#49120 commented on
11 hours agoJul 29, 2026 • new comments - [ROCm] Enable AITER and FP8 inference on GFX120x
#43615 commented on
1 hour agoJul 30, 2026 • new comments - [Quantization] Support native Quark W4A16 INT4/UINT4 exports in vLLM
#48606 commented on
last weekJul 24, 2026 • new comments - [Hardware][XPU] Register matmul and linear batch-invariant kernels for XPU
#49209 commented on
2 days agoJul 27, 2026 • new comments - [ROCm] DeepSeek-V4-Pro PD Disaggregation through MORI IO KV Connector on AMD GPUs
#48989 commented on
yesterdayJul 29, 2026 • new comments - [Core] Consume SleepModeBackend capability flags in the worker suspend/resume path
#47500 commented on
2 days agoJul 27, 2026 • new comments - [Model] Enable LoRA support for tower and connector in Voxtral
#45697 commented on
yesterdayJul 29, 2026 • new comments - [Misc] Add unit test for decode_attention Triton kernels
#47862 commented on
2 days agoJul 27, 2026 • new comments - [ROCm][feature] Add new moe backends supporting int4/int8 weight-only…
#49086 commented on
2 days agoJul 27, 2026 • new comments - [2/N][Feat][Perf] Add new warmup infrastructure for JITs. Add predicate filtering for JIT warmup, and migrate Inkling FA4
#49315 commented on
1 hour agoJul 30, 2026 • new comments - [Kernel] ReplaySSM: cache SSM inputs for faster Gated DeltaNet standard decode
#48792 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Fix structured outputs FSM desync with pipeline parallelism on the V1 model runner
#45015 commented on
yesterdayJul 29, 2026 • new comments - [Bugfix] Support non-gated MoE in online quantization and Marlin MoE tile padding
#48028 commented on
18 hours agoJul 29, 2026 • new comments - [feature] shuffle safetensor weight files
#42837 commented on
last weekJul 23, 2026 • new comments - [Frontend][Core][Spec Decode] Per-request acceptance stats in OpenAI API responses
#48915 commented on
yesterdayJul 29, 2026 • new comments - [XPU][UT]Convert awq-packed MoE qweight to the gptq-equivalent layout on xpu
#48552 commented on
5 days agoJul 24, 2026 • new comments - Enable gfx1250 ROCm architecture
#46516 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix][Spec Decode] Fix DSpark bonus-anchor query width accounting
#48909 commented on
5 days agoJul 25, 2026 • new comments - [ROCm][CI] Loosen wvSplitKrc skinny-GEMM test tolerance for bf16 (gfx950)
#49309 commented on
yesterdayJul 28, 2026 • new comments - [CPU][Zen] Route BF16 MoE inference through zentorch on AMD
#44201 commented on
1 hour agoJul 30, 2026 • new comments - perf: Cache staging buffer in structured output to fix memory regression (#49013)
#49168 commented on
3 hours agoJul 29, 2026 • new comments - [Frontend] Add the p-less Sampling Method for LLM Decoding
#45288 commented on
7 hours agoJul 29, 2026 • new comments - [Bugfix][ROCm] AITER MLA: size MTP verification decode metadata for real qlen/dtype
#45227 commented on
2 hours agoJul 30, 2026 • new comments - [Perf] Add per-phase cold-start startup span benchmark harness
#49037 commented on
2 days agoJul 28, 2026 • new comments - perf: pre-allocate buffer to avoid per-call zeroing of MLA KV counter buffer
#49447 commented on
11 hours agoJul 29, 2026 • new comments - [Frontend] Add vllm snapshot create and CRIU restore for serve, default-on in the Docker image
#48996 commented on
20 hours agoJul 29, 2026 • new comments - feat(frontend): session id plumbing into requests
#48048 commented on
17 hours agoJul 29, 2026 • new comments - [Spec Decode] Add PARD-2 parallel draft model support
#49406 commented on
6 hours agoJul 29, 2026 • new comments - [KERNEL][ROCm]Native HIP MXFP4(Compressed+Quark) (dense + MoE) for RDNA3
#46676 commented on
yesterdayJul 28, 2026 • new comments - [Do not merge!] [Build] Migrate vendored DeepGEMM from pybind to TORCH_LIBRARY (abi3)
#48962 commented on
yesterdayJul 29, 2026 • new comments - [CPU][Zen] Route Int8 MoE inference through zentorch on AMD
#44834 commented on
1 hour agoJul 30, 2026 • new comments - [Bugfix] Emit a valid media type from encode_{audio,image,video}_url
#49056 commented on
2 days agoJul 28, 2026 • new comments - [Perf] Integrate flash-maxsim Triton kernels for late-interaction scoring
#40337 commented on
3 days agoJul 26, 2026 • new comments - Detect ROCm wheel variant from environment for precompiled wheels.
#49365 commented on
2 days agoJul 28, 2026 • new comments - [Frontend] Cohere chat v2 api support
#47189 commented on
2 hours agoJul 30, 2026 • new comments - Enable return_routed_experts support with CPU KV offload
#45635 commented on
12 hours agoJul 29, 2026 • new comments - [DCP] Add FlashInfer fused A2A backend for decode context parallelism
#48248 commented on
yesterdayJul 28, 2026 • new comments - [Misc] Add unit test for _fwd_kernel_ep_scatter_1 and _fwd_kernel_ep_…
#46064 commented on
yesterdayJul 28, 2026 • new comments - [ROCm][CI] Upgrading ROCm AITER MLA coverage
#40948 commented on
33 minutes agoJul 30, 2026 • new comments - [Feature] Default EPLB num_redundant_experts to minimum valid value
#30614 commented on
8 hours agoJul 29, 2026 • new comments - [bugfix] fix minicpmv mm prompt placeholder parse error
#49162 commented on
8 hours agoJul 29, 2026 • new comments - [Perf] [Feat] [ROCm] Add densemha support to ROCm AITER Sparse MLA
#49263 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Make metadata send non-blocking in GroupCoordinator.isend_tensor_dict
#49274 commented on
15 hours agoJul 29, 2026 • new comments - [Perf][ROCm] Dual-stream decode with hipgraphs
#48223 commented on
4 hours agoJul 29, 2026 • new comments - [Core][hybrid attention] group-aware recovery for KV load failures on hybrid models
#48216 commented on
10 hours agoJul 29, 2026 • new comments - [Model][LoRA] Add tower/connector LoRA support for Ultravox
#48215 commented on
yesterdayJul 28, 2026 • new comments - [Bugfix][Model] MiMo-V2: support TP > num_kv_heads for the fused FP8 QKV projection
#46755 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix][Parser] kimi_k2: route streaming through arg_converter so schema type coercion applies
#49318 commented on
yesterdayJul 28, 2026 • new comments - [Feature][Whisper] Native word-level timestamps (cross-attention + DTW)
#47664 commented on
2 days agoJul 28, 2026 • new comments - [rl] Stateful Trainer Send: IPC [2/N]
#48981 commented on
2 days agoJul 28, 2026 • new comments - [ROCM][Bugfix] Deferring to TRITON_ATTN backend for Prefix lm models instead of ROCM_ATTN and fixed minor bug in Gemma3 Multimodal code
#46754 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Accept logprobs=-1 in the Completion API
#46175 commented on
2 days agoJul 28, 2026 • new comments - [WIP][Spec Decode] DSpark confidence-scheduled verification
#47808 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix][Distributed] Disable fused allreduce when NVLink multicast is unavailable
#48075 commented on
3 days agoJul 27, 2026 • new comments - [XPU][Attention] Widen prefill BLOCK_M to 32 for power-of-2 GQA (~2x prefill on BMG)
#47246 commented on
4 days agoJul 25, 2026 • new comments - [Bugifx][INC] Fix INC quantization method selection for non-quantized layers
#47237 commented on
last weekJul 23, 2026 • new comments - [ROCm]Migrating Deepseek V3.2 to vllm/models/deepseek_v32/
#47207 commented on
1 hour agoJul 30, 2026 • new comments - [Spec Decode] Add D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding
#47131 commented on
last weekJul 23, 2026 • new comments - [Kernel][XPU] Tensor-descriptor operand loads for Triton W8A8 scaled_mm
#47205 commented on
last weekJul 23, 2026 • new comments - [Rust Frontend] Match Python union-type coercion precedence
#47203 commented on
5 hours agoJul 29, 2026 • new comments - [CI] Add GSM8K accuracy configs for large NVFP4/INT4 MoEs (4xB200)
#47171 commented on
2 days agoJul 27, 2026 • new comments - [Build] Fix CUDA arch coverage checks and scoped kernel feature flags
#47149 commented on
yesterdayJul 28, 2026 • new comments - [Attention] Add KVarN: calibration-free variance-normalized KV-cache quantization backend
#46812 commented on
last weekJul 23, 2026 • new comments - [Rust Frontend] Add Rust Apertus tool parser
#46813 commented on
15 hours agoJul 29, 2026 • new comments - [BugFix] Fix decode-phase blocks not stored in CPU KV cache offloading
#46824 commented on
4 days agoJul 25, 2026 • new comments - [codex] Fix oversized video inputs in multimodal loader
#46836 commented on
5 days agoJul 25, 2026 • new comments - [Bugfix][Core] Fix request-bound KV cache sizing
#46841 commented on
yesterdayJul 28, 2026 • new comments - [CI] Mooncake PD integration tests
#46844 commented on
8 hours agoJul 29, 2026 • new comments - [Bugfix] Fix MiniMax-M3 compressed-tensors FP8 MoE SwiGLU params
#46845 commented on
13 hours agoJul 29, 2026 • new comments - [Kernel] Manual activation+quant fusion via QuantizedActivation
#46864 commented on
last weekJul 23, 2026 • new comments - [XPU] Low-latency all-to-all Expert-Parallellism backend for batched MoE
#46871 commented on
2 hours agoJul 30, 2026 • new comments - [Core] Use FlashInfer workspace sizing helper
#46883 commented on
10 hours agoJul 29, 2026 • new comments - [MoE] [MoE Refactor] Migrate int8 w4a8int8 oracle 37753
#46901 commented on
yesterdayJul 28, 2026 • new comments - [CPU][Bugfix] Build cpu_fused_moe on Apple Silicon
#46907 commented on
2 days agoJul 27, 2026 • new comments - [Hybird][PrefixCache] Pre-copy-free align prefix cache
#46912 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix][GB10] Fix negative CUDA graph memory estimate on unified-memory GPUs (#44740)
#46932 commented on
last weekJul 23, 2026 • new comments - Update flashinfer compute support to 7.5
#46943 commented on
last weekJul 24, 2026 • new comments - fix: clear shared_experts output when forward fails
#46950 commented on
1 hour agoJul 30, 2026 • new comments - [XPU] Unify XPU RMSNorm kernels with vllm_c and drop redundant XPU-specific implementation
#46981 commented on
4 hours agoJul 29, 2026 • new comments - [ROCm] Pass vLLM gfx942 FP8 FNUZ bound to DeepEP builds
#46989 commented on
last weekJul 24, 2026 • new comments - [Spec][V2] Support MTP speculative decoding under pipeline parallelism
#46994 commented on
yesterdayJul 29, 2026 • new comments - [Bugfix][MiniMax-M3] Fix garbled outputs and streaming reasoning parser on NVFP4 MoE backends for SM 12.0a (RTX PRO 6000 Blackwell)
#47001 commented on
7 hours agoJul 29, 2026 • new comments - [Bugfix] Handle list slot_mapping in remaining KV-cache consumers
#47013 commented on
2 days agoJul 27, 2026 • new comments - [ROCm][DistInf] Enable vLLM DI CI with buildkite/slurm
#47030 commented on
5 hours agoJul 29, 2026 • new comments - [BugFix] Improve DCP error message with actionable remedy
#47041 commented on
4 days agoJul 25, 2026 • new comments - Revert "[MoE Refactor] Standardize Humming MoE experts + utilities" (#43373)
#47097 commented on
last weekJul 23, 2026 • new comments - [XPU] fix collecting oneccl version info
#47104 commented on
5 days agoJul 24, 2026 • new comments - [Kernel] Support Nvfp4 Cutedsl Moe Swiglu-oai and Relu2(non-gated) Activation
#47106 commented on
5 hours agoJul 29, 2026 • new comments - Bump astral-sh/setup-uv from 7.6.0 to 9.0.0
#47108 commented on
16 hours agoJul 29, 2026 • new comments - [Bugfix] Fix misclassification of prefill as uniform decode for hybrid models
#47123 commented on
3 days agoJul 27, 2026 • new comments - [Quantization][Autoround][XPU] Add W4A16(moe) / MXFP4(linear/moe) Support
#47124 commented on
2 hours agoJul 29, 2026 • new comments - [XPU][Attention] Enable tensor-descriptor Q load for non-power-of-2 GQA (~1.9x Qwen2-7B)
#47248 commented on
4 days agoJul 25, 2026 • new comments - [MM][CG] Support ViT full CUDA graph for Idefics3 and SmolVLM
#47625 commented on
4 days agoJul 26, 2026 • new comments - [Attention] TRITON_MLA_SPARSE backend for SM80/SM121 sparse MLA (rebase & takeover of #38476)
#47629 commented on
4 days agoJul 26, 2026 • new comments - [Bugfix][LoRA] Guard None group members in expand_packed_lora (partial LoRA on Qwen3.5/3.6 GatedDeltaNet)
#47640 commented on
3 days agoJul 26, 2026 • new comments - Fix gemma4 unified multi audio
#47652 commented on
26 minutes agoJul 30, 2026 • new comments - [Quantization] Marlin FP4: report missing CUDA custom ops in is_supported
#47673 commented on
last weekJul 23, 2026 • new comments - [Quantization] NVFP4 W4A16: fall back to weight-only emulation when Marlin is unavailable
#47674 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix `--data-parallel-start-rank 0` being treated as unset in `create_engine_config`
#47692 commented on
5 days agoJul 25, 2026 • new comments - [Rust Frontend] Beam search support for completions and chat completions
#47705 commented on
15 hours agoJul 29, 2026 • new comments - Add FP8 W8A8 Block kernel configs for NVIDIA GeForce RTX 4090D
#47732 commented on
last weekJul 23, 2026 • new comments - fix: bad invalid expert IDs and scaling factor in FlashInfer all-to-all communication integration
#47733 commented on
3 days agoJul 27, 2026 • new comments - [ROCm][CI][Bugfix] Update mrope triton kernel to support disabling neox style
#47740 commented on
2 hours agoJul 30, 2026 • new comments - [Draft] Add NexusQuant E8 lattice KV cache quantization backend
#47742 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Disable custom all-reduce on consumer Blackwell (sm_12x)
#47751 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix] Prevent encoder cache deadlock with multiple images in Mamba align mode
#47771 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Add DSV4-Flash-DSpark gsm8k and pipe kv_quant_mode through SWA MLA unify
#47776 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Support quantized eh_proj in HYV3 MTP draft model
#47792 commented on
last weekJul 23, 2026 • new comments - Handle missing vLLM metadata in Triton import
#47793 commented on
4 days agoJul 26, 2026 • new comments - [ROCm][DSV4][Perf] Step 1/N: Port ATOM attention into vLLM: port single-node single-token-prediction case
#47796 commented on
last weekJul 23, 2026 • new comments - [do not land] [wip] Add Helion call-site routing + route per_token_group_fp8_quant
#47799 commented on
last weekJul 23, 2026 • new comments - [Core][Hardware][NVIDIA] Add custom all-reduce suspend hooks
#47806 commented on
2 hours agoJul 30, 2026 • new comments - refactor: streamline DeepSeek V4 mHC warmup and remove token-size cap
#47807 commented on
4 hours agoJul 29, 2026 • new comments - Bump astral-sh/setup-uv from 7.6.0 to 9.0.0
#47814 commented on
16 hours agoJul 29, 2026 • new comments - [Bugfix][Scheduler] Skip lookahead allocation for running prefill chunks
#47822 commented on
last weekJul 24, 2026 • new comments - [Models] Fix Qwen3.5 MTP with GPTQ/AWQ quantization
#47828 commented on
last weekJul 24, 2026 • new comments - [Core][KV Cache] Priority-aware KV cache eviction for priority scheduling
#47837 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Reject whitespace-only bad_words in SamplingParams
#47841 commented on
6 hours agoJul 29, 2026 • new comments - [Frontend] Add video decode cache
#47856 commented on
2 days agoJul 28, 2026 • new comments - Fix video temporal padding token estimates
#47876 commented on
2 hours agoJul 30, 2026 • new comments - [Doc] Fix broken link, stale permalinks, and missing docstrings
#47887 commented on
15 hours agoJul 29, 2026 • new comments - [Kernel][ROCm][Perf] FlyDSL decode-attention kernel for 4-bit TurboQuant KV cache
#47896 commented on
last weekJul 23, 2026 • new comments - [ROCm]: Bump torch 2.12, torchvision, torchaudio, triton 3.7
#47917 commented on
yesterdayJul 29, 2026 • new comments - Fix Qwen3.5 text causal LM hybrid support
#47252 commented on
2 hours agoJul 30, 2026 • new comments - [Rust Frontend] add functiongemma tool parser support
#47254 commented on
12 hours agoJul 29, 2026 • new comments - [Bugfix][Model] Harden DiffusionGemma self-conditioning matmul against torch.compile shape mismatch (#47129)
#47278 commented on
2 days agoJul 28, 2026 • new comments - [Frontend] Add detokenization streaming derender for disaggregated serving
#47301 commented on
1 hour agoJul 30, 2026 • new comments - [Bugfix][Kernel] Re-bias exponent in BF16 NVFP4 Marlin scale dequant
#47315 commented on
3 days agoJul 27, 2026 • new comments - [ROCm][Perf] Enable fused indexer-Q RoPE+quant kernel for DeepSeek/GLM sparse attention
#47335 commented on
4 days agoJul 26, 2026 • new comments - [Rust Frontend]: Added olmo3 reasoning parser
#47345 commented on
2 hours agoJul 30, 2026 • new comments - [Attention] Overlap sparse MLA indexer with native CUDA streams
#47355 commented on
yesterdayJul 29, 2026 • new comments - [Rust Frontend] Add TLS certificate hot-reload (`--enable-ssl-refresh`)
#47362 commented on
13 hours agoJul 29, 2026 • new comments - fix(logging): redact structured-outputs schema in crash dumps
#47368 commented on
5 days agoJul 25, 2026 • new comments - Revert "[CPU][Perf]Added tanh AOR for faster gelu activations." (#44639)
#47369 commented on
2 days agoJul 27, 2026 • new comments - fix(metrics): fix TPOT accounting for MTP multi-token outputs
#47382 commented on
2 days agoJul 27, 2026 • new comments - Revert "[MoE] Plumb gemm1_alpha/beta/clamp_limit into TRT-LLM FP8 MoE" (#45723)
#47431 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Structured Output] Let terminal grammars stop under min_tokens
#47489 commented on
last weekJul 24, 2026 • new comments - [DSV4] Support MTP for TP16
#47492 commented on
last weekJul 23, 2026 • new comments - [Rust Frontend] add gigachat3 tool parser
#47501 commented on
15 hours agoJul 29, 2026 • new comments - [KVConnector] Guard lmcache_mp_connector state transition with num_external_tokens
#47505 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Gemma-4 k_eq_v x compressed-tensors: propagate shard aliases
#47507 commented on
last weekJul 23, 2026 • new comments - [Bugfix] 'Already borrowed' by pooling tokenizer in StructuredOutputManager
#47509 commented on
2 days agoJul 28, 2026 • new comments - DFlash SWA — resolved for personal build
#47511 commented on
last weekJul 23, 2026 • new comments - [Quantization] Fix NVFP4 per-half global scale for fused gate_up_proj
#47515 commented on
last weekJul 23, 2026 • new comments - [Attention][TurboQuant] Preserve configured FlashAttention version
#47531 commented on
last weekJul 23, 2026 • new comments - fix(quant): forward MiniMax M3 SwiGLU parameters to FusedMoEQuantConfig in int8 paths
#47552 commented on
last weekJul 23, 2026 • new comments - [Core] Optimize sliding-window prefix-cache miss scan
#47556 commented on
yesterdayJul 28, 2026 • new comments - [Frontend][Model] Add language control and metadata for Qwen3-ASR realtime ASR
#47563 commented on
2 days agoJul 27, 2026 • new comments - [Kernel] ReplaySSM: cache SSM inputs instead of state for faster standard and speculative decode (Mamba2 + GDN)
#47576 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fall back to V1 model runner when UVA is unavailable (WSL2)
#47579 commented on
3 days agoJul 26, 2026 • new comments - [Perf][Frontend] Add opt-in incremental prompt-encoding cache for multi-turn chat
#47583 commented on
3 days agoJul 27, 2026 • new comments - [ModelRunner V2] Support custom logits processors
#47585 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix] [EPLB] Avoid Silent Accuracy Corruption in Quark MXFP4 EPLB
#47588 commented on
last weekJul 23, 2026 • new comments - Fix: detect reasoning end with accepted MTP tokens
#47617 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Reject NVFP4 MoE checkpoints with missing per-expert scales
#45320 commented on
last weekJul 23, 2026 • new comments - Add Kimi-K2.5 NVFP4 CuTe DSL decode kernels
#45642 commented on
2 days agoJul 28, 2026 • new comments - [Core] shm_broadcast: honor shutdown in acquire_write to avoid dead-worker writer hang
#45644 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Deterministic MoE combine (reduce_scatterv) under VLLM_BATCH_INVARIANT
#45683 commented on
5 days agoJul 24, 2026 • new comments - [KVConnector] Remove `VLLM_DISABLE_REQUEST_ID_RANDOMIZATION` and auto-set it for PD
#45688 commented on
5 hours agoJul 29, 2026 • new comments - [Misc] Add and enable Triton kernel unit tests on XPU
#45694 commented on
last weekJul 24, 2026 • new comments - [ROCm][DSv4] Add AITER MXFP4-MXFP8 W4A8 MoE backend
#45711 commented on
last weekJul 23, 2026 • new comments - [ROCm] Convert MXFP8 MoE weights to block FP8 on gfx94x
#45726 commented on
last weekJul 23, 2026 • new comments - [Quantization] Extend ModelOpt mixed precision and NVFP4 runtime formats
#45735 commented on
last weekJul 23, 2026 • new comments - [NVFP4] Support clamped SwiGLU-OAI (SwigluBias) on FlashInfer-CUTLASS MoE
#45738 commented on
last weekJul 23, 2026 • new comments - Turboquant native fp8 v4 store
#45748 commented on
last weekJul 23, 2026 • new comments - [BugFix] Add kimi_k2 to MLA allowlist for EAGLE3 draft model detection
#45764 commented on
2 days agoJul 28, 2026 • new comments - [fuse_norm_quant] Support mixed input/weight dtype in fused RMSNorm+FP8 quant kernels (fixes Qwen3.5-FP8)
#45770 commented on
last weekJul 23, 2026 • new comments - [DiffusionGemma-26B-A4B-it FP8][XPU][Quant] Enable native XPU FP8 MoE for per-token dynamic activations
#45776 commented on
last weekJul 23, 2026 • new comments - [DRAFT] OPT: MergeAttention shared memory for LSE loading
#45778 commented on
5 hours agoJul 29, 2026 • new comments - [DiffusionGemma-26B-A4B-it-NVFP4 ][XPU][Quant] Register NVFP4 emulation linear kernel for XPU
#45779 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Parser engine: count reasoning tokens when <think> is pre-filled in the prompt
#45787 commented on
yesterdayJul 28, 2026 • new comments - [ROCm][MLA] Fuse RoPE + MLA KV-cache write via AITER Triton kernel
#45798 commented on
last weekJul 24, 2026 • new comments - [PD][Core] Add partial prefix caching for hybrid model in PD disaggregation
#45804 commented on
14 hours agoJul 29, 2026 • new comments - fix: report stop_sequence stop_reason in Anthropic Messages API
#45807 commented on
2 days agoJul 28, 2026 • new comments - [Refactor] Make Int8ScaledMMLinearLayerConfig use QuantKey
#45821 commented on
last weekJul 23, 2026 • new comments - TRITON_ATTN support for KV cache dtype fp8 on sm75 to pre-sm89
#45829 commented on
4 days agoJul 26, 2026 • new comments - Minimax m3 gfx950 mxfp4
#45838 commented on
last weekJul 23, 2026 • new comments - [Kernel] Manual TP fusion via ResidualStream (AR+RMSNorm+quant)
#45855 commented on
last weekJul 23, 2026 • new comments - [ROCm] Minimax m3 mxfp4 on 45794
#45894 commented on
last weekJul 23, 2026 • new comments - fix(kimi-k2): defer reasoning→tool transition when boundary token tex…
#45909 commented on
2 days agoJul 28, 2026 • new comments - [Compiler] Normalize reshape shapes before RMS quant fusion
#45910 commented on
last weekJul 23, 2026 • new comments - [Perf][Kernel][ROCm] Add Triton split-KV paged decode fallback for gfx12
#45916 commented on
1 hour agoJul 30, 2026 • new comments - Pass quant config to Qwen3Next embeddings
#45994 commented on
last weekJul 23, 2026 • new comments - [KV Connector][Mooncake]Support DeepSeek V4 Mooncake PP/PD shared KV metadata
#46004 commented on
last weekJul 23, 2026 • new comments - [Bugfix][MoE] Preserve unquantized weight storage on ROCm
#46009 commented on
last weekJul 23, 2026 • new comments - [Core] Drop stale scheduler requests before scheduling
#46011 commented on
5 days agoJul 25, 2026 • new comments - Reduce prompt logprobs memory usage
#45327 commented on
2 days agoJul 27, 2026 • new comments - [Kernel] [vLLM IR] [Perf] Add support for mixed input/weight dtype to rms norm quant fusion ops
#45331 commented on
last weekJul 23, 2026 • new comments - [Model][Spec Decode] MiMo-V2.5-Pro-FP4 DFlash speculative decoding + AL fix
#45343 commented on
last weekJul 23, 2026 • new comments - [Quant] Port RMSNormQuantFusionPass to manual fusion (fp8 static per-tensor + dynamic per-token)
#45364 commented on
last weekJul 23, 2026 • new comments - [Feature] Fused K-RoPE + static FP8 per-tensor KV Cache write
#45370 commented on
5 hours agoJul 29, 2026 • new comments - [Kernel][XPU] Enable per-token-group quant kernel tests on XPU
#45382 commented on
last weekJul 23, 2026 • new comments - docs: clarify Gemma 4 MTP assistant quantization
#45420 commented on
last weekJul 23, 2026 • new comments - docs: add TurboQuant KV cache presets
#45422 commented on
last weekJul 23, 2026 • new comments - [Kernel] Single-read fast path for fused RMSNorm + dynamic per-token FP8 quant
#45428 commented on
20 hours agoJul 29, 2026 • new comments - fix: reset KV load recompute placeholders
#45430 commented on
2 days agoJul 27, 2026 • new comments - [Draft][Rust Frontend] Route top-N tool-call families through the Dynamo parser
#45451 commented on
2 days agoJul 27, 2026 • new comments - Opt-in two-stream + launch-elision latency optimizations for Kimi-K2.6-NVFP4 decode on B300 (TP=4)
#45452 commented on
2 days agoJul 27, 2026 • new comments - [Perf] Reuse topk SparseMatrix routing metadata in GPT-OSS MoE forward
#45457 commented on
19 hours agoJul 29, 2026 • new comments - Hardware-agnostic model definition for DeepSeek V4
#45470 commented on
4 days agoJul 25, 2026 • new comments - [ROCm][CI] LM Eval coverage for Kimi K2 Thinking and MiniMax M2.7
#45481 commented on
3 days agoJul 27, 2026 • new comments - Preserve tensor dtypes when loading quantized tensorizer models
#45504 commented on
last weekJul 23, 2026 • new comments - fix: cache bad_words tokenization to avoid 'Already borrowed' errors under concurrency
#45522 commented on
5 hours agoJul 29, 2026 • new comments - [Model][Quant] compressed-tensors WNA16 input embeddings + tied embedding (lm_head) support
#45535 commented on
2 days agoJul 28, 2026 • new comments - Fix cp_gather_cache block table bounds
#45537 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix SamplingParams repr, docstrings, top_k validation order, and convert StructuredOutputsParams to msgspec.Struct
#45541 commented on
5 hours agoJul 29, 2026 • new comments - Add Kimi video chunk splitting
#45545 commented on
2 days agoJul 28, 2026 • new comments - [ROCm][Kernel] Extend skinny gemm N=5 to N=8 cases on GFX12 (RDNA4) using SWMMAC optimization
#45559 commented on
2 hours agoJul 30, 2026 • new comments - [GPT-OSS] Strict tool call and constrained decoding for Harmony
#45560 commented on
5 hours agoJul 29, 2026 • new comments - [Bugfix] cumem: validate VA/handle invariants and recover from wake-time cuMemMap failures
#45565 commented on
3 days agoJul 27, 2026 • new comments - [Attention] Porting MLARoPEKVCacheCatFusionPass to manual fusion
#45573 commented on
2 days agoJul 28, 2026 • new comments - fix disable_sliding_window for strict hf gemma4 config
#45578 commented on
4 days agoJul 25, 2026 • new comments - [BugFix] Fix MXFP8 checkpoint loading when quant_method collides with online shorthand
#45582 commented on
last weekJul 23, 2026 • new comments - [Bugfix][V1] Warm up all top-k/top-p Triton sampler kernel specializations
#45601 commented on
last weekJul 23, 2026 • new comments - [BugFix] Auto-disable custom all-reduce when cumem allocator (sleep mode) is active
#45611 commented on
11 hours agoJul 29, 2026 • new comments - refactor: rename flashinfer_utils.py to flashinfer_moe.py
#45618 commented on
last weekJul 23, 2026 • new comments - [BugFix] Gate level-2 sleep buffer restore on the weights wake tag
#45619 commented on
yesterdayJul 28, 2026 • new comments - [Bugfix][Kernel][ROCm] Fix Wave32 LDS overflow in top-k merge launch
#46012 commented on
2 hours agoJul 30, 2026 • new comments - [LocateAnything-3B][XPU] Add locateanything model on support on XPU
#46395 commented on
2 days agoJul 28, 2026 • new comments - [Quantization] Align online fp8_ptpc/block_fp8/mxfp8 weight quantization with llm-compressor (compressed-tensors) export
#46400 commented on
last weekJul 23, 2026 • new comments - [rust][tool-parser] Add streaming invariant tests
#46416 commented on
2 days agoJul 27, 2026 • new comments - [Bugfix][Mamba2] Fix assert crash when prefill-reclassified-as-decode occurs with no concurrent spec tokens
#46424 commented on
19 hours agoJul 29, 2026 • new comments - [rust][server] Add generative scoring route
#46427 commented on
3 days agoJul 27, 2026 • new comments - [Model][Attention] DiffusionGemma: NVFP4 KV cache via FlashInfer VO-split + per-request causal grouping (sm120)
#46443 commented on
last weekJul 23, 2026 • new comments - [Bugfix][TurboQuant] Add continuation guard to fast-path to prevent prefix K/V loss under prefix caching
#46461 commented on
last weekJul 23, 2026 • new comments - [Scheduler] Fix KeyError in PP2 when last stage finishes request via tool-call parser
#46476 commented on
2 days agoJul 27, 2026 • new comments - [ROCm] cap mem_get_info total to local VRAM to fix 192 GiB over-allocation (#36890)
#46477 commented on
4 days agoJul 25, 2026 • new comments - [ROCm] honour skip_kv_gather in AITER MLA sparse prefill chunk loop (#40018)
#46478 commented on
4 days agoJul 25, 2026 • new comments - [Attention][MLA] FlashMLA sparse: DCP on the fp8_ds_mla mixed-batch path + MTP (stacked on #46076)
#46514 commented on
8 hours agoJul 29, 2026 • new comments - [Bugfix] Fix IPC handle leak in CustomAllreduce during CUDA graph memory profiling
#46515 commented on
5 hours agoJul 29, 2026 • new comments - [HelionLinearBackend][3/N] Add HelionFP8ScaledMMLinearKernel
#46522 commented on
2 days agoJul 28, 2026 • new comments - Enable humming wNaM assymetric quant (zero_point) wih compressed-tensors
#46528 commented on
yesterdayJul 28, 2026 • new comments - [codex] Optimize FP8 group 128 quantization for Hopper
#46541 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Use block_k for block-wise FP8 activation group_shape
#46593 commented on
last weekJul 23, 2026 • new comments - Support Vit CudaGraph for v2
#46606 commented on
yesterdayJul 28, 2026 • new comments - [Bugfix] Parse compact sentence-transformers pooling_mode (>=5.4.0)
#46614 commented on
1 hour agoJul 30, 2026 • new comments - fix(quant): resolve unquantized embedding method and key mismatch in inc path for MiniMax-M3
#46630 commented on
last weekJul 23, 2026 • new comments - [Spec Decode] Add FlashInfer CuteDSL non-causal decode path for DFlash
#46638 commented on
last weekJul 23, 2026 • new comments - [Kernel] Support batch invariance for WNA16 Marlin MoE
#46639 commented on
yesterdayJul 29, 2026 • new comments - Fix gpt-oss-20b NVFP4 (ModelOpt) inference
#46645 commented on
last weekJul 23, 2026 • new comments - [ModelRunner v2] Enable by default for all non-pooling models
#46646 commented on
yesterdayJul 28, 2026 • new comments - [Core][V1] Support trace_decode_token_ids for deterministic decode replay
#46701 commented on
4 hours agoJul 29, 2026 • new comments - [CI] Add registry layer cache to x86 CUDA release image builds
#46711 commented on
8 hours agoJul 29, 2026 • new comments - [ROCm][DSV4] B-preshuffle the attention fp8 projections
#46720 commented on
yesterdayJul 29, 2026 • new comments - [Feat] Support thinking_token_budget in Model Runner V2
#46727 commented on
5 hours agoJul 29, 2026 • new comments - [Kernel][MoE] Integrate TokenSpeed Mxfp4 MOE Kernel
#46732 commented on
last weekJul 23, 2026 • new comments - [ROCm] [Feat] TokenSpeed MHA integration for GPTOSS
#46742 commented on
last weekJul 23, 2026 • new comments - Recover from P0/P1 processor cache drift (#46747)
#46747 commented on
20 hours agoJul 29, 2026 • new comments - [WIP][Feature] A new 2-bit KV cache quantisation backend that cuts 5x memory than FP16 (Oscar-2)
#46774 commented on
last weekJul 23, 2026 • new comments - [Feature] NVFP4 dispatch for fused RoPE quantization
#46031 commented on
yesterdayJul 28, 2026 • new comments - [Bugfix] Add support for SWA draft models in speculative decoding
#46032 commented on
12 hours agoJul 29, 2026 • new comments - [Misc] Add unit test for l2norm_fwd_kernel1 and l2norm_fwd_kernel2
#46059 commented on
last weekJul 23, 2026 • new comments - [Bugfix][TurboQuant] Fix dangling decode scratch when workspace grows after cudagraph capture
#46067 commented on
last weekJul 23, 2026 • new comments - [Config] Reject negative values for max_logprobs and long_prefill_token_threshold
#46068 commented on
2 days agoJul 28, 2026 • new comments - [Misc] Add unit test for per_token_quant_int8 kernel
#46073 commented on
last weekJul 23, 2026 • new comments - [Compile] Add aot_eager backend for piecewise compilation
#46085 commented on
3 days agoJul 27, 2026 • new comments - [Kernel] Start SWA/chunked KV-tile loop at first allowed key (#44575)
#46087 commented on
4 days agoJul 25, 2026 • new comments - [ROCm][Perf] TP-shard Lightning indexer prefill
#46103 commented on
2 days agoJul 28, 2026 • new comments - fix: resolve vLLM performance and API issues
#46120 commented on
last weekJul 23, 2026 • new comments - [ROCm][Perf] FlyDSL BF16 MoE for MiniMax-M3 MXFP8 emulation (gfx942) via --moe-backend aiter
#46123 commented on
20 hours agoJul 29, 2026 • new comments - [Bugfix] Make Kimi's tool parser accept numeric only tool call IDs
#46127 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Fix Humming kernel weight conversion for unpacked CompressedTensors W4A8 INT checkpoints
#46136 commented on
last weekJul 23, 2026 • new comments - [ROCm] Enable RDNA3 W4A16 GEMM kernels on gfx1151 (Strix Halo)
#46186 commented on
last weekJul 23, 2026 • new comments - [ROCm][Perf] Avoid fp32 round-trip dequant in fp8 KV paged decode
#46191 commented on
last weekJul 23, 2026 • new comments - [Bugfix][V1] Warm up KV block zeroing before JIT monitor
#46215 commented on
5 hours agoJul 29, 2026 • new comments - [Kernel] ReplaySSM opt-in AR decode for Mamba2 (#46187)
#46239 commented on
5 days agoJul 25, 2026 • new comments - [Hybrid] Decoupling block_size from allocation block size
#46251 commented on
3 days agoJul 27, 2026 • new comments - [Kernel] Add fused SiLU+Mul+PerTokenQuant CUDA kernel
#46273 commented on
last weekJul 23, 2026 • new comments - fix: resolve issue #35344
#46297 commented on
5 hours agoJul 29, 2026 • new comments - fix: resolve issue #26211
#46318 commented on
last weekJul 23, 2026 • new comments - fix: resolve issue #23826
#46320 commented on
last weekJul 23, 2026 • new comments - [Quantization][INC] Support hybrid INT4+FP8 AutoRound checkpoints via maybe_update_config
#46322 commented on
2 days agoJul 28, 2026 • new comments - fix(cudagraph): align spec-decode capture sizes for PIECEWISE mode
#46324 commented on
19 hours agoJul 29, 2026 • new comments - HiSparse: host-resident sparse-MLA decode hot-buffering + GLM-5.2 indexCache opts (vLLM v0.24.0)
#46326 commented on
last weekJul 24, 2026 • new comments - [Attention][Quantization] NVFP4 KV cache on consumer/SoC Blackwell (sm120/sm121) for Gemma 3/4 via FlashInfer FA2
#46329 commented on
yesterdayJul 29, 2026 • new comments - [V1][Core] Add non-blocking get_output_nowait() to engine core client
#46331 commented on
2 days agoJul 28, 2026 • new comments - [Doc] Fix %% rendering in CLI reference for --safetensors-load-strategy
#46335 commented on
3 days agoJul 26, 2026 • new comments - [rust][tool-parser] Add step3 tool parser
#46369 commented on
16 hours agoJul 29, 2026 • new comments - [Feature] Initial support for fault tolerant ep using scale-down
#46370 commented on
4 hours agoJul 29, 2026 • new comments - fix: guard init_fp8_kv_scales against unallocated tensors during wake up after sleep
#46371 commented on
5 hours agoJul 29, 2026 • new comments - fix(spec_decode): bypass embedding dim check for MTP speculative decoding methods
#49036 commented on
2 days agoJul 27, 2026 • new comments - Make `load_weights` completely optional
#49038 commented on
2 days agoJul 27, 2026 • new comments - [Bugfix] Propagate VLLM_MARLIN_INPUT_DTYPE into INC/AutoRound quant path
#49039 commented on
last weekJul 23, 2026 • new comments - [Docs] quantized_kvcache: add measured performance section, replace deprecated gated example model
#49050 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Reject negative --device-ids indices instead of selecting the wrong device
#49053 commented on
4 days agoJul 26, 2026 • new comments - [Bugfix][MiniMax-M3] Avoid NaNs for empty sparse decode rows
#49054 commented on
2 days agoJul 27, 2026 • new comments - [Quant] Add online NVFP4 dense-linear quantization (W4A16 + W4A4)
#49060 commented on
last weekJul 23, 2026 • new comments - [Bugfix][KV Connector] Propagate EAGLE state across merged Mooncake store groups
#49069 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Forward FlashInfer skip ops to sparse MLA warmup
#49072 commented on
yesterdayJul 29, 2026 • new comments - [Core][V1] Enforce max_num_partial_prefills and max_long_partial_prefills in V1 scheduler
#49075 commented on
3 days agoJul 26, 2026 • new comments - TRITON_ATTN support for KV cache dtype fp8 on SM75 to pre-SM89
#49077 commented on
4 days agoJul 26, 2026 • new comments - [Perf] fused_moe: load B weights via TMA tensor descriptor on SM90+
#49078 commented on
11 hours agoJul 29, 2026 • new comments - [Feature] Add IndexCache support for DeepSeek V4 on Hopper
#49085 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Guard scaled_fp4_quant against non-contiguous input
#49092 commented on
last weekJul 23, 2026 • new comments - Improve Marlin MoE wide-N alignment
#49099 commented on
2 days agoJul 28, 2026 • new comments - [Misc] Bump `openai` to `>=2.25.0` to support namespace tools types
#49104 commented on
yesterdayJul 29, 2026 • new comments - [Bugfix]: move FlashInfer MoE autotune before memory profiling + logging
#49115 commented on
yesterdayJul 29, 2026 • new comments - [Bugfix][Spec Decode] DSpark: build draft under its own model/quant config (NVFP4 target corrupts MXFP4 draft experts)
#49133 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Kernel] Fix persistent top-k histogram reuse after short rows
#49139 commented on
2 days agoJul 28, 2026 • new comments - Enable MLA RoPE KV-cache fusion by default
#49142 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix] Prevent stale partial prefix cache hashes
#49145 commented on
last weekJul 23, 2026 • new comments - [Perf] Avoid per-step pinned staging allocation in apply_grammar_bitmask
#49150 commented on
2 days agoJul 28, 2026 • new comments - [Multimodal] Reorganize video decoder backends
#49155 commented on
4 days agoJul 25, 2026 • new comments - [ROCm] Added warmup for AITER unified attention
#49170 commented on
18 hours agoJul 29, 2026 • new comments - [Perf] Skip logits and sampling for unfinished prefills
#49171 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Disable DeepGEMM on GH200 to workaround illegal memory access
#49183 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Test] Quantize q to fp8 in test_flashmla_dense_fp8_decode_unified_slot_view
#49191 commented on
last weekJul 23, 2026 • new comments - [Perf][MoE] Eliminate staging from NCCL symmetric reduce-scatter
#49194 commented on
6 hours agoJul 29, 2026 • new comments - [Bugfix] Fix MiniMax M3 index_topk kernel for non-power-of-2 num_idx_heads (#49157)
#49199 commented on
3 hours agoJul 29, 2026 • new comments - [Refactor][Model Loader] Unify the weight loading lifecycle
#49201 commented on
4 days agoJul 26, 2026 • new comments - fix: resolve silent request skipping in PRIORITY scheduling
#49206 commented on
3 hours agoJul 29, 2026 • new comments - [Bugfix][Quantization] Respect explicit WNA16 MoE backend
#49207 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix humming crash on ParallelLMHead (output_partition_sizes, has_bias)
#48856 commented on
last weekJul 23, 2026 • new comments - fix: NVFP4 quantization out_dtype should match model dtype, not torch default
#48861 commented on
last weekJul 23, 2026 • new comments - reasoning: treat no-think-block output as content, not reasoning
#48865 commented on
3 days agoJul 26, 2026 • new comments - [Model][Quant] Fused WNA16 GEMM for tied quantized lm_head logits
#48870 commented on
2 days agoJul 27, 2026 • new comments - [Core] Add opt-in eager PyNccl TP all-reduce split for PIECEWISE CUDA graphs
#48877 commented on
3 days agoJul 27, 2026 • new comments - [LoRA][Perf] Zero-slice early-exit for LoRA kernels
#48887 commented on
2 days agoJul 28, 2026 • new comments - [DO NOT REVIEW][Kernel][Helion] Eager call-site routing for per_token_group_fp8_quant
#48888 commented on
2 hours agoJul 30, 2026 • new comments - [HelionLinearBackend][2/N] Add helion_cutlass_hybrid_scaled_mm c++ kernel
#48889 commented on
2 days agoJul 28, 2026 • new comments - [ROCm] Enable AITER by default
#48890 commented on
2 days agoJul 28, 2026 • new comments - [Model Runner V2][Spec Decode] Add multi-layer MTP speculator
#48892 commented on
13 minutes agoJul 30, 2026 • new comments - [Attention] Add direct symmetric-memory DCP A2A
#48897 commented on
2 days agoJul 28, 2026 • new comments - [Kernel][Helion] Retune grouped quant schedules
#48900 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Reload] Preserve unmanaged tensor addresses across weight reload
#48902 commented on
last weekJul 23, 2026 • new comments - Hiprune gemma4
#48903 commented on
5 days agoJul 24, 2026 • new comments - [Perf][MRV2] use FlashInfer AIR for for top-p rejection sampling
#48928 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Model] Fix MiniMax-M3 NVFP4 inference correctness
#48929 commented on
4 hours agoJul 29, 2026 • new comments - [Model]Fuse MiniMax-M3 dense-layer KV-cache insert into qknorm+rope k…
#48935 commented on
5 days agoJul 24, 2026 • new comments - [Benchmark] KV Cache Offload Benchmark — Block Copy Performance
#48936 commented on
2 days agoJul 28, 2026 • new comments - [ROCm] Support MiniMax-M3 NVFP4 SwiGLU-OAI
#48939 commented on
last weekJul 24, 2026 • new comments - [Spec Decode] Context-length-aware K in DSD (RFC #48627): extend num_speculative_tokens_per_batch_size with a ctx axis
#48944 commented on
2 days agoJul 27, 2026 • new comments - [ROCm][Model] Add Inkling BF16 support
#48954 commented on
2 hours agoJul 30, 2026 • new comments - [ROCm][MoE] Avoid redundant AITER MXFP8 MoE output copy
#48960 commented on
12 hours agoJul 29, 2026 • new comments - [Kernel][Helion] Retune dynamic quant schedules on H100
#48964 commented on
last weekJul 23, 2026 • new comments - [Test] e2e hybrid-Mamba prefix-cache corruption regression tests (#43559)
#48970 commented on
5 hours agoJul 29, 2026 • new comments - [Feature] Add mamba ssu kernel autotune at warmup time
#48980 commented on
last weekJul 24, 2026 • new comments - [Bugfix][Structured Output] Wire whitespace_pattern through V1 backends
#48982 commented on
last weekJul 23, 2026 • new comments - [Kernel] Opt-in FP8 requant of bf16 attention proj on NVFP4 checkpoints
#48983 commented on
19 hours agoJul 29, 2026 • new comments - [Kernel][Helion] Add packed per-token group FP8 quant
#48991 commented on
last weekJul 23, 2026 • new comments - [Helion] Route fusion-only kernels to Helion during CUDA-graph capture
#48995 commented on
5 days agoJul 25, 2026 • new comments - [XPU] Support sequence parallelism for block fp8 on XPU
#49009 commented on
yesterdayJul 28, 2026 • new comments - fix(v1): avoid false shutdown failures on clean exit
#49034 commented on
4 days agoJul 26, 2026 • new comments - fix: handle missing parent modules in _has_module
#49035 commented on
4 days agoJul 26, 2026 • new comments - [ModelOpt] Redesign the LinearMethod classes using the generic QuantKey-driven method
#49381 commented on
2 days agoJul 28, 2026 • new comments - [Kernel] Add FlashInferW4A16NvFp4LinearKernel (FlashInfer mm_bf16_fp4)
#49382 commented on
last weekJul 23, 2026 • new comments - DO NOT MERGE
#49383 commented on
yesterdayJul 29, 2026 • new comments - [WIP][Model Runner V2][Spec Decode] Enable rejection sampling with draft logits by default, capped to draft_logits_buffer_gb
#49386 commented on
2 days agoJul 28, 2026 • new comments - [Misc] Remove deprecated calculate_kv_scales runtime KV scale calculation
#49389 commented on
5 days agoJul 25, 2026 • new comments - [Perf] Raise Blackwell CUDA graph capture default to 1024
#49390 commented on
last weekJul 24, 2026 • new comments - [CI] Add opt-in Hugging Face Hub monitoring
#49401 commented on
last weekJul 23, 2026 • new comments - [XPU][DSv4] Use SYCL deepseek_inv_rope_fp8_quant kernel in DeepSeek-V4 o_proj
#49402 commented on
last weekJul 23, 2026 • new comments - [CPU] [Feat] Add native AMX-FP8 attention impl for Diamond Rapids
#49410 commented on
last weekJul 24, 2026 • new comments - [Docs] Clarify ASR audio chunking is non-overlapping
#49414 commented on
9 hours agoJul 29, 2026 • new comments - [XPU] [CI] Add MiniMaxM2ForCausalLM and NemotronParseForConditionalGeneration to model init test
#49416 commented on
11 hours agoJul 29, 2026 • new comments - [ROCm][Bugfix][Quark] Map weight/input quantizer scale names in AutoWeightsLoader
#49420 commented on
yesterdayJul 28, 2026 • new comments - [chore] delete useless code
#49424 commented on
last weekJul 24, 2026 • new comments - feat(mamba): add FlashInfer SSD prefill backend
#49425 commented on
5 days agoJul 25, 2026 • new comments - [Draft][Bugfix][Parser] Strip content whitespace at tool-call boundaries in streaming to match non-streaming
#49426 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix pipeline parallelism for Kimi-Linear
#49430 commented on
2 days agoJul 28, 2026 • new comments - [Draft] Simplify encoder cuda graph implementation
#49432 commented on
2 days agoJul 28, 2026 • new comments - [Model] Add support for Nanbeige4.2
#49433 commented on
2 days agoJul 27, 2026 • new comments - [Bugfix] Fix SM100 fp8_ds_mla cache scales
#49435 commented on
5 hours agoJul 29, 2026 • new comments - [chore] log process manager shutdown with more details
#49437 commented on
3 days agoJul 26, 2026 • new comments - [Bugfix][Core] Prevent streaming session deadlock under KV pressure
#49439 commented on
3 days agoJul 27, 2026 • new comments - Fix speculative drafter access on non-final pipeline parallel ranks
#49442 commented on
last weekJul 23, 2026 • new comments - Fix PP sampled token broadcast for speculative decoding
#49443 commented on
last weekJul 23, 2026 • new comments - [Misc] Enable test_silu_mul_fp8_quant_deep_gemm on XPU
#49444 commented on
last weekJul 23, 2026 • new comments - [Core] Add `max_num_queued_reqs` and `max_num_queued_tokens` for queue size management
#49445 commented on
12 hours agoJul 29, 2026 • new comments - [Bugfix] Fix multimodal streaming-session prefix rehashing
#49448 commented on
5 days agoJul 24, 2026 • new comments - [CPU] Add MLA backend so DeepSeek-V2/V3 can run on CPU
#49453 commented on
10 hours agoJul 29, 2026 • new comments - [Perf][Model Loader] Reload layout-identical weights directly
#49459 commented on
4 days agoJul 26, 2026 • new comments - [Bugfix][Structured Output] Cap grammar-compile executor workers to avoid CFS throttling in containers
#49461 commented on
2 days agoJul 28, 2026 • new comments - [TurboQuant][TritonCPU] Enable triton-cpu path for Turbo-quant algorithm
#49465 commented on
last weekJul 24, 2026 • new comments - [Core] Balance padding vs group count when grouping hybrid KV cache layers
#49472 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix MiniMax-M3 ModelOpt FP8 MoE SwiGLU params
#49473 commented on
last weekJul 23, 2026 • new comments - [CI/Build][The Rock] Use model_class_overrides so spawned worker can use test PredictableLlamaForCausalLM class when worker spawned using Python 3.14
#49218 commented on
1 hour agoJul 30, 2026 • new comments - [KV-offload][FS]: Batching for read/write threads
#49225 commented on
52 minutes agoJul 30, 2026 • new comments - [Bugfix][Structured Output] Mask request stop tokens in xgrammar until grammar terminates
#49227 commented on
last weekJul 24, 2026 • new comments - [Bugfix]: ensure previous_text and tokens id accumulation in streaming state for named and required tool choices
#49228 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Validate NIXL speculative config compatibility
#49230 commented on
54 minutes agoJul 30, 2026 • new comments - [Bugfix][V1] Reserve CUDA graph memory in V2 GPU model runner
#49233 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] glm47 parser: tolerate missing opening <arg_value> tag
#49249 commented on
4 days agoJul 26, 2026 • new comments - [Kernel][MoE] Add opt-in DeepGEMM BF16 grouped-GEMM MoE backend (contiguous + masked)
#49272 commented on
4 days agoJul 26, 2026 • new comments - [Core] Fix EngineDeadError on late KV transfer completion
#49278 commented on
2 hours agoJul 30, 2026 • new comments - [Rust Frontend] Allow media placeholder target to be a list of tokens
#49279 commented on
2 hours agoJul 30, 2026 • new comments - [Bug Fix] Indexer Cache need indexer num equal to attention num
#49286 commented on
last weekJul 24, 2026 • new comments - [Bugfix][Attention] Keep model-dtype query for FlashInfer builders serving non-causal attention
#49293 commented on
2 days agoJul 27, 2026 • new comments - [XPU] Support Sequence Parallelism for mxfp8 on XPU
#49303 commented on
yesterdayJul 28, 2026 • new comments - Add vllm:kv_offload_cpu_total_blocks capacity metric
#49307 commented on
2 days agoJul 27, 2026 • new comments - [MoE][Kernel] Add optional HPC BF16xFP32 router GEMM
#49312 commented on
yesterdayJul 28, 2026 • new comments - [Quark] Online-requant unquantized layers via quantization_config ove…
#49313 commented on
3 days agoJul 27, 2026 • new comments - fix(sampler): fall back to native sampling when flashinfer is absent
#49314 commented on
5 days agoJul 25, 2026 • new comments - [Feature][MoE]W4A16 cold-expert CPU offload via in-graph HIP gather (experimental)
#49321 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Kernel] Fixes issue #49290: MRotaryEmbedding Triton kernel hardcodes Neox pairing for GPT-J style models
#49323 commented on
yesterdayJul 29, 2026 • new comments - [XPU][LoRA] Fix fused MoE LoRA shrink BLOCK_SIZE_N for small ranks
#49327 commented on
yesterdayJul 28, 2026 • new comments - [Frontend] Add --max-waiting-queue-length to bound the waiting queue
#49330 commented on
last weekJul 23, 2026 • new comments - [Test] Force enforce_eager for pythia-70m in test_models to avoid compile-vs-eager drift
#49332 commented on
2 days agoJul 27, 2026 • new comments - [Bugfix] Fix DBRX MoE weight loading after FusedMoE/MoERunner refacto…
#49336 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix][KV Connector][NIXL] Support matching DCP layouts with PCP
#49342 commented on
yesterdayJul 28, 2026 • new comments - fix: accept empty tool call list in case of none required tool_choice
#49346 commented on
12 hours agoJul 29, 2026 • new comments - [Doc] Add Crusoe Managed Inference deployment guide
#49353 commented on
2 days agoJul 28, 2026 • new comments - [BugFix] bound FlashMLA sparse decode intermediate tensors size
#49357 commented on
2 hours agoJul 30, 2026 • new comments - [Attention][MiniMax-M3] Add opt-in FlashInfer TRTLLM-GEN sparse decode
#49358 commented on
last weekJul 23, 2026 • new comments - [DONOTMERGE][ROCm]: bump AITER to 0.1.19
#49361 commented on
yesterdayJul 29, 2026 • new comments - [Perf] Split-reduction batch-invariant log_softmax for small-batch decode
#49367 commented on
3 days agoJul 26, 2026 • new comments - [Bugfix][ROCm] Fix ROCM_AITER_FA & ROCM_AITER_UNIFIED_ATTN QK-Norm+RoPE+KVCache fusion for the packed KV-cache [BLOCKS, HEADS, BLOCK_SIZE, 2*HEAD_DIM] layout
#49373 commented on
1 hour agoJul 30, 2026 • new comments - [ROCm][CI] Add More AITER quantization/MoE kernel tests
#49375 commented on
last weekJul 23, 2026 • new comments - Fa4 fp8 kv dequant integration clean
#48192 commented on
last weekJul 23, 2026 • new comments - [Quantization] Fix misleading NVFP4 Marlin linear warning on FP4-native GPUs
#48199 commented on
last weekJul 23, 2026 • new comments - [Refactor]: StructuredOutputManager x Speculative Decoding Refactor
#48200 commented on
last weekJul 24, 2026 • new comments - [CPU] Add CPU-tuned autotune configs for FLA (GDN) Triton kernels
#48212 commented on
11 hours agoJul 29, 2026 • new comments - [CPU] Add CPU-tuned autotune configs for Mamba2/SSD Triton kernels
#48213 commented on
11 hours agoJul 29, 2026 • new comments - [Bugfix] Fix pooling input buffer race across chunked prefill steps
#48214 commented on
14 hours agoJul 29, 2026 • new comments - [Bugfix] Fix cumulative Whisper timestamp drift from variable-length audio chunks
#48225 commented on
2 days agoJul 28, 2026 • new comments - [KV Connector][Mooncake] Skip lookup for locally reused prefix
#48230 commented on
last weekJul 23, 2026 • new comments - [Attention][TurboQuant] Optimize 4-bit GQA decode for group size 4
#48235 commented on
last weekJul 23, 2026 • new comments - fix: Domino inference patches for speculators-format checkpoints
#48241 commented on
5 days agoJul 24, 2026 • new comments - [Core] Skip redundant draft token alloc + sampling
#48244 commented on
last weekJul 24, 2026 • new comments - [Perf][ROCm] Add AITER custom AG/RS
#48247 commented on
3 days agoJul 27, 2026 • new comments - Support MLA properly in the Transformers modeling backend
#48250 commented on
yesterdayJul 28, 2026 • new comments - [KV Connector][4/N][NIXL] Recover all dedup'd HMA pool members (FA + multi-group SSM) under PP
#48263 commented on
1 hour agoJul 30, 2026 • new comments - [Model] Add GraniteSWA and GraniteMoeSWA
#48270 commented on
yesterdayJul 29, 2026 • new comments - Pin RMSNorm block size under batch-invariant mode
#48272 commented on
2 hours agoJul 30, 2026 • new comments - [Model] Add Nemotron Ministral masked block-diffusion LM
#48279 commented on
12 hours agoJul 29, 2026 • new comments - [RFC] Heterogeneous rank-to-GPU mapping + Qwen3.5/3.6 GGUF enablement
#48280 commented on
2 days agoJul 28, 2026 • new comments - Skip layerwise reload for non-quantized models
#48284 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Tool Parser] Stream tool calls that arrive whole in a single delta in llama3_json, jamba and ernie45 parsers
#48295 commented on
4 days agoJul 25, 2026 • new comments - [Bugfix][Frontend] Flatten list message content when echoing in streaming chat completions
#48339 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix][Tool Parser] InternLM2: stream tool calls that arrive whole in a single delta
#48348 commented on
4 days agoJul 26, 2026 • new comments - [Bugfix] Hermes tool parser: parse tool call JSON by object boundary, not literal </tool_call>
#48353 commented on
last weekJul 23, 2026 • new comments - feat: extended EPLB support for Mistral Large 3 and additional MoE backends
#48355 commented on
10 hours agoJul 29, 2026 • new comments - [BugFix] Honor drop_eagle_block in MambaManager
#48375 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Spec Decode] DSpark: store draft KV cache in model dtype for MLA-only cache layouts; fail fast on DCP
#48381 commented on
4 days agoJul 26, 2026 • new comments - [Spec Decode] DFlash/DSpark draft support under decode context parallelism
#48392 commented on
4 days agoJul 26, 2026 • new comments - [Bugfix] Warn when prefix caching is inert on hybrid models (block_size > max_model_len)
#48402 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Size the sparse-indexer expanded block table from the runner's block-table width
#48404 commented on
3 hours agoJul 29, 2026 • new comments - [Bugfix][MM] Fix MiniCPM-V placeholder replacement and image processor loading on Transformers v5
#48413 commented on
5 days agoJul 24, 2026 • new comments - [KV Connector] Canonical CPU layout for parallelism-agnostic KV offload
#48414 commented on
3 days agoJul 26, 2026 • new comments - [Bugfix] V1: fix allowed_token_ids_mask aliasing in InputBatch.swap_states
#48419 commented on
last weekJul 23, 2026 • new comments - [KV Cache] Native simulator
#47922 commented on
2 days agoJul 27, 2026 • new comments - [Core] Pre-size cudagraph output staging buffers to the max capture descriptor
#47925 commented on
3 days agoJul 27, 2026 • new comments - Account scheduled spec slots on empty output rows
#47928 commented on
2 days agoJul 28, 2026 • new comments - [EC Connector] P2P NIXL + CPU EC Connector
#47941 commented on
last weekJul 23, 2026 • new comments - [MoE] Add flashinfer.moe_ep (NCCL-EP) all2all backends: flashinfer_ep_low_latency / flashinfer_ep_high_throughput
#47948 commented on
9 hours agoJul 29, 2026 • new comments - [Bugfix] Report finish_reason='length' for tool calls truncated by max_tokens in streaming
#47963 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix][Tool Parser] Return early for unlisted tool names in InternLM2 parser
#47968 commented on
4 days agoJul 26, 2026 • new comments - Support DeepSeek-V4 AMD Quark NVFP4 with emulation kernel
#47972 commented on
9 hours agoJul 29, 2026 • new comments - [Misc] Add unit test for fused_recurrent_gated_delta_rule kernel
#47976 commented on
last weekJul 23, 2026 • new comments - [Perf] SM120 PCIe serving stack: SP/async-TP enablement, FlashInfer spec-decode FULL cudagraphs, and PCIe-safe multi-GPU comms
#47979 commented on
16 hours agoJul 29, 2026 • new comments - [Bugfix] Handle E8M0 block scales in CUTLASS and Triton FP8 linear kernels
#47988 commented on
last weekJul 23, 2026 • new comments - [Rust Frontend] add mistral reasoning parser
#48013 commented on
9 hours agoJul 29, 2026 • new comments - Fix/qwen3 vl text lora noop
#48022 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Kernel] Make Marlin MoE route alignment deterministic
#48032 commented on
2 hours agoJul 30, 2026 • new comments - [DSv4] Remove sparse-MLA q-head padding for FlashInfer >=0.6.14
#48047 commented on
last weekJul 23, 2026 • new comments - [ROCm] Don't re-quantize bf16 MLA kv_b_proj to fp4/fp8 for GLM MoE DSA
#48051 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Hardware] DeepSeek-V4 o_proj fp8 einsum NaNs on SM12x: use SM90-style raw f32 block scales
#48052 commented on
last weekJul 23, 2026 • new comments - [BugFix][Mooncake] Use global data_parallel_index for the DP engine index
#48061 commented on
2 days agoJul 27, 2026 • new comments - [Bug fix] Always pass launch_pdl kwarg in fused_inv_rope_fp8_quant (DeepSeek-V4)
#48086 commented on
last weekJul 23, 2026 • new comments - [Render][1/n] Paged shared memory storage for multimodal.
#48088 commented on
4 days agoJul 26, 2026 • new comments - fix(speech_to_text): use real per-chunk offsets for segment timestamps
#48104 commented on
2 days agoJul 28, 2026 • new comments - [OOT] Add OOT support for fp8_block linear kernel
#48105 commented on
2 days agoJul 27, 2026 • new comments - [Bugfix][XPU] Fix Mamba state pointer overflow
#48109 commented on
10 hours agoJul 29, 2026 • new comments - Support online C128 compression for DeepSeek V4
#48119 commented on
last weekJul 23, 2026 • new comments - [Hybrid] Stage the postprocess inputs with a single loop over the request list
#48120 commented on
2 days agoJul 27, 2026 • new comments - [Frontend] Expose canonical KV cache group metadata
#48121 commented on
4 days agoJul 26, 2026 • new comments - [Core] Add Min-k sampling: temperature-invariant logit-space truncation
#48142 commented on
12 hours agoJul 29, 2026 • new comments - [Structured Output] Compose auto tool-call grammar with response_format schema
#48157 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix] Migrate Olmo3 reasoning parser to the streaming parser engine
#48160 commented on
last weekJul 24, 2026 • new comments - [ROCm][Bugfix] Pad block-FP8 MoE intermediate size for TP when not divisible by block_n
#48173 commented on
last weekJul 23, 2026 • new comments - [Bugfix][TurboQuant] Preserve KV cache dtype in reshape KV tensors
#48177 commented on
4 days agoJul 25, 2026 • new comments - [Core] Make the DeepGEMM JIT cache key portable across install layouts
#48190 commented on
3 days agoJul 27, 2026 • new comments - [Kernel] TRITON_MLA SWA
#48605 commented on
yesterdayJul 29, 2026 • new comments - [Bugfix][NVFP4/FP8 MoE] Fix gated MoE crash on unaligned intermediate
#48624 commented on
last weekJul 23, 2026 • new comments - [XPU] Fix Dockerfile.xpu: add /opt/venv/lib to LD_LIBRARY_PATH
#48634 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix][LoRA] Avoid block_n=128 for lora_expand on Hopper
#48638 commented on
5 days agoJul 24, 2026 • new comments - [ROCm][CI] Reuse equivalent ROCm CI images
#48646 commented on
13 hours agoJul 29, 2026 • new comments - fix: Add FP8 type dispatch to per_token_group_quant_8bit for transformers backend
#48653 commented on
last weekJul 23, 2026 • new comments - [Kernel] Gemma-4 FA4 FP8 Kernel
#48666 commented on
last weekJul 23, 2026 • new comments - [V1][Metrics] Preserve prefix-cache stats on zero-output steps
#48668 commented on
18 hours agoJul 29, 2026 • new comments - test(quantization): cover tied lm_head/embed_tokens when lm_head excluded from ModelOpt
#48673 commented on
last weekJul 23, 2026 • new comments - [Misc] Remove `override_attention_dtype`
#48684 commented on
last weekJul 23, 2026 • new comments - [Model] Add minimal native RWKV7 serving support
#48686 commented on
3 days agoJul 26, 2026 • new comments - [MRV2][Spec Decode] Adaptive Speculative Decoding - Initial Support
#48692 commented on
yesterdayJul 28, 2026 • new comments - [Bugfix] Avoid FlashAttention 3 with batch invariance on Hopper
#48694 commented on
5 days agoJul 25, 2026 • new comments - [Bugfix] Share FlashInfer B12x MoE workspace across layers
#48698 commented on
last weekJul 23, 2026 • new comments - [BugFix] Fix Jamba tool parser deleting all occurrences of the streamed prefix on tool transition
#48709 commented on
4 days agoJul 25, 2026 • new comments - Synchronize distributed FlashInfer autotuning and persist cache
#48714 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Core] Reserve the null block in auto-fit max_model_len
#48724 commented on
3 days agoJul 26, 2026 • new comments - [Perf][Kernel] Fused DSA indexer Top-k kernel (LiteTopk)
#48726 commented on
5 days agoJul 24, 2026 • new comments - [ROCm][Perf] Use AITER tgemm for DeepSeek V4 compressors
#48727 commented on
yesterdayJul 29, 2026 • new comments - [Perf] Improve `--linear-backend` filtering
#48735 commented on
2 days agoJul 27, 2026 • new comments - [Bugfix] Update shared expert config during elastic EP reconfiguration
#48756 commented on
last weekJul 23, 2026 • new comments - [PD][NixlPush][Bugfix] Fix prefix caching
#48758 commented on
49 minutes agoJul 30, 2026 • new comments - [WIP][XPU][Test]add xpu yaml
#48761 commented on
14 hours agoJul 29, 2026 • new comments - [Model] Add Inkling multi-depth MTP support [5/N]
#48768 commented on
5 days agoJul 25, 2026 • new comments - [Doc] Expand ModelOpt NVFP4 docs: hardware support, MoE serving, accuracy evaluation
#48782 commented on
3 days agoJul 27, 2026 • new comments - [Rust Frontend] Add hy_v3 reasoning parser
#48800 commented on
15 hours agoJul 29, 2026 • new comments - [Spec Decode][V1] Warm Eagle and DFlash/DSpark spec-decode Triton kernels at startup
#48804 commented on
2 hours agoJul 30, 2026 • new comments - Fix MTP Mamba align prefix cache retention
#48815 commented on
last weekJul 23, 2026 • new comments - Perf/h20 moe config e256 n512
#48825 commented on
14 hours agoJul 29, 2026 • new comments - [Bugfix][Multimodal] Fix M-RoPE media mapping for chunked prefill and multi-media chunks
#48835 commented on
11 hours agoJul 29, 2026 • new comments - [ROCm][CI] Loosen block-FP8 fused MoE test tolerance for large-K shapes
#48847 commented on
8 hours agoJul 29, 2026 • new comments - [Bugfix][LoRA] Add embedding_modules for Qwen3.5 CausalLM
#48850 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix Qwen3-Omni crash on video with no audio track when use_audio_in_video=True
#48420 commented on
5 days agoJul 24, 2026 • new comments - [Bugfix][Quantization] Run block kernel post-processing for ModelOpt FP8_PB_WO
#48422 commented on
last weekJul 23, 2026 • new comments - fix(sampling_params): check top_k type before comparing it to -1
#48426 commented on
1 hour agoJul 30, 2026 • new comments - [ROCm][Quant] Requantize serialized MXFP8 linears to FP8 PTPC
#48427 commented on
2 days agoJul 28, 2026 • new comments - [CI/Build] Regression test: CPU MoE activation table must not require a config context
#48430 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Preserve Marlin runtime tensor storage across weight reload
#48438 commented on
last weekJul 24, 2026 • new comments - Revert "[Bugfix] Guard mixed-dtype allreduce RMSNorm quant fusions" (#48330)
#48447 commented on
last weekJul 23, 2026 • new comments - [ROCm][P/D] Fix MoRIIO READ prefix-cache block matching
#48471 commented on
3 days agoJul 27, 2026 • new comments - Update vendored LMCache connector for current IPCCacheServerKey API
#48479 commented on
last weekJul 23, 2026 • new comments - [Core] Waiting-Queue-Informed LRU for Prefix Cache Eviction
#48488 commented on
last weekJul 23, 2026 • new comments - [Performance] Add Triton kernel for Gemma3n sparse GELU
#48498 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix] Initialize tracing in API server worker processes
#48502 commented on
2 days agoJul 27, 2026 • new comments - feat(mxfp4): add MXFP4 expert cache support to DeepSeek V4
#48505 commented on
last weekJul 23, 2026 • new comments - fix(models): pass quant_config to eh_proj in MTP layers to prevent silent precision loss
#48506 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Pass quant_config to DiffusionGemma's ParallelLMHead
#48521 commented on
last weekJul 23, 2026 • new comments - [Bugfix] GLM 5.2: Fix skip_topk on the fused DSA path (MTP index share)
#48528 commented on
5 days agoJul 25, 2026 • new comments - [Bugfix][V1] Preserve per-group eviction order in deferred block free
#48532 commented on
last weekJul 23, 2026 • new comments - [Bugfix][KV-transfer] MoRIIO: per-layer READ-completion barrier in wait_for_layer_load
#48534 commented on
15 hours agoJul 29, 2026 • new comments - [Frontend] Add diarized_json support for MOSS-Transcribe-Diarize
#48543 commented on
4 hours agoJul 29, 2026 • new comments - feat: integrate ScaleSweep NVFP4 quantization kernel
#48548 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Initialize MiniMax M3 reasoning from prompt mode
#48550 commented on
last weekJul 24, 2026 • new comments - [XPU] Route INC WNA16 MoE to oracle backend instead of bf16 dequant
#48555 commented on
12 hours agoJul 29, 2026 • new comments - Bump astral-sh/setup-uv from 7.6.0 to 9.0.0
#48556 commented on
16 hours agoJul 29, 2026 • new comments - [Kernel][Test] Avoid NULL_BLOCK_ID in GDN sigmoid-gating test state indices
#48557 commented on
last weekJul 23, 2026 • new comments - fix(moe): add metaclass bridge to FusedMoE factory to preserve isinstance() type checks for out-of-tree plugins under PyTorch 2.12.0
#48565 commented on
last weekJul 24, 2026 • new comments - [Misc] Add unit test for chunk_fwd_kernel_o kernel
#48571 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Spec Decode] Warm spec-sized logits all-gather during init
#48572 commented on
16 hours agoJul 29, 2026 • new comments - [Bugfix][Kernel] Pass the correct expert count to WNA16 MoE block config
#48574 commented on
11 hours agoJul 29, 2026 • new comments - [Rust Frontend] Migrate to hf-hub 1.0
#48575 commented on
14 hours agoJul 29, 2026 • new comments - [Rust Frontend] Add support for truncate_prompt_tokens and truncation_side
#48584 commented on
2 hours agoJul 30, 2026 • new comments - [Bugfix][Frontend] Handle None/NaN/Inf logprob values when using FP8 quantization
#48585 commented on
last weekJul 23, 2026 • new comments - [MRV2] Profile flashinfer top-k/top-p sampler to fix warmup OOM
#48601 commented on
yesterdayJul 28, 2026 • new comments - [Bug]: `+rotary_embedding` error with DeepSeek-V3.2-NVFP4
#40587 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: XQA decode under FULL cudagraph capture silently corrupts attention output (fp8 and nvfp4 KV)
#49010 commented on
16 hours agoJul 29, 2026 • new comments - [Bug]: MiniMax-M3 streaming reasoning parser still leaks `<mm:think>` into content on latest vLLM
#46042 commented on
16 hours agoJul 29, 2026 • new comments - [RFC]: Sparse attention KV cache offloading to support longer sequence length
#33980 commented on
17 hours agoJul 29, 2026 • new comments - [Bug]: v0.19.1 Crash with CUDA invalid argument / Segfault when using KV Offloading + EAGLE3 + Expert Parallel (on 8x H20 141GB)
#40259 commented on
18 hours agoJul 29, 2026 • new comments - [Bug]: Eagle 2/3 acceptance length regression over time
#41838 commented on
18 hours agoJul 29, 2026 • new comments - [Feature]: Add SpectralQuant KV Cache Compression (builds on TurboQuant #38479)
#43475 commented on
19 hours agoJul 29, 2026 • new comments - [Feature]: Add INT8 Support for KV Cache Quantization (Currently FP8-Only)
#33480 commented on
yesterdayJul 29, 2026 • new comments - [Bug]: Gemma 4 KVCache CPU offloading broken
#42348 commented on
yesterdayJul 29, 2026 • new comments - [Bug][ROCm] Triton kernel_paged_attention_2d illegal memory access (hipErrorLaunchFailure) on gfx950 (MI350X) under long-context decode
#48043 commented on
yesterdayJul 29, 2026 • new comments - [RFC]: Multi-tier KV offloading via the vLLM offloading connector
#38260 commented on
yesterdayJul 29, 2026 • new comments - [Bug]: [kernel] MRotaryEmbedding Triton kernel hardcodes Neox-style rotation, producing wrong results for GPT-J style models (like GLM-OCR).
#49290 commented on
yesterdayJul 29, 2026 • new comments - [Bug]: Wrong timestamps if audio > 30s
#32588 commented on
yesterdayJul 29, 2026 • new comments - [Bug]: Hybrid Mamba PD-disagg under pipeline-parallel (PP>1) emits degenerate output — FullAttention region_group_ids collapse to the Mamba group under HMA
#46407 commented on
yesterdayJul 29, 2026 • new comments - [Bug]: 100% cpu usage on 3 cores on every node when using ray distributed pipeline parallel
#21231 commented on
yesterdayJul 28, 2026 • new comments - [Bug]: Decode instance segfaults on NIXL `loadRemoteMD` after prefill pod restarts in P/D disaggregation
#49238 commented on
yesterdayJul 28, 2026 • new comments - [Bug] DeepSeekV4-Flash produces incorrect output with inline system messages after PR #46025 when `preserved in-place`
#46710 commented on
yesterdayJul 28, 2026 • new comments - [Feature]: Sentence transformers embeddings support
#17493 commented on
yesterdayJul 28, 2026 • new comments - [vllm IR]: Remove `QuantFP8` in favour of direct `ir.ops` calls
#40617 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Running DeepSeek-V4 fails with CUDA error: unspecified launch failure in synchronize_input_prep
#40952 commented on
15 hours agoJul 29, 2026 • new comments - [RFC]: Refactor PassManager infrastructure
#40953 commented on
15 hours agoJul 29, 2026 • new comments - [Bug] Online FP8 quantization ignores logical_widths on MergedColumnParallelLinear
#41022 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Failed to apply MiniCPMVProcessor
#41073 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: vLLM only prints access logs, not performance statistics logs (v0.1.dev15830+g8d599d76a with deepseek-V4-flash)
#41081 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Qwen3_next a/b are not contiguous
#41112 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: unsupported architecture
#41117 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: aiohttp tracing leaks image urls and cannot be disabled
#41146 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Mimo v2.5 model loading fails from s3/remote locations
#41172 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: `sharded_state` load fails for FP8 models: `_filter_subtensors` drops `q_scale/k_scale/v_scale/prob_scale` parameters
#41174 commented on
15 hours agoJul 29, 2026 • new comments - [RFC]: Expert Weight Backup for Elastic EP
#41204 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: MoRIIO does not support heterogenous TP
#41211 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: Update default Python version in pre-built docker images
#41264 commented on
15 hours agoJul 29, 2026 • new comments - [Performance]: Encode performance of vLLM
#41267 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: V1 + Ray multi-node pipeline parallel `KeyError` at KV-cache init due to missing `global_rank` update
#41287 commented on
15 hours agoJul 29, 2026 • new comments - [Refactor] Merge `select_gpt_oss_mxfp4_moe_backend` and `select_mxfp4_moe_backend`
#41291 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Qwen3.5 NVFP4 models crash on ARM64 GB10 DGX Spark (CUDA illegal instruction during generation)
#35519 commented on
2 days agoJul 28, 2026 • new comments - [RFC]: Add Helion linear backend for vLLM
#46526 commented on
2 days agoJul 28, 2026 • new comments - [RFC]: Tensor descriptor (TD) adoption strategy for vLLM Triton kernels
#42545 commented on
2 days agoJul 28, 2026 • new comments - [Perf] ~2x decode throughput regression for structured outputs since #45424: apply_grammar_bitmask staging rewrite (bisected to commit, file, and hunk)
#49013 commented on
2 days agoJul 28, 2026 • new comments - vLLM crashes on OpenShift (gemma4-unified-cu129)
#44548 commented on
2 days agoJul 27, 2026 • new comments - [RFC]: Support ViT Full CUDA Graph (Tracker)
#38175 commented on
2 days agoJul 27, 2026 • new comments - [RFC]: Context-length-aware speculative token scheduling — extending num_speculative_tokens_per_batch_size with a context-length axis
#48627 commented on
2 days agoJul 27, 2026 • new comments - [RFC]: Add PPL and KLD to VLLM
#35962 commented on
2 days agoJul 27, 2026 • new comments - [Bug]: Engine core dies on KV-connector load failure for hybrid-KV models — `_update_requests_with_invalid_blocks` assumes a single KV cache group
#45474 commented on
2 days agoJul 27, 2026 • new comments - [Bug]: XPU TP=2 on dual Intel Arc Pro B70 (Battlemage): GP fault + xe BCS engine reset reproduces in intel/vllm:0.17.0-xpu on Ubuntu 24.04 HWE 6.17
#41663 commented on
2 days agoJul 27, 2026 • new comments - [RFC]: Per-instance EPLB metrics
#30696 commented on
2 days agoJul 27, 2026 • new comments - Expose public get_layer_params(config) helper for Gemma4 per-layer attention/FFN params
#48661 commented on
2 days agoJul 27, 2026 • new comments - [RFC]: Context-Aware KV-Cache Retention API (Prioritized Evictions)
#37003 commented on
2 days agoJul 27, 2026 • new comments - [Bug]: BF16 NVFP4 Marlin produces garbled output on GPUs without native FP4 support
#34694 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: Qwen 3.5 fails to load from GGUF
#38122 commented on
3 days agoJul 27, 2026 • new comments - [RFC]: Triton Kernel Dispatcher for Multi-Platform Support
#45133 commented on
5 days agoJul 24, 2026 • new comments - [RFC]: Intel Quantization Support Roadmap (H1 2026)
#37979 commented on
yesterdayJul 28, 2026 • new comments - [Bug][XPU]: zeMemOpenIpcHandle INVALID_ARGUMENT on 2x Arc B50 (Battlemage) TP=2
#48953 commented on
yesterdayJul 28, 2026 • new comments - [RFC]: GDS support for filesystem KV-cache offloading
#48504 commented on
yesterdayJul 28, 2026 • new comments - [Feature]: Deduplicate replicated MLA KV across TP ranks in native offloading
#47929 commented on
yesterdayJul 28, 2026 • new comments - [Bug][XPU] Mamba align-mode prefix caching crashes: "Overflow when unpacking long long" storing state.data_ptr()
#48059 commented on
yesterdayJul 28, 2026 • new comments - [Feature]: nvfp4 KV cache on SM120 — flashinfer ships the kernels, vLLM isn't wired to them (working prototype, 245K ctx on a 5090)
#49011 commented on
2 days agoJul 28, 2026 • new comments - [Docs] Document NIXL KV connector metrics aggregation semantics
#41230 commented on
2 days agoJul 28, 2026 • new comments - Optimize --help performance: Avoid torch import during help display
#33741 commented on
2 days agoJul 28, 2026 • new comments - [Bug]: V1 engine workers die after idle period (SystemError: PyCFunction / EngineDeadError) — TP=2, multiprocessing
#35104 commented on
2 days agoJul 28, 2026 • new comments - [RFC]: Revamp Ray Distributed Executor Backend (from Ray team)
#35848 commented on
2 days agoJul 28, 2026 • new comments - [Bug]: UMA Memory Profiling Misattributes OS Page Cache and Fails in Concurrent Deployments
#35920 commented on
2 days agoJul 28, 2026 • new comments - [Doc]: Inconsistent hash notation in Prefix Caching "Time 5" diagram
#35992 commented on
2 days agoJul 28, 2026 • new comments - [Bug]: Improve `--kv-cache-dtype` behavior when checkpoint specifies `kv_cache_quant_algo`
#34752 commented on
2 days agoJul 28, 2026 • new comments - [Roadmap] Rust Frontend Feature Parity
#44280 commented on
2 days agoJul 28, 2026 • new comments - [RFC] External post-generation classifier hook API
#43999 commented on
2 days agoJul 28, 2026 • new comments - [RFC]: Make Triton kernel unit tests hardware-agnostic and cover untested kernels
#48480 commented on
2 days agoJul 28, 2026 • new comments - [Bug]: prefix-caching: inconsistent completions
#5543 commented on
2 days agoJul 28, 2026 • new comments - Fix ASR verbose segment timestamps for split chunk boundaries
#43855 commented on
20 hours agoJul 29, 2026 • new comments - Fix merge conflict.
#33188 commented on
last weekJul 24, 2026 • new comments - [Feature] Add tuning script and config files for Mamba SSM
#33084 commented on
4 days agoJul 26, 2026 • new comments - [Fix] Include list index in multimodal validation error messages
#32985 commented on
last weekJul 24, 2026 • new comments - [Bench] Add interactive mode and sample caching to vllm bench serve
#32814 commented on
4 days agoJul 26, 2026 • new comments - [Quark] Support online block-diagonal rotations in dense GEMM layers
#32272 commented on
last weekJul 24, 2026 • new comments - [V1][Hybrid] GatedDeltaNet Automatic Prefix Caching (`all`-mode)
#26807 commented on
5 days agoJul 25, 2026 • new comments - [LoRA] Gemma3n LoRA support
#24003 commented on
4 days agoJul 26, 2026 • new comments - [Feature]: Performance Optimization for Deepseek V4
#45861 commented on
46 minutes agoJul 30, 2026 • new comments - [Bug]: Gemma4 MTP speculative decoding crashes at engine init on 0.25.1 — "a and b must have same reduction dim" (regression from 0.21.0)
#48848 commented on
1 hour agoJul 30, 2026 • new comments - [Bug]: GLM 5.2 MTP didn't work in AMD MI300x
#48568 commented on
3 hours agoJul 29, 2026 • new comments - [Bug]: Gemma 4: Unsloth LoRA adapters are ignored during inference despite successful loading
#41754 commented on
6 hours agoJul 29, 2026 • new comments - [Bug]: Including two images in a single request with Qwen3.6-27B causes an infinite rollback.
#47738 commented on
7 hours agoJul 29, 2026 • new comments - `NixlPushMode` (WRITE) Roadmap - Reliability Issue Inventory
#48633 commented on
8 hours agoJul 29, 2026 • new comments - [RFC]: Multimodal Asynchronous Collaborative Architecture Sidecar
#49288 commented on
10 hours agoJul 29, 2026 • new comments - [Bug]: EngineDeadError with Kimi-K2.6 model using vLLM 0.20.2
#42363 commented on
11 hours agoJul 29, 2026 • new comments - [Performance]: Remove duplicate cudagraph capture in elastic EP
#42107 commented on
11 hours agoJul 29, 2026 • new comments - [Bug]: DeepSeek-V4-Flash hangs after ~6 requests with cudagraph_mode=FULL_AND_PIECEWISE + chunked prefill on SM 12.x (GB10)
#40969 commented on
12 hours agoJul 29, 2026 • new comments - [Bug] Fix MLPSpeculatorConfig missing num_attention_heads attribute
#34163 commented on
2 days agoJul 28, 2026 • new comments - Support MP backend for elastic EP scale-down
#34075 commented on
last weekJul 24, 2026 • new comments - [Anthropic API] Fix 6 protocol compliance bugs in Messages endpoint
#34053 commented on
last weekJul 24, 2026 • new comments - [Feature][Scheduler] Add split prefix caching feature to eliminate bf16 GEMM tiling divergence across cache-hit/miss paths
#34046 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Fix illegal memory access in AWQ-Marlin with CUDA graphs (Fixes #32834)
#33896 commented on
last weekJul 24, 2026 • new comments - fix: Qwen3ReasoningParser - handle prompt prefix format for Thinking models
#33866 commented on
last weekJul 24, 2026 • new comments - Fix empty content when max_tokens truncates reasoning end token
#33764 commented on
1 hour agoJul 30, 2026 • new comments - [HelionLinearBackend][1/N] Add Helion kernel for scaled_mm
#33651 commented on
2 days agoJul 28, 2026 • new comments - [Frontend] Adding a callback to allow direct access to engine output
#33623 commented on
last weekJul 24, 2026 • new comments - Changing the gating TG group composition
#33608 commented on
last weekJul 24, 2026 • new comments - Fix: Corrected timestamp if audio > 30s
#33564 commented on
last weekJul 24, 2026 • new comments - [Benchmark] Add vllm bench iterations for prefill/decode measurement
#33438 commented on
last weekJul 24, 2026 • new comments - feat(pooling/score): Enhance RerankRequest with optional instruction field and multimodal input support
#33387 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Fix inconsistent embeddings between AsyncLLM and LLM
#33385 commented on
last weekJul 24, 2026 • new comments - [torch.compile] Avoid graph fragmentation in unified_kv_cache_update
#33355 commented on
last weekJul 24, 2026 • new comments - [Frontend] Add structured_outputs field to ResponsesRequest
#33249 commented on
last weekJul 24, 2026 • new comments - Cpu binding to improve GPU performance and utilize idle CPU cores
#33222 commented on
last weekJul 24, 2026 • new comments - [Usage]: DSpark much slower than no-spec on single B300 (DeepSeek-V4-Flash) — config check, or not effective on a saturated batch yet?
#49369 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: GLM-4.1V fails to start at tensor-parallel size 32 — vision tower head-count mismatch
#49368 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: `MooncakeStoreConnector` KV Offload hangs the vLLM engine on a GDN model (`Qwen3.6-35B-A3B`)
#49360 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: MTP speculative decoding is broken with pipeline parallelism (PP>1) — three distinct failures
#49355 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Qwen3.5 / Qwen3.6 hybrid-GDN LoRA: vLLM `enable_lora` does not change generations vs base, while HF Peft does
#49354 commented on
15 hours agoJul 29, 2026 • new comments - Zero JIT compilation during runtime
#49349 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).
#46611 commented on
15 hours agoJul 29, 2026 • new comments - [RFC]: Logprobs/Logits Semantics and Determinism Across the vLLM Ecosystem
#42259 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Qwen3.5-9B answer !!!!!!!!!
#38077 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: RCCL RDNA3 gfx1100 Tp2 ROCM at startup
#38587 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Hang During CUDA Graph Capture on ROCM in 0.19
#39010 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Qwen3.5 Text Only Model (Qwen3_5ForCausalLM)
#39231 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: SubSpec — Lossless Training-Free Speculative Decoding for CPU-Offloaded LLMs via Quantized Substitute Draft
#39427 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Thinking token budget not enforced with MTP speculative decoding (works without MTP)
#39573 commented on
15 hours agoJul 29, 2026 • new comments - [RFC]: Async parallel startup for EngineCore processes in DP/TP scenarios
#39678 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Mistral3 text-only startup fails when text_config.architectures is None
#40318 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: Support full BlockPool state dump as KV events for external router
#40363 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: TurboQuant `_continuation_prefill` OOMs and kills engine at long-context prefill (~185K actual tokens)
#40420 commented on
14 hours agoJul 29, 2026 • new comments - [Performance]: Logprob divergence between vLLM and transformers on a fine-tuned Qwen3.5-VL mobile-use model
#47425 commented on
14 hours agoJul 29, 2026 • new comments - [Bug]: BatchPrefillWithPagedKVCache failed with error invalid argument
#49463 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: StructuredOutputManager sizes grammar-compile ThreadPoolExecutor by host CPU count — cgroup-unaware and uncapped, causing CFS throttling and decode stalls on Kubernetes
#49460 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: V1 streaming-session rebuild leaves stale prefix-cache block hashes and can produce incorrect output
#49449 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: Re-enable `TRITON_UNFUSED` MXFP4 MOE backend
#49446 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: Support for GLM-5.2 NVFP4 on NVIDIA RTX PRO 6000 (SM120 Architecture)
#49421 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: DeepSeek-V4-Flash-DSpark fails to launch with DSpark speculative decoding on SM120 Pro6000D (works fine when disabled)
#49418 commented on
15 hours agoJul 29, 2026 • new comments - [RFC]: KV offload event path refactor — provenance-carrying events and key-only removals
#49413 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Streaming vs non-streaming content whitespace mismatch at tool-call boundaries (ParserEngine: qwen3/nemotron_v3/seed_oss/glm47_moe/gemma4)
#49412 commented on
15 hours agoJul 29, 2026 • new comments - [RFC] [Feature] [Experimental] Support GigaToken Accelerated Tokenizer Mode
#49411 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]: enhance pcp with vllm
#49404 commented on
15 hours agoJul 29, 2026 • new comments - [Docs]: Post-#30201 SharedStorage→ExampleConnector rename still lacks migration docs and stale public examples
#49399 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Qwen3-Omni crashes at startup with "Tensor on device meta" during profile_run when limit_mm_per_prompt disables images
#49384 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Poolside Laguna-S-2.1 poolside_v1 reasoning parser falls back to IdentityReasoningParser and fails auto initialization
#49379 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Token truncation leaves stale Request.block_hashes and can cause incorrect KV cache hits
#49377 commented on
15 hours agoJul 29, 2026 • new comments - [Usage]: DeepSeek-V4-Flash on single B300 — auto-enabled VLLM_USE_BREAKABLE_CUDAGRAPH caps throughput (disabling gives ~1.6x); is it safe, and is torch.compile support planned?
#49370 commented on
15 hours agoJul 29, 2026 • new comments - [Bug]: Wen vllm engine core ready timeout because deepgemm warmup, apiserver exit,but engine core keep running
#32116 commented on
last weekJul 24, 2026 • new comments - [Bug]: MTP speculative decoding produces corrupted output at concurrency >= 4 (V1 engine)
#35288 commented on
last weekJul 24, 2026 • new comments - [Feature]: Someone please upstream this gfx1201/RDNA4 FP8 Patch into vllm-rocm
#28649 commented on
last weekJul 24, 2026 • new comments - [Bug] MiniMax-M2.7 multi-node TP=4: NCCL collective deadlock (all ranks spin at SM~96%/mem=0%/~15W)
#46097 commented on
last weekJul 24, 2026 • new comments - [Bug]: sample_tokens RPC timeout with GLM-5.2-FP8 + DSpark speculative decoding, TP=8 across 2 nodes (Blackwell GB200)
#48752 commented on
last weekJul 24, 2026 • new comments - [Bug]: V1 structured outputs: a malformed grammar request after a valid one crashes EngineCore
#43920 commented on
last weekJul 24, 2026 • new comments - [RFC]: Return extracted hidden states in the generation response
#48743 commented on
last weekJul 24, 2026 • new comments - [Bug]: Native KV offloading with prefix caching changes tool visibility after a long shared-prefix request on Qwen3.6
#49127 commented on
last weekJul 24, 2026 • new comments - [Bug]: GLM-5.2 (DSA sparse MLA) + fp8_ds_mla — sparse indexer off-by-one crashes concurrent decode at max_model_len >= ~325K
#46074 commented on
last weekJul 23, 2026 • new comments - [Bug]: Speculative Decoding + Structured Output(tool call)组合下,decode 阶段出现秒级卡顿
#49002 commented on
last weekJul 23, 2026 • new comments - Recurring CUDA kernel hang on 2x DGX Spark (GB10, sm_12.1) with MiniMax-M2.7-NVFP4, TP=2 across 2 nodes
#41725 commented on
last weekJul 23, 2026 • new comments - [Bug]: Distributed inference hanging on a 2 node DGX spark cluster with Mistral 3.5 Medium 128B with TP=2
#42354 commented on
last weekJul 23, 2026 • new comments - [Tracking Issue]: NIXL P/D Disaggregation for Hybrid Models
#40017 commented on
last weekJul 23, 2026 • new comments - Upgrade to Transformers v5
#38379 commented on
last weekJul 23, 2026 • new comments - [Bug]: Deploying the GLM5.2-nvfp4 model using the 0.24.0 image , the model occasionally outputs "!!!!!!!!!!!!!!!!!" in the thinking phase,
#47367 commented on
last weekJul 23, 2026 • new comments - [Feature]: shuffle safetensor weight files
#42840 commented on
last weekJul 23, 2026 • new comments - [RFC]: Partial Cache Hits for Hybrid Models
#45702 commented on
last weekJul 23, 2026 • new comments - [RFC]: Design a shared warmup infrastructure for JITs
#47456 commented on
5 days agoJul 24, 2026 • new comments - Model Runner V2 Remaining TODOs
#47172 commented on
5 days agoJul 24, 2026 • new comments - [RFC]: Rust front-end
#40846 commented on
5 days agoJul 24, 2026 • new comments - [Roadmap] Minimax M3
#45668 commented on
5 days agoJul 24, 2026 • new comments - [Bug]: runai_safetensors_weights_iterator yields tensors in nondeterministic order, breaking FP8 inference on some platforms
#38991 commented on
5 days agoJul 24, 2026 • new comments - [Bug]: Crash on Transcription (size for tensor a must match the size of tensor b) with reproduce
#39202 commented on
5 days agoJul 24, 2026 • new comments - [Bug] Inkling on sm_121a (GB10): unclamped q-row in rel-bias score-mod gather causes deterministic illegal address (coredump evidence); + aux-stream KV-write race in fused_qkvr_prep
#49049 commented on
5 days agoJul 24, 2026 • new comments - [Bug][Spec Decode] num_speculative_tokens_per_batch_size + MTP speculator fails full CUDA graph decode capture (InputBatch.make_dummy assert)
#48494 commented on
last weekJul 24, 2026 • new comments - [Bug]: vLLM does not support DeepSeek series on RTX PRO 6000/SM120
#26211 commented on
last weekJul 24, 2026 • new comments - [RFC]: vLLM IR: A Functional Intermediate Representation for vLLM
#32358 commented on
last weekJul 24, 2026 • new comments - [Usage]: How to use VLLM to infer the MXFP4A16 model exported by LLM Compressor
#32856 commented on
last weekJul 24, 2026 • new comments - [Usage]: 求助:vllm 在线部署qwen3-vl-Embedding模型,产出结果和离线transformer调用结果不一致是什么原因呢?vllm=0.14.0
#33167 commented on
last weekJul 24, 2026 • new comments - [Enhancement]: Qwen3-ASR realtime endpoint produces degraded output — stateless segments, no cross-segment context, raw format leaks
#35767 commented on
last weekJul 24, 2026 • new comments - [Bug]: Failed to import Triton kernels. Please make sure your triton version is compatible. Error: cannot import name 'SparseMatrix' from 'triton_kernels.tensor'
#36004 commented on
last weekJul 24, 2026 • new comments - [Feature]: fused SiLU + fp8 block quantized kernel in Helion
#36972 commented on
last weekJul 24, 2026 • new comments - [Feature]: Upstream DGX spark improvements from Avarok-Cybersecurity/dgx-vllm
#37141 commented on
last weekJul 24, 2026 • new comments - [RFC]: Active Coordination and Two-Zone Scheduling Mechanism for KV Cache in Long-Running Agents
#37168 commented on
last weekJul 24, 2026 • new comments - [BUG] Port-allocation race between ApiServer processes in hybrid-LB mode (ZMQError: Address already in use)
#40443 commented on
last weekJul 23, 2026 • new comments - [RFC]: Hybrid checkpoint ABI for non-KV prefix resume
#40533 commented on
last weekJul 23, 2026 • new comments - [RFC]: Fault tolerant EPLB
#40567 commented on
last weekJul 23, 2026 • new comments - [Performance]: DeepSeek-V3.2 performance on 8xH20 is not match with official data
#40592 commented on
last weekJul 23, 2026 • new comments - [Bug][ROCm]: NIXL not available logs when using MoRI connector
#40593 commented on
last weekJul 23, 2026 • new comments - [Usage]: Missing vllm:num_requests_running and other metrics in vLLM v0.18 when deploying Qwen3.5 27B
#40605 commented on
last weekJul 23, 2026 • new comments - [vllm IR]: Replace`per_token_group_quant_fp8` with `ir.ops.dynamic_group_quant_fp8`
#40607 commented on
last weekJul 23, 2026 • new comments - [RFC]: Unified Device Capability Abstraction for Cross-Platform Feature Detection
#40620 commented on
last weekJul 23, 2026 • new comments - [Bug]: KeyError on model.layers.N.self_attn.attn during initialize_attn_backend with pipeline_parallel_size=4 (V1 engine + Ray)
#40649 commented on
last weekJul 23, 2026 • new comments - [Bug]: CUBLAS_STATUS_EXECUTION_FAILED during CUDA graph compilation of BF16 vision encoder on NVIDIA Jetson AGX Thor (vLLM 0.19.0 regression)
#40661 commented on
last weekJul 23, 2026 • new comments - [Bug] MultiConnector bypasses the deprecated-signature shim and breaks old-style child connectors
#40690 commented on
last weekJul 23, 2026 • new comments - [Bug]: The size of tensor a (34) must match the size of tensor b (63) at non-singleton dimension 1
#40716 commented on
last weekJul 23, 2026 • new comments - [Model Runner V2][Bug]: The _gumbel_sample_kernel exhibits poor performance on H800.
#40755 commented on
last weekJul 23, 2026 • new comments - [RFC]: Layerwise and Sparse KV cache offloading to support longer sequence length
#48203 commented on
last weekJul 23, 2026 • new comments - [Feature]: Add cache_salt support for Anthropic Messages API
#46688 commented on
last weekJul 23, 2026 • new comments - [Bug]: RuntimeError: UVA is not available
#47292 commented on
last weekJul 23, 2026 • new comments - [Usability] Engine fails to start when available Mamba cache blocks < max_num_seqs (hybrid models + LoRA) — suggest auto-clamping with a warning
#49064 commented on
last weekJul 23, 2026 • new comments - [Feature]: Migration from Model Runner v1 to Model Runner v2
#41286 commented on
last weekJul 23, 2026 • new comments - [Bug]: Qwen3.6-35B-A3B (GDN hybrid) intermittently livelocks under load on GB10/SM121 — GPU 96% util, 0 tok/s, no crash, no Xid
#49203 commented on
last weekJul 23, 2026 • new comments - [Bug]: TP=2 deadlock on dual AMD R9700 (gfx1201/RDNA4) — GPUs spin at 100%, inference blocked
#40980 commented on
last weekJul 23, 2026 • new comments - [Bug][XPU] compressed-tensors FP8 W8A8 (dynamic) generates garbage output on Intel Arc Pro B70 (Battlemage)
#48058 commented on
last weekJul 23, 2026 • new comments - [Performance]: Qwen3.5 native MTP can be slower than no-MTP CUDA graph baseline despite good acceptance
#47277 commented on
last weekJul 23, 2026 • new comments - [CI Failure]: Gemma3 OOMs with transformers backend
#37736 commented on
last weekJul 23, 2026 • new comments - [Bug]: chart-helm does not support configuring shared memory (`/dev/shm`)
#37982 commented on
last weekJul 23, 2026 • new comments - [Feature]: support affinity settings in helm chart
#38308 commented on
last weekJul 23, 2026 • new comments - [Bug]: vLLM fails to start with LMCache + Qwen3-Coder-Next-FP8 (nightly image)
#38700 commented on
last weekJul 23, 2026 • new comments - [Feature]: Use torch.compile Dynamo to see full trace for model forward pass
#39215 commented on
last weekJul 23, 2026 • new comments - [Bug]: Minimax-M3 running crase
#49147 commented on
last weekJul 23, 2026 • new comments - [Bug]: 0.19.0 rocm+7900xtx: Failed to infer device type
#39378 commented on
last weekJul 23, 2026 • new comments - [Bug] Regression: GPTQ models fail to load on Intel XPU in v0.19.0 (missing XPU branches in gptq.py)
#39474 commented on
last weekJul 23, 2026 • new comments - Guided vLLM install mission in KubeStellar Console
#39954 commented on
last weekJul 23, 2026 • new comments - [Bug][ROCm] MI355 + AITER MXFP4 MOE: `Unsupported kernel config for moe heuristic dispatch`
#40008 commented on
last weekJul 23, 2026 • new comments - [Bug]: MTP draft head TP allgather deadlock under sustained long-context load (GLM-5.1-FP8)
#40345 commented on
last weekJul 23, 2026 • new comments - [Feature]: [parity with CUDA] PD disagg recipes on vllm
#40421 commented on
last weekJul 23, 2026 • new comments - [Doc]: vLLM 0.23.0 + Qwen3-VL-8B-Instruct-FP8 on RTX 5080 (Blackwell) - Engine initializes but generate() hangs silently
#46619 commented on
3 days agoJul 27, 2026 • new comments - [bug/perf] V4-Pro hangs ~60 min in post-shard-load weight materialization without --safetensors-load-strategy prefetch on EXT4
#40988 commented on
3 days agoJul 27, 2026 • new comments - [Feature][FP8] Opt-in `ParallelLMHead` quantization in legacy `Fp8Config` (parity with AWQ-Marlin / GPTQ-Marlin / cpu_wna16)
#40999 commented on
3 days agoJul 27, 2026 • new comments - [Doc]: Embed Agent Friendly Code Score Badge
#41001 commented on
3 days agoJul 27, 2026 • new comments - [Feature]: FP8 inference fails on Ampere GPUs (RTX A6000, SM 8.6) due to unsupported default fp8e4nv (E4M3FN) format
#41014 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: AssertionError in sampler.py:383
#41031 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: Incorrect Transport for NIXL in Docker image on cu129
#41048 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: Performance Regression since 0.22 on RDNA2
#45372 commented on
3 days agoJul 27, 2026 • new comments - [Usage]: DeepSeek-V4: --kv-cache-dtype auto fails silently on Blackwell SM120 — should auto-resolve to fp8_ds_mla
#47174 commented on
3 days agoJul 27, 2026 • new comments - [RFC] Weight Reload Correctness for RL
#48312 commented on
3 days agoJul 27, 2026 • new comments - [Doc] Fix broken link, stale permalinks, and missing docstrings
#47886 commented on
3 days agoJul 26, 2026 • new comments - [Bug]: Accuracy drops ~20% when `--enable-prefix-caching` is used together with MTP speculative decoding (Qwen3.6 35B-A3B)
#43559 commented on
3 days agoJul 26, 2026 • new comments - [Bug]: Partial LoRA on Qwen3.5/Qwen3.6 GatedDeltaNet (in_proj_qkv without in_proj_z) crashes in expand_packed_lora — regression from #37912
#47639 commented on
3 days agoJul 26, 2026 • new comments - [Bug]: MTP does not mask embedding on position 0
#33299 commented on
3 days agoJul 26, 2026 • new comments - [Feature][Cleanup]: Unify `vllm.utils.flashinfer` and `vllm.model_executor.layers.quantization.utils.flashinfer_utils`
#31414 commented on
3 days agoJul 26, 2026 • new comments - [Bug]: Regression ; v0.24.0 -> v0.25.1 / v0.26.0 breaks MinimaxM2.7 TP2 on NVIDIA GH200NVL2
#49158 commented on
3 days agoJul 26, 2026 • new comments - [Bug]: MXFP4 MoE Marlin fallback crashes with illegal memory access on Hopper (DeepSeek V4, TP+EP) — non-tile-aligned intermediate size
#47769 commented on
3 days agoJul 26, 2026 • new comments - [Performance]: Official FP8 checkpoint ~40% slower at bs=1 than runtime fp8 quant of same model (Qwen3-4B, RTX 4090, v0.24.0)
#48035 commented on
3 days agoJul 26, 2026 • new comments - [Performance]: Model download at CLI entry: prior art, a first measurement, and a go/no-go protocol
#48194 commented on
3 days agoJul 27, 2026 • new comments - [Roadmap] Cold Start Q3 2026
#48193 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: Qwen1 use_logn_attn may be unsupported in vLLM
#36880 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: [ROCm][gfx1151] Engine Core segfaults in libhsa-runtime64.so when loading Qwen3-VL-32B-AWQ on AMD Ryzen AI MAX+ 395
#37151 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: responses API, combining of message and tool call
#37167 commented on
3 days agoJul 27, 2026 • new comments - [Community] RTX 5090 (Blackwell sm_120) + WSL2 2.7.0: CUDA graphs work — benchmarks + full config
#37242 commented on
3 days agoJul 27, 2026 • new comments - [Feature]: Mamba `DS` conv state layout | Support speculative decoding with `mamba_cache_mode=align`
#38898 commented on
3 days agoJul 27, 2026 • new comments - [Feature]: Built-in request multiplexer to let `vllm serve` use all available GPUs without external proxies
#38923 commented on
3 days agoJul 27, 2026 • new comments - [RFC]: PR de-dup/Similarity-Check CI workflow ?
#39694 commented on
3 days agoJul 27, 2026 • new comments - [RFC]: Automatic test target determination for CI
#39884 commented on
3 days agoJul 27, 2026 • new comments - [RFC]: Unified ModelOpt Quantization in vLLM
#40182 commented on
3 days agoJul 27, 2026 • new comments - [Feature]: Support per-deployment leader election ID and pod-level scoping for LoRA adapters in multi-deployment namespace setups
#40748 commented on
3 days agoJul 27, 2026 • new comments - [Feature]: vllm support audio otel tracing?
#40781 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: RMSNormGated input_guard breaks torch.compile dynamo tracing
#40919 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: profile_cudagraph_memory() ignores GPU memory clamp on sliced GPUs (HAMi/MIG/MPS) — --gpu-memory-utilization is inert with AutoRound INT4 + fp8_e5m2 KV + FlashInfer + CUDA graphs
#40937 commented on
3 days agoJul 27, 2026 • new comments - [Usage]: We are using vLLM version 0.19.1. When attempting to run DeepSeek-V4-Flash with a 32k context window across eight RTX 4090 GPUs, we encountered an error indicating that the `transformers` library needed to be updated. We then updated the library using the command `uv pip install --no-cache-dir git+https://github.com/huggingface/transformers.git`, but the error persisted as shown below:
#40954 commented on
3 days agoJul 27, 2026 • new comments - [Bug]: Enhance KV cache load error handling with detailed error codes / information
#40978 commented on
3 days agoJul 27, 2026 • new comments - [Bug][Tracking Issue]: NaNs in CUDA Graph padding regions corrupt activations in some per-token kernels
#40047 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: For Qwen3.5 serise, Large benchmark gap (~10 points) between offline vLLM inference and vLLM API server despite aligned input_ids
#40699 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Scheduling deadlock in _mamba_block_aligned_split with multiple large multimodal inputs on hybrid Mamba models
#40707 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: vLLM TP=4 corrupts non-English output on PCIe consumer GPUs
#40725 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: runai_streamer loads both Ministral consolidated and HF sharded safetensors
#40765 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Workspace allocation failure when combining Decode Context Parallelism (DCP) with EAGLE3 speculative decoding
#40791 commented on
5 days agoJul 25, 2026 • new comments - [RFC]: [RL - WIP] Checkpoint -> Kernel Format transition graph
#40822 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Qwen3-VL-MoE NVFP4 checkpoint (un-BMM'd per-expert format) fails load with IndexError: Dimension out of range
#40885 commented on
5 days agoJul 25, 2026 • new comments - [RFC]: StructuredOutputManager x Speculative Decoding Refactor
#48197 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: vllm Requests stuck indefinitely
#33099 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Gemma 4 E4B extremely slow on v0.19.0 forced TRITON_ATTN fallback yields ~9 tok/s on RTX 4090 (vs ~100+ tok/s for comparable Llama 3B)
#38887 commented on
5 days agoJul 25, 2026 • new comments - [RFC]: CUDA Checkpoint/Restore for Near-Zero Cold Starts
#34303 commented on
5 days agoJul 25, 2026 • new comments - [Feature]: GLM 5.2 Performance Optimization
#46654 commented on
5 days agoJul 25, 2026 • new comments - [Performance]: DeepSeek-V3.2 Performance Optimization Tracking
#31473 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: vllm 0.23.0 and 0.24.0 - Qwen3.6-35B-A3B-FP8 - Fails generating code- "400 Unterminated string starting at"
#47761 commented on
5 days agoJul 25, 2026 • new comments - [Bug] vLLM >= 0.18.0 NCCL segfault (cuMemCreate) with TP>1 on RTX 4090 (SM 89)
#38967 commented on
5 days agoJul 25, 2026 • new comments - [Bug] ValueError: Following weights were not initialized from checkpoint (Gemma 4 models with KV sharing)
#44788 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Qwen3.5 CUDA Illegal Memory Access in GDN Kernel
#34948 commented on
4 days agoJul 26, 2026 • new comments - [CI Failure]: Multiple tests failing with assert output_size is not None
#43350 commented on
4 days agoJul 26, 2026 • new comments - [Bug][vllm-omni] Qwen3-TTS crashes under concurrent TTS with ref_context_size mismatch
#44933 commented on
4 days agoJul 26, 2026 • new comments - DeepseekV32IndexerBackend requires --block-size 64 — not documented or auto-detected
#48286 commented on
4 days agoJul 26, 2026 • new comments - [Feature]: [Batch-Invariant-Kernels] GDN_ATTN does not support batch-invariant mode for Qwen3.5/Qwen3.6 GDN models
#48613 commented on
4 days agoJul 26, 2026 • new comments - [RFC]: Standalone Rust renderer
#49047 commented on
4 days agoJul 26, 2026 • new comments - [Bug]:MiMo Code mimo_v2_omni.py ERROR
#47864 commented on
4 days agoJul 26, 2026 • new comments - [Bug]: moe_wna16_marlin_gemm applies wrong per-row topk weights (mul_topk_weights=True) at gpt-oss NVFP4 MoE shapes — corrupt output
#48895 commented on
4 days agoJul 25, 2026 • new comments - [RFC]: Session-centric KV-cache orchestration over typed session identity
#48501 commented on
4 days agoJul 25, 2026 • new comments - [Bug]: v0.25.0 regression - Qwen3.5 FP8 on H200 crashes during CUDA graph capture with CUDA illegal memory access
#49311 commented on
4 days agoJul 25, 2026 • new comments - SIGSEGV in torch::jit::invokeOperatorFromPython (transpose) with NVFP4 + DFlash + torch.compile on Blackwell SM120
#48234 commented on
4 days agoJul 25, 2026 • new comments - [SM120][GLM-5.1] NVFP4 DCP/MTP stack tracker
#37113 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: When use_audio_in_video is enabled in qwen3-omni, the output may exhibit issues such as empty or repetitive output.
#38351 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Disaggregate prefill script cannot work due to inconsistent request id between P node and D node.
#38808 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Gemma4 vision encoder crashes with ValueError: Expected hidden_size to be 5376, but found: 72
#39061 commented on
5 days agoJul 25, 2026 • new comments - [Bug]: Gemma 4 MoE (26B-A4B-it) crashes at startup — AssertionError: top_k is None in MoEMixin.recursive_replace
#39066 commented on
5 days agoJul 25, 2026 • new comments - [vLLM IR] Propagate IR op names to torch profiler annotations
#39363 commented on
5 days agoJul 25, 2026 • new comments - [Formatting] Collapse multi-line arg lists where possible
#43449 commented on
2 days agoJul 28, 2026 • new comments - [XPU][MoE] Enable MoE kernel unit tests on Intel XPU
#43407 commented on
4 days agoJul 26, 2026 • new comments - [Kernel] Support UE8M0 scales in fused SiLU block quant
#43399 commented on
last weekJul 23, 2026 • new comments - [Perf][ROCm] MoE int4 weight repacking for GPTQ/AWQ Triton kernel
#43389 commented on
last weekJul 23, 2026 • new comments - Rdt weight sync
#43375 commented on
2 days agoJul 27, 2026 • new comments - [ROCm] Add per-call decode budget to sparse-MLA indexer
#43327 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix][DSV4] MTP draft model: detect BF16 MTP on disk + skip quant_config
#43319 commented on
5 days agoJul 24, 2026 • new comments - [Perf][Core] Use power-of-two-choices for DP engine selection at large DP
#43279 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix][compressed-tensors] Wrap `is_static_input_scheme` with `bool()` for `input_quant=None` schemes
#43248 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Gemma4MoE] Fix AutoRound quantized Gemma4 MoE loading
#43227 commented on
last weekJul 23, 2026 • new comments - [Misc] Add exponential distribution to multi-turn benchmark
#43217 commented on
7 hours agoJul 29, 2026 • new comments - [Feature] Add External LB Mode Support for Elastic EP Scaling
#43202 commented on
20 hours agoJul 29, 2026 • new comments - Refactor CT NVFP4 linear schemes around reusable weight/activation bu…
#43165 commented on
last weekJul 23, 2026 • new comments - [Frontend] Add readiness endpoint and engine health checks
#43122 commented on
2 days agoJul 28, 2026 • new comments - [Core][WIP] Check for GPU<->CPU sync during CI
#43107 commented on
last weekJul 23, 2026 • new comments - [Feature] Triton kernel dispatcher
#43048 commented on
5 days agoJul 24, 2026 • new comments - [ROCm] Cpu offload for ROCm 7.13+ to align the hipMemcpyBatchAsync params
#43018 commented on
4 days agoJul 26, 2026 • new comments - Port MLA attention + quantization fusion to manual fused kernel
#43772 commented on
last weekJul 23, 2026 • new comments - [Bugfix][TurboQuant] Fix CUDA graph capture crash with spec-decode + chunked-prefill (#40807)
#43747 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Quantization] Refuse block-FP8 in MarlinFP8.can_implement
#43722 commented on
last weekJul 23, 2026 • new comments - [Kernel][MoE] Add Sonic-MoE unquantized fused-experts backend
#43686 commented on
last weekJul 23, 2026 • new comments - [Kernel] Warm up hybrid GDN/Mamba/MRoPE kernels
#43642 commented on
7 hours agoJul 29, 2026 • new comments - [Kernel] Add prepared-input fast path with MiniMax-M2 top-k/act quant fusion
#43592 commented on
last weekJul 23, 2026 • new comments - Fix custom allreduce warmup values
#43588 commented on
7 hours agoJul 29, 2026 • new comments - [Feature]Support Sequence Parallel (SP) for qwen3_vl
#43580 commented on
yesterdayJul 28, 2026 • new comments - fix(turboquant): use SDPA prefill fallback on pre-Ampere GPUs
#43577 commented on
last weekJul 23, 2026 • new comments - [Draft][Kernel] Port ActivationQuantFusionPass to manual fusion (#43501)
#43560 commented on
last weekJul 23, 2026 • new comments - [Quant] Manual fusion for static FP8 attention output quant
#43542 commented on
last weekJul 23, 2026 • new comments - [Draft] Migrate bitsandbytes support to OOT plugin
#43529 commented on
7 hours agoJul 29, 2026 • new comments - config: add compile_factors() and refactor compute_hash() as thin wrappers
#43527 commented on
4 days agoJul 26, 2026 • new comments - [Bugfix] Assert KV connector ext_tokens contract in scheduler
#43517 commented on
2 days agoJul 28, 2026 • new comments - [Core] Add prefix-cache-aware routing for internal DP load balancing
#43483 commented on
7 hours agoJul 29, 2026 • new comments - [Bugfix] hadacore_transform: respect inplace parameter to fix garbage outputs with QuIP transforms
#43462 commented on
last weekJul 23, 2026 • new comments - [Offloader] Skip per-tensor scale parameters from CPU offload
#43453 commented on
last weekJul 23, 2026 • new comments - [Perf] Skip blocking GPU->CPU sync of num_accepted_tokens in hybrid+a…
#42574 commented on
yesterdayJul 29, 2026 • new comments - Fix async KV loads counting toward scheduler request limit
#42568 commented on
2 days agoJul 28, 2026 • new comments - integrate nccl2
#42565 commented on
last weekJul 24, 2026 • new comments - [Kernel][DSv4] Add sparse-gather variant of dequantize_and_gather_k_cache
#42504 commented on
last weekJul 23, 2026 • new comments - enable split_group for pytorch process group creation
#42471 commented on
last weekJul 24, 2026 • new comments - [Quant][Prototype] Hoist static FP8 input quant into model code
#42469 commented on
last weekJul 23, 2026 • new comments - Kimi NVFP4 specialized model
#42458 commented on
2 days agoJul 28, 2026 • new comments - Remove unused CT formats and checks
#42450 commented on
last weekJul 23, 2026 • new comments - Support weight-only quant for MX formats
#42447 commented on
last weekJul 23, 2026 • new comments - [Quantization][ModelOpt] W4A16 NVFP4 fused MoE + --override-activation-dtype flag
#42428 commented on
last weekJul 23, 2026 • new comments - [Misc] Move FlashInfer MoE helpers into utils package
#42378 commented on
last weekJul 23, 2026 • new comments - Fix false substring match in INCConfig.get_layer_config
#42367 commented on
last weekJul 23, 2026 • new comments - Create `test_kimi_k2_thinking_nvfp4.py` for accuracy check
#42351 commented on
2 days agoJul 28, 2026 • new comments - [MoE Refactor] Refactor FusedMoEQuantConfig into pre/post-weight loading types.
#42344 commented on
last weekJul 23, 2026 • new comments - [MoE Refactor] Mk construct
#42252 commented on
last weekJul 23, 2026 • new comments - [ROCm] Normalize fp8 scales through float32
#42247 commented on
last weekJul 23, 2026 • new comments - [Kernels] Integration for DeepGEMM MegaMoE kernel
#40833 commented on
last weekJul 23, 2026 • new comments - [Quantization] add hadamard transform + online quantization support to humming
#42997 commented on
last weekJul 23, 2026 • new comments - Resolve silu mul quant padded NaN corruption correctness
#42984 commented on
last weekJul 23, 2026 • new comments - Revert "[Perf] Wire silu_and_mul_per_block_quant into TritonFP8MoE (MiniMax-M2)" (#42497)
#42973 commented on
last weekJul 23, 2026 • new comments - [ModelRunnerV2] Support prompt embeds
#42963 commented on
last weekJul 23, 2026 • new comments - [WIP] Abstract preprocess for k25
#42940 commented on
2 days agoJul 28, 2026 • new comments - [ROCm][DSv4] Functional fixes for DeepSeek V4 on MI300X (gfx942)
#42893 commented on
last weekJul 23, 2026 • new comments - Implement MLA Decode + dynamic fp8 quant Helion kernel and backend
#42886 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Reduce sparse-MLA + indexer workspace bounds on SM_120 (consumer Blackwell)
#42856 commented on
last weekJul 23, 2026 • new comments - [Quantization] Add ModelOpt FP8/NVFP4 weight-only embedding methods
#42791 commented on
last weekJul 23, 2026 • new comments - [MM][CG] Enable encoder CUDA Graph for MiniCPM-V
#42785 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] _triton_mrope_forward: support GPT-J-style rotation (#42016)
#42765 commented on
yesterdayJul 29, 2026 • new comments - [Model][Hardware][AMD][Kernel]: Part 2/2 -> Enable e2e QK Norm + RoPE + KV Cache runtime fusion for Qwen3-30B-A3B on ROCM_ATTN
#42755 commented on
9 hours agoJul 29, 2026 • new comments - [LoRA][Gemma4] Support vision tower LoRA
#42662 commented on
last weekJul 24, 2026 • new comments - [CT] Fix weight processing for FP8 Marlin
#42657 commented on
last weekJul 23, 2026 • new comments - [Dist]support for NCCL’s new shrink and grow features.
#42622 commented on
5 days agoJul 25, 2026 • new comments - [Misc] Upgrade typos from v1.43.5 to v1.46.1 and fix new findings
#42618 commented on
last weekJul 23, 2026 • new comments - [Quant][Prototype] Manual fusion AR+RMS+Quant
#42597 commented on
last weekJul 23, 2026 • new comments - [MyPy] Fix mypy errors in kimi model family (3/5 files)
#45296 commented on
2 days agoJul 28, 2026 • new comments - [Attention] Mamba attention module refactor - Final part
#44857 commented on
3 days agoJul 27, 2026 • new comments - Add SM120 NVFP4 KV cache support
#44851 commented on
4 days agoJul 26, 2026 • new comments - [Core] Enable KimiLinear (KDA/GDN + MLA) PD Separation via NIXL
#44848 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix][DeepSeekV4] Add BF16 MTP O-proj fallback for unquantized draft weights
#44847 commented on
last weekJul 23, 2026 • new comments - [Bugfix][CI] Retry cached HF tokenizer load after transport failures
#44820 commented on
4 days agoJul 25, 2026 • new comments - Revert "[Bugfix][MoE] Snapshot max_cudagraph_capture_size into FusedMoEConfig" (#44613)
#44754 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Gracefully shut down GPU workers when engine-core startup times out or is interrupted (#32116)
#44751 commented on
4 days agoJul 26, 2026 • new comments - [Bugfix][Model] GraniteMoE: load FP8_DYNAMIC expert weight_scale tensors
#44739 commented on
20 hours agoJul 29, 2026 • new comments - [Test][Quantization] Restrict test_machete_mm gate to Hopper family
#44695 commented on
last weekJul 23, 2026 • new comments - Fix RMS quant fusion for mixed dtype fused add RMSNorm
#44693 commented on
last weekJul 23, 2026 • new comments - Whisper : use `<|startofprev|>` instead of `<|prev|>`
#44662 commented on
yesterdayJul 28, 2026 • new comments - [ROCm][Bugfix] Add quantization compatibility guard for Fused Shared Expert in DeepSeek-V2/V3/Kimi-K2
#44651 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix][Quantization] Fp8 family: match modules_to_not_convert by substring
#44628 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Don't crash engine on malformed output-socket frame (#44486)
#44627 commented on
20 hours agoJul 29, 2026 • new comments - [Feature] KVarN: Variance-Normalized KV-Cache Quantization (4-bit K, 2-bit V)
#44581 commented on
last weekJul 23, 2026 • new comments - [MoE Refactor] Combine CompressedTensorsWNA16MarlinMoEMethod with CompressedTensorsWNA16MoEMethod
#44570 commented on
2 days agoJul 28, 2026 • new comments - [KV Cache Quantization] SQuat: Subspace-orthogonal KV Cache Quantization
#44564 commented on
last weekJul 23, 2026 • new comments - [MM][CG] Support ViT full CUDA graph for Ernie-4.5-VL image inference
#45254 commented on
2 days agoJul 27, 2026 • new comments - [V1][Spec Decode] Relaxed acceptance for thinking-phase tokens (port of #22238)
#45229 commented on
5 days agoJul 25, 2026 • new comments - [Core][KV-transfer] MoRIIO: multi-node TP prefill→decode dispatch via published host list
#45228 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix][ROCm][MoE] MoRI: pass num_qp_per_pe/quant_type explicitly; preserve router top-k for finalize
#45225 commented on
last weekJul 23, 2026 • new comments - [Quantization] Fix misleading NVFP4 Marlin MoE fallback warning on GPUs with native FP4
#45192 commented on
last weekJul 23, 2026 • new comments - [Kernel] Fuse per-group FP8 dynamic quant into Triton attention epilogue
#45151 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix] Clean up ModelOpt LM head state before tying weights
#45145 commented on
last weekJul 23, 2026 • new comments - [Kernel] NVIDIA-tuned tile configs + PID swizzling for triton_scaled_mm (up to 1.6x on H800, 2.0x on L20)
#45126 commented on
last weekJul 23, 2026 • new comments - fix: compensate timestamp drift in Whisper chunked transcription
#45115 commented on
2 days agoJul 28, 2026 • new comments - [ROCm] Fix AMD build from shuffle mask dtype error while compiling `silu_and_mul_per_block_quant_kernel`
#45055 commented on
last weekJul 23, 2026 • new comments - llmd+vllm+mori-ep(inter node wide-ep)+mori-io(write) for 2p2d with dp=ep=16 tp=1
#45043 commented on
2 days agoJul 27, 2026 • new comments - [Kernel][FP8][Test] Add CUDA qkv_padded_fp8_quant for ViT FP8
#45001 commented on
last weekJul 23, 2026 • new comments - [Misc] Add KDA scaled-dot KKT test
#44982 commented on
last weekJul 23, 2026 • new comments - [Test][V1] Add sleep/wake correctness regression test for hybrid GDN/…
#44972 commented on
8 hours agoJul 29, 2026 • new comments - [ROCm][CI] Gating more ROCm tests
#44969 commented on
2 days agoJul 28, 2026 • new comments - [Doc] Mention RTX 5090 in CUDA install and INT8 quantization Blackwel…
#44935 commented on
last weekJul 23, 2026 • new comments - fix(structured_output): pass new_token_ids to should_advance() to fix MTP spec-decode off-by-one
#44927 commented on
last weekJul 24, 2026 • new comments - [KV Offload] Async CPU memory pinning for SimpleCPUOffloadConnector
#44127 commented on
yesterdayJul 28, 2026 • new comments - [CPU] Fix simulated multi-NUMA env parsing
#44114 commented on
20 hours agoJul 29, 2026 • new comments - [Bugfix] Quark: set moe_quant_config in QuarkW8A8Int8MoEMethod
#44085 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Expand packed module names in GPTQ modules_in_block_to_quantize
#44072 commented on
last weekJul 23, 2026 • new comments - [XPU] Add PyTorch fallback for per-block SiLU-and-Mul FP8 quantization
#44034 commented on
last weekJul 23, 2026 • new comments - [Kernel][Helion][1/N] Add Helion kernel for silu_and_mul_dynamic_per_token_quant
#44018 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Quantization] Fix OCP MX emulation MoE crash with fp8 activations (w_*_a_fp8)
#43983 commented on
yesterdayJul 28, 2026 • new comments - Kimi K2.5/2.6 LoRA adapter loading
#43952 commented on
2 days agoJul 28, 2026 • new comments - [Manual Fusion][ROCm] Port RMS + Group Quant Fused Op for Qwen3 on ROCm Platform
#43944 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Split mixed reasoning/content streaming deltas
#43938 commented on
3 hours agoJul 29, 2026 • new comments - Support Hybrid&Mamba in kv transfer
#43935 commented on
3 hours agoJul 29, 2026 • new comments - b12x nvfp4 w4a16 use a16 fix
#43929 commented on
last weekJul 23, 2026 • new comments - [ROCm][Perf] DSv3.2: fuse indexer Q-RoPE+quant + K-norm/RoPE/quant/cache
#43907 commented on
last weekJul 23, 2026 • new comments - feat(benchmark): support custom system prompt, request_id tracking, and length filtering controls in bench serve
#43904 commented on
5 days agoJul 24, 2026 • new comments - [PERF] Vedic 4-bit quantized matmul for MLA attention — 3.5x CPU speedup
#43895 commented on
last weekJul 23, 2026 • new comments - Fix MLA kv_b_proj.weight access on quantized ColumnParallelLinear
#43889 commented on
last weekJul 23, 2026 • new comments - [kv_connector] Introduce explicit KVCache offloading connector
#43868 commented on
last weekJul 23, 2026 • new comments - [ROCm][MLA] AITER FP8 ASM prefill backend
#44544 commented on
2 days agoJul 27, 2026 • new comments - [Bugfix] Two independent fixes uncovered while running Quark-quantized MoE checkpoints on ROCm.
#44542 commented on
last weekJul 23, 2026 • new comments - [WIP][Kernel][CuTeDSL] Quant Scaled MM Per (Tensor/token/channel) FP8/INT8 kernel in CuTeDSL
#44501 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix FP8 MoE double memory allocation for non-gated models
#44498 commented on
last weekJul 23, 2026 • new comments - [Metrics][Spec Decoding] Add per-request acceptance rate Prometheus histogram
#44487 commented on
last weekJul 23, 2026 • new comments - [4/N][KV-Cache Layout Refactor] Standardize KV cache layout
#44458 commented on
last weekJul 23, 2026 • new comments - [ROCm][Compile] Fuse RMSNorm + MXFP4 quant via AITER Triton kernels (DeepSeek-R1)
#44437 commented on
last weekJul 23, 2026 • new comments - [Profiler] Get Detailed MultiModal Data Preprocessing Timing Stats
#44404 commented on
14 hours agoJul 29, 2026 • new comments - fix(tracing): dual-emit legacy and current OTel GenAI token attributes
#44373 commented on
2 days agoJul 28, 2026 • new comments - [Perf][Feat] Add generic cuteDSL LL FP32 router (GEMM)
#44343 commented on
20 hours agoJul 29, 2026 • new comments - [Quantization][FP8] Replace bare asserts with descriptive exceptions in fp8.py
#44342 commented on
last weekJul 23, 2026 • new comments - [Misc] Add unit test for fused_gdn_gating_kernel
#44315 commented on
5 days agoJul 24, 2026 • new comments - [MoE Refactor] Add MoEKernelOracle ABC base class and migrate Unquantized oracle
#44231 commented on
last weekJul 23, 2026 • new comments - [Model] Add LocateAnything support
#44182 commented on
2 days agoJul 27, 2026 • new comments - [Core] feat: Widen Request.priority from int to float
#44166 commented on
3 days agoJul 27, 2026 • new comments - fix(internlm2_tool_parser): prevent ValueError when output has multiple tool-call markers
#44140 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix][Frontend] Avoid KeyError in tool_choice validation for tools missing a function key
#44138 commented on
2 hours agoJul 30, 2026 • new comments - [TurboQuant] Add NIAH long-context eval for TurboQuant KV cache
#40122 commented on
5 days agoJul 25, 2026 • new comments - Feature/intern s1 parser
#40115 commented on
5 days agoJul 25, 2026 • new comments - skip tokens do not span a full block in remove_skipped_blocks
#40019 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix level-2 sleep/wake/reload with enable_lora=True
#39935 commented on
yesterdayJul 29, 2026 • new comments - [Docs] Add Whisper Multi-LoRA documentation
#39911 commented on
last weekJul 23, 2026 • new comments - [ROCm] route known-bad gfx9 ROCM_ATTN mfma4 shapes to Triton
#39849 commented on
last weekJul 23, 2026 • new comments - [ROCm][DX] Clarify collect_env ROCm reporting
#39817 commented on
15 hours agoJul 29, 2026 • new comments - [Config] Introduce RuntimeDefault fields for runtime initialized config values with correct static types
#39800 commented on
5 days agoJul 24, 2026 • new comments - [ROCm] Use unified decode fallback for sliding-window AITER FA
#39640 commented on
15 hours agoJul 29, 2026 • new comments - [Bugfix] Fix Triton stream capture error on A100 in GDN attention with MTP speculative decoding
#39483 commented on
last weekJul 23, 2026 • new comments - Fix: Propagate child process startup errors to the frontend
#39345 commented on
15 hours agoJul 29, 2026 • new comments - fix: respect `VLLM_CONFIGURE_LOGGING` in V1 subprocesses
#39152 commented on
10 hours agoJul 29, 2026 • new comments - [EP] Fault tolerance: automatic elastic scale-down on DP engine death
#38862 commented on
15 hours agoJul 29, 2026 • new comments - [Kernel fusion] QK Norm + RoPE + Cache + Quant
#38621 commented on
5 days agoJul 24, 2026 • new comments - [Feature] TRITON_MLA_SPARSE backend for SM8x/11x/12x DSA Sparse MLA Support
#38476 commented on
3 days agoJul 27, 2026 • new comments - [Docs] Add vLLM CI overview documentation for contributors
#38458 commented on
8 hours agoJul 29, 2026 • new comments - fix(tokenizer): skip reasoning_effort when None in Mistral tokenizer
#38448 commented on
15 hours agoJul 29, 2026 • new comments - [Attention][TurboQuant] Optimize k8v4 decode attention with GQA head grouping
#40792 commented on
last weekJul 23, 2026 • new comments - Enable Expert Parallel Load Balancing (EPLB) for Kimi K2.5/2.6 and Marlin Kernel
#40752 commented on
last weekJul 23, 2026 • new comments - [EC Connector] Mooncake EC Connector for Distributed Encoder-Cache Transfer
#40695 commented on
16 hours agoJul 29, 2026 • new comments - [Bugfix] Fix double off-by-one in align_to_block_size in example_connector
#40684 commented on
last weekJul 23, 2026 • new comments - [Bugfix] fix incorrect apply_interleaved_rope in mrope under torch.compile
#40680 commented on
5 days agoJul 25, 2026 • new comments - [Bugfix] Fix load balancer waiting count
#40634 commented on
last weekJul 23, 2026 • new comments - Fix FireRedASR2 hallucination on non-speech audio
#40619 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Gemma-4: Add bnb QuantState alias hook on k_eq_v to load Gemma-4 BNB 4-bit weights
#40606 commented on
last weekJul 23, 2026 • new comments - docs: remove outdated LD_PRELOAD instructions [CPU-Backend]
#40603 commented on
last weekJul 23, 2026 • new comments - fix: correct typo 'Hyrbid' to 'Hybrid' in test-amd.yaml
#40598 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Remove duplicate size_k divisibility check in get_moe_wna16_block_config
#40547 commented on
last weekJul 23, 2026 • new comments - compilation: cap cudagraph sizes for TP Marlin on Ampere
#40385 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Gemma4] Fix vision fp16 overflow causing <pad> output
#40347 commented on
yesterdayJul 28, 2026 • new comments - [Bug] Fix shm_broadcast PyCFunction descriptor corruption under JIT loads
#40303 commented on
last weekJul 23, 2026 • new comments - [ROCm][ViT] Detect Triton-AMD kernels at their new aiter location
#40289 commented on
yesterdayJul 28, 2026 • new comments - Fix DP internal LB first-request routing drift
#40285 commented on
last weekJul 23, 2026 • new comments - [Kernel] Support fused_moe tuning with gemma-4-26B-A4B-it
#40181 commented on
last weekJul 23, 2026 • new comments - Fix MoE models in EP mode on Ascend
#35345 commented on
last weekJul 24, 2026 • new comments - [KVConnector] scheduler: Add HMA support for KV load recovery
#35223 commented on
last weekJul 24, 2026 • new comments - [Bugfix] FlashInfer incompatible with sleep mode
#35124 commented on
2 days agoJul 28, 2026 • new comments - [Bugfix] Separate speculator layers into dedicated KV cache group
#35062 commented on
yesterdayJul 29, 2026 • new comments - Add engine KV checkpoint rewind and lifecycle tools
#35046 commented on
last weekJul 24, 2026 • new comments - feat: add reasoning/thinking support to Anthropic /v1/messages endpoint
#35035 commented on
last weekJul 24, 2026 • new comments - [Feature] Add per-request attention capture to the OpenAI-compatible API
#35014 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Detect missing shard files in quantized checkpoints
#34930 commented on
last weekJul 24, 2026 • new comments - [Build] Add cuda 12.8 release wheel
#34867 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Add is_blackwell_class() for SM121/GB10 DGX Spark support
#34822 commented on
last weekJul 24, 2026 • new comments - [WIP][Attention] Enable masked MHA for topk sparse attention using FA4
#34744 commented on
3 hours agoJul 29, 2026 • new comments - [Draft] amd-quark online fp8 and mxfp4 quantization
#34515 commented on
last weekJul 24, 2026 • new comments - [Benchmark] Support recursive trace discovery and JAX trace formats
#34475 commented on
last weekJul 24, 2026 • new comments - [Bugfix] Fix assertion error in _dummy_run for MTP speculative decoding
#34474 commented on
last weekJul 24, 2026 • new comments - fix(openai): correct timestamp drift for long audio chunking (#32588)
#34461 commented on
last weekJul 24, 2026 • new comments - add log of raw request when crashing
#34221 commented on
last weekJul 24, 2026 • new comments - [Frontend] Add per-request stream interval control with time-based throttling
#34191 commented on
last weekJul 24, 2026 • new comments - [Model Runner v2] E/P/D disaggregation support
#38390 commented on
last weekJul 24, 2026 • new comments - [Feature]: fused RMSNorm + fp8 block quantized kernel in Helion
#38211 commented on
15 hours agoJul 29, 2026 • new comments - [Doc] Update example docs to include Nemotron Super v3 and Nano 4B
#38117 commented on
15 hours agoJul 29, 2026 • new comments - [Frontend] Add /v1/responses/render endpoint and refactor responses preprocessing
#38084 commented on
last weekJul 24, 2026 • new comments - [Feature] RIY: Runtime expert pruning for MoE models
#37824 commented on
2 days agoJul 27, 2026 • new comments - [MoE] Unify MoE oracles with class structure
#37776 commented on
15 hours agoJul 29, 2026 • new comments - [ROCm] Add AITER fused kernel support for DeepSeek MLA attention
#37515 commented on
15 hours agoJul 29, 2026 • new comments - [Bugfix] Sync xgrammar termination state after failed token acceptance
#37506 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix weight transfer tests using stale envs cache - CI test failures
#37482 commented on
5 days agoJul 25, 2026 • new comments - [Bugfix] Lower spec decode match threshold from 66% to 60% to increase chances of test pass on CI
#37477 commented on
5 days agoJul 25, 2026 • new comments - jais: only enable ALiBi when position_embedding_type == "alibi"
#37474 commented on
15 hours agoJul 29, 2026 • new comments - [Responses API] tool_choice support (auto / required / none) for GPT-OSS
#37433 commented on
15 hours agoJul 29, 2026 • new comments - [Doc] Fix inconsistent hash notation in Prefix Caching diagram
#37208 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix] Fix TOCTOU race in KV block allocator causing prefix-cache block theft
#37164 commented on
5 days agoJul 24, 2026 • new comments - [Perf][Utility] Update RMSNorm benchmark script
#36899 commented on
last weekJul 24, 2026 • new comments - [Feature] Add `response_prefix` parameter to audio transcription/translation endpoints
#36018 commented on
5 days agoJul 24, 2026 • new comments - [Core] Add score mode with perplexity and KLD computation
#35961 commented on
yesterdayJul 29, 2026 • new comments - [Bugfix][V1] Warm attention through dummy profiles
#42215 commented on
7 hours agoJul 29, 2026 • new comments - DeepSeekv4 ROCm Optimization
#41601 commented on
last weekJul 23, 2026 • new comments - [ROCm] Fix TurboQuant shape mismatch on non-power-of-2 head_dim
#41597 commented on
last weekJul 23, 2026 • new comments - [Misc] Reduce `time vllm --help` and `time vllm serve --help` to <1s warm
#41518 commented on
last weekJul 23, 2026 • new comments - [ROCm][Deepseekv4] DeepseekV4 Mi300 support
#41451 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Handle NaN in QuantFP8 Native Forward
#41427 commented on
last weekJul 23, 2026 • new comments - [Attention][TurboQuant] Sparse V tile-skip (opt-in)
#41422 commented on
last weekJul 23, 2026 • new comments - [Attention][TurboQuant] Pre-bake Lloyd-Max centroids for common (d, bits) shapes
#41418 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Attention][TurboQuant] Pad head_dim to power-of-2 for WHT
#41414 commented on
last weekJul 23, 2026 • new comments - [torch.compile] typing/debug cleanups: standardise LoadConfig.compute_hash, Path-typed traced_files, richer cache_key_factors.json
#41381 commented on
2 days agoJul 28, 2026 • new comments - Add opt-in FP8 vocab embedding support
#41365 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Prevent stale multiproc RPC deadlines from becoming unbounded waits
#41357 commented on
3 hours agoJul 29, 2026 • new comments - [CVE Backport] Handle `trust_remote_code` for transformers backend (releases/v0.12.0)
#41311 commented on
15 hours agoJul 29, 2026 • new comments - [Bugfix] V1 Ray: sync `global_rank` in `adjust_rank` for PP > 1
#41298 commented on
15 hours agoJul 29, 2026 • new comments - [CI] Wire EAGLE3 acceptance length tests into spec_decode nightly lanes
#41281 commented on
10 hours agoJul 29, 2026 • new comments - [Attention] Remove unused slot mapping from FlashAttention metadata
#41272 commented on
15 hours agoJul 29, 2026 • new comments - [Bugfix][Gemma4] Fix k_eq_v full-attention projection layout
#41253 commented on
15 hours agoJul 29, 2026 • new comments - [Feature]Support Dynamic Chunked Pipeline Parallel with Profiling
#41247 commented on
yesterdayJul 28, 2026 • new comments - [Bugfix] Fix streaming stop handling during tool parsing
#42213 commented on
7 hours agoJul 29, 2026 • new comments - [Bugfix][V1][MoE] Warm up WNA16 MoE Triton kernels
#42193 commented on
7 hours agoJul 29, 2026 • new comments - [Frontend] Add serve warmup for completions, chat, and embeddings
#42101 commented on
11 hours agoJul 29, 2026 • new comments - Inv rope fp8 quant
#42033 commented on
last weekJul 23, 2026 • new comments - [Core] [Quantization] Pass `out_dtype` to linear methods
#42001 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Extend compressed-tensors ignore for Qwen3.5 MTP experts
#41994 commented on
last weekJul 23, 2026 • new comments - fix(model_loader): deterministic suffix mapping for Gemma4 MoE quantization
#41988 commented on
last weekJul 23, 2026 • new comments - fixed kimi k2.5 dflash
#41917 commented on
3 days agoJul 27, 2026 • new comments - [Perf][MoE] RTX 5090 Triton MoE optimizations: tuned config, fused SiLU+Mul+FP8 quant, unrolled topk sum
#41915 commented on
last weekJul 23, 2026 • new comments - [New Model][Nvidia] Add SM12x support for DeepSeek V4 Flash with essential fixes
#41834 commented on
yesterdayJul 29, 2026 • new comments - [Attention][MLA] Add Triton-fused TurboQuant decode backend
#41803 commented on
last weekJul 23, 2026 • new comments - [Bugfix][Model] Fix DeepSeek V4 scale_fmt default for non-canonical quant configs
#41791 commented on
last weekJul 23, 2026 • new comments - [Quantization] Remove duplication in select_mxfp4_moe_backend
#41764 commented on
last weekJul 23, 2026 • new comments - Fix hipErrorCapturedEvent on RoCM when using LoRA with cuda graphs
#41756 commented on
3 days agoJul 27, 2026 • new comments - [MoE Refactor] Optional router_logits argument
#41747 commented on
2 days agoJul 28, 2026 • new comments - [WIP][Model]Add parakeet tdt 0.6b v3 model
#41708 commented on
20 hours agoJul 29, 2026 • new comments - [Bug] Skip KVConnector lookup when request opts out of prefix cache
#41625 commented on
2 days agoJul 28, 2026 • new comments - [Kernel] Add MixFP4 W4A16 Marlin inference support
#41004 commented on
last weekJul 23, 2026 • new comments - [FP8] Add opt-in ParallelLMHead dispatch to Fp8Config
#41000 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix FlashInfer AllReduce benchmark initialization and workspace size
#40970 commented on
3 days agoJul 27, 2026 • new comments - fix(rocm): detect AMD APU and fix VRAM reporting for unified memory
#40963 commented on
3 days agoJul 27, 2026 • new comments - [MoE] Add FlashInfer CUTLASS FP8 ptpc MoE wiring
#40957 commented on
last weekJul 23, 2026 • new comments - [ROCm][CI] Upgrade ROCm quantized MoE coverage
#40943 commented on
5 days agoJul 25, 2026 • new comments - [ROCm][CI] Upgrade quantized FP4 kernels coverage
#40939 commented on
last weekJul 23, 2026 • new comments - [ROCm][CI] Move ROCm AITER quantization tests
#40938 commented on
last weekJul 23, 2026 • new comments - [WIP]Support DeepSeek V4 flash on SM120 with Triton fallback
#40929 commented on
last weekJul 23, 2026 • new comments - [Kernel] Tune default fp8 block-scaled Triton config for M<=8 decode
#40925 commented on
last weekJul 23, 2026 • new comments - [Feature] Add --pre-warm-nccl to reduce first-request TTFT on multi-GPU
#40918 commented on
3 days agoJul 27, 2026 • new comments - [Bugfix][Spec-Decode] TurboQuant K+1 spec-verify routing (fixes #40880)
#40914 commented on
last weekJul 23, 2026 • new comments - [ROCm][DSv4] Share AITER decode dequant + fp8-cast buffers across layers
#40909 commented on
last weekJul 23, 2026 • new comments - fix(gemma4): remap compressed-tensors AWQ MoE keys in _weight_iterator
#40886 commented on
5 days agoJul 25, 2026 • new comments - [NIXL][TURBOQUANT] Support turboquant in NIXL KV connector
#40858 commented on
5 days agoJul 25, 2026 • new comments - [Spec Decode] Inherit online quantization for draft models
#40849 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Fix `NameError` on `AsyncEngineArgs.quantization_config` forward reference
#40847 commented on
last weekJul 23, 2026 • new comments - WIP: b12x updates
#41243 commented on
last weekJul 23, 2026 • new comments - Refactor Dockerfile and update release pipeline configuration
#41231 commented on
15 hours agoJul 29, 2026 • new comments - [Kernel] Add H20-3e FP8 block-scaled GEMM tuned configs for DeepSeek-V4-Flash expert shapes
#41220 commented on
last weekJul 23, 2026 • new comments - Fix sharded_state load for FP8 models with aliased scale keys
#41179 commented on
15 hours agoJul 29, 2026 • new comments - [security] Change VLLM_MEDIA_URL_ALLOW_REDIRECTS default to False
#41167 commented on
15 hours agoJul 29, 2026 • new comments - [DP] External DP LB compatibility with all GPU visibility
#41164 commented on
5 days agoJul 25, 2026 • new comments - First draft of LLM assisted test target determination
#41139 commented on
15 hours agoJul 29, 2026 • new comments - fix(kv-cache): allow TurboQuant on hybrid models
#41123 commented on
last weekJul 23, 2026 • new comments - improve dynamic_per_token_scaled_fp8_quant_kernel_strided kernel
#41115 commented on
last weekJul 23, 2026 • new comments - [ROCm][CI] Extended Fused MoE and FP8 MoE test support
#41100 commented on
last weekJul 23, 2026 • new comments - [Bugfix] Raise clear error when --enable-eplb on non-EPLB model
#41097 commented on
15 hours agoJul 29, 2026 • new comments - [Bugfix] avoid loading errors for dense qwen3_next models
#41087 commented on
15 hours agoJul 29, 2026 • new comments - [Bugfix] Fix GLM45 reasoning token counting
#41077 commented on
5 days agoJul 25, 2026 • new comments - [Bugfix][Quantization] GGUF: sort merged-column slots by index, not stream order
#41057 commented on
last weekJul 23, 2026 • new comments - Fix quant benchmarking script
#41030 commented on
last weekJul 23, 2026 • new comments - [Quantization] Per-shard FP8 scaling for MergedColumnParallelLinear
#41021 commented on
last weekJul 23, 2026 • new comments - Tree decoding segment
#41005 commented on
3 days agoJul 27, 2026 • new comments