vLLM
VerifiedHigh-throughput, low-latency LLM serving engine powered by PagedAttention.
vLLM Chronological Timeline
v0.31.0rc3: [Model Runner V2] Support randomized dummy inputs (#58411)
<p>Signed-off-by: Robert Shaw <a href="mailto:robshaw@redhat.com">robshaw@redhat.com</a><br> Co-authored-by: Robert Shaw <a href="mailto:robshaw@redhat.com">robshaw@redhat.com</a><br> Co-authored-by: Claude Opus 5.5 <a h...
proto-v0.4.0
The vLLM project has released version 0.4.0 of its proto package. This update marks a specific version increment for the protocol-related components within the vLLM ecosystem.
v0.31.0rc2
<p>[Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse r…</p>...
v0.31.0rc1: [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)
The vLLM build process now explicitly skips the snapshot runtime component when targeting CUDA 12.x container images. This change modifies the CI/CD pipeline configuration to optimize build behavior for specific NVIDIA CUDA versions.
v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)
vLLM has introduced support for MI355 dense NVFP4 and MoRI kernel mirrors within the ROCm CI pipeline. This update enables specific hardware-accelerated precision formats for AMD Instinct MI355 accelerators.
v0.30.0
The v0.30.0 release updates the build process for CPU-based images. It specifically hardens the fetch mechanism for the triton-cpu sleef submodule.
v0.30.0: [Build] Fix DeepGEMM CUDA 12.9 release builds (#57554)
This update resolves build failures specifically affecting DeepGEMM integration within the vLLM framework when using CUDA 12.9. It ensures compatibility for users deploying vLLM on the latest NVIDIA CUDA toolkit version.
v0.30.0rc2
This release introduces a bugfix for the NIXL component to prevent unnecessary receive reports. It specifically targets notification-only requests to optimize network traffic and processing overhead.
v0.30.0rc1: [Bugfix] Isolate supplemental FlashInfer BF16 autotuning (#57285)
This release introduces an isolation mechanism for FlashInfer BF16 autotuning processes. It specifically addresses potential conflicts in supplemental tuning configurations during the vLLM execution pipeline.
proto-v0.3.0
The vllm-proto package has been updated to version 0.3.0. This release marks a specific iteration of the protocol-related components within the vLLM ecosystem.
proto-v0.2.0: vllm-proto 0.2.0
The vllm-proto package has been updated to version 0.2.0. This release is formally validated by pull request #56538 and commit fa2a26f.
v0.29.1rc0
vLLM v0.29.1rc0 introduces dual-key Gumbel-max watermarking support specifically for speculative decoding workflows. This implementation enables cryptographically verifiable provenance for model outputs generated through speculative execution pipelines.
proto-v0.1.0
The vLLM project has released vllm-proto version 0.1.0. This release introduces the initial implementation of the proto package within the vLLM ecosystem.
v0.29.0
Model Runner V2 (MRV2) is now the default execution engine for all models in vLLM. This release introduces CUDA graph memory profiling for KV cache auto-sizing and batch-sharded sampling to reduce per-step logits memory usage.
v0.29.0rc6
vLLM v0.29.0rc6 introduces a default dense prefix cache configuration for hybrid model architectures. This update addresses issue #55 to ensure consistent memory management across mixed-model deployments.
v0.29.0rc5
<p>[Core] Default prefix_cache_retention_interval to dense for Mamba + E…</p>...
v0.29.0rc4: [Bugfix] Avoid sync in TRT-LLM ragged prefill
<p>Generated-by: Codex <a href="mailto:codex@openai.com">codex@openai.com</a></p> <p>Signed-off-by: Codex <a href="mailto:codex@openai.com">codex@openai.com</a></p>...
Complete vLLM Change Log Index
| Date | Change Title | Type | Impact | Details |
|---|---|---|---|---|
| Oct 1, 2026 | v0.31.0rc3: [Model Runner V2] Support randomized dummy inputs (#58411) | feature | 8/10 | View ➔ |
| Sep 30, 2026 | proto-v0.4.0 | api | 5/10 | View ➔ |
| Sep 30, 2026 | v0.31.0rc2 | feature | 8/10 | View ➔ |
| Sep 29, 2026 | v0.31.0rc1: [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118) | feature | 4/10 | View ➔ |
| Sep 23, 2026 | v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281) | feature | 8/10 | View ➔ |
| Sep 21, 2026 | v0.30.0 | feature | 3/10 | View ➔ |
| Sep 21, 2026 | v0.30.0: [Build] Fix DeepGEMM CUDA 12.9 release builds (#57554) | feature | 6/10 | View ➔ |
| Sep 18, 2026 | v0.30.0rc2 | api | 4/10 | View ➔ |
| Sep 17, 2026 | v0.30.0rc1: [Bugfix] Isolate supplemental FlashInfer BF16 autotuning (#57285) | feature | 6/10 | View ➔ |
| Sep 17, 2026 | proto-v0.3.0 | api | 5/10 | View ➔ |
| Sep 16, 2026 | proto-v0.2.0: vllm-proto 0.2.0 | feature | 6/10 | View ➔ |
| Sep 12, 2026 | v0.29.1rc0 | feature | 8/10 | View ➔ |
| Sep 11, 2026 | proto-v0.1.0 | product_launch | 5/10 | View ➔ |
| Sep 10, 2026 | v0.29.0 | feature | 9/10 | View ➔ |
| Sep 8, 2026 | v0.29.0rc6 | feature | 7/10 | View ➔ |
| Sep 8, 2026 | v0.29.0rc5 | feature | 8/10 | View ➔ |
| Sep 4, 2026 | v0.29.0rc4: [Bugfix] Avoid sync in TRT-LLM ragged prefill | feature | 8/10 | View ➔ |