Back to Live Feed
vLLM logo

vLLM

Verified
Artificial Intelligence • Global
Official Site

High-throughput, low-latency LLM serving engine powered by PagedAttention.

Tracked Changes
17
Pricing Shifts
0
Features & Launches
14
Last Verified Event
Oct 1, 2026

vLLM Chronological Timeline

2026
featureOct 1, 2026
96% Verified

v0.31.0rc3: [Model Runner V2] Support randomized dummy inputs (#58411)

<p>Signed-off-by: Robert Shaw <a href="mailto:robshaw@redhat.com">robshaw@redhat.com</a><br> Co-authored-by: Robert Shaw <a href="mailto:robshaw@redhat.com">robshaw@redhat.com</a><br> Co-authored-by: Claude Opus 5.5 <a h...

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with v0.31.0rc3: [Model Runner V2] Support randomized dummy inputs (#58411).
View full change record & proof ➔
apiSep 30, 2026
96% Verified

proto-v0.4.0

The vLLM project has released version 0.4.0 of its proto package. This update marks a specific version increment for the protocol-related components within the vLLM ecosystem.

Before: vllm-proto version 0.3.x or earlier.
After: vllm-proto version 0.4.0.
View full change record & proof ➔
featureSep 30, 2026
96% Verified

v0.31.0rc2

<p>[Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse r…</p>...

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with v0.31.0rc2.
View full change record & proof ➔
featureSep 29, 2026
96% Verified

v0.31.0rc1: [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)

The vLLM build process now explicitly skips the snapshot runtime component when targeting CUDA 12.x container images. This change modifies the CI/CD pipeline configuration to optimize build behavior for specific NVIDIA CUDA versions.

Before: The snapshot runtime was included by default in all vLLM build images, including those based on CUDA 12.x.
After: The snapshot runtime is explicitly excluded from the build process for CUDA 12.x images.
View full change record & proof ➔
featureSep 23, 2026
96% Verified

v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)

vLLM has introduced support for MI355 dense NVFP4 and MoRI kernel mirrors within the ROCm CI pipeline. This update enables specific hardware-accelerated precision formats for AMD Instinct MI355 accelerators.

Before: Lack of validated CI support for MI355-specific dense NVFP4 and MoRI kernel configurations.
After: Validated CI support for MI355 dense NVFP4 and MoRI kernel mirrors, enabling optimized inference for these specific formats.
View full change record & proof ➔
featureSep 21, 2026
96% Verified

v0.30.0

The v0.30.0 release updates the build process for CPU-based images. It specifically hardens the fetch mechanism for the triton-cpu sleef submodule.

Before: The triton-cpu sleef submodule fetch process was susceptible to intermittent failures during CPU image builds.
After: The triton-cpu sleef submodule fetch process is hardened to ensure consistent and reliable dependency resolution during CPU image builds.
View full change record & proof ➔
featureSep 21, 2026
96% Verified

v0.30.0: [Build] Fix DeepGEMM CUDA 12.9 release builds (#57554)

This update resolves build failures specifically affecting DeepGEMM integration within the vLLM framework when using CUDA 12.9. It ensures compatibility for users deploying vLLM on the latest NVIDIA CUDA toolkit version.

Before: vLLM builds failed when attempting to compile DeepGEMM components using the CUDA 12.9 toolkit.
After: vLLM builds successfully compile DeepGEMM components when using the CUDA 12.9 toolkit.
View full change record & proof ➔
apiSep 18, 2026
96% Verified

v0.30.0rc2

This release introduces a bugfix for the NIXL component to prevent unnecessary receive reports. It specifically targets notification-only requests to optimize network traffic and processing overhead.

Before: NIXL generated receive reports for all request types, including notification-only requests.
After: NIXL suppresses receive reports for notification-only requests, reducing unnecessary network overhead.
View full change record & proof ➔
featureSep 17, 2026
96% Verified

v0.30.0rc1: [Bugfix] Isolate supplemental FlashInfer BF16 autotuning (#57285)

This release introduces an isolation mechanism for FlashInfer BF16 autotuning processes. It specifically addresses potential conflicts in supplemental tuning configurations during the vLLM execution pipeline.

Before: FlashInfer BF16 autotuning parameters were potentially shared or not fully isolated, leading to configuration conflicts.
After: FlashInfer BF16 autotuning is now isolated, ensuring supplemental tuning configurations do not interfere with primary kernel execution.
View full change record & proof ➔
apiSep 17, 2026
96% Verified

proto-v0.3.0

The vllm-proto package has been updated to version 0.3.0. This release marks a specific iteration of the protocol-related components within the vLLM ecosystem.

Before: vllm-proto version 0.2.x or earlier.
After: vllm-proto version 0.3.0.
View full change record & proof ➔
featureSep 16, 2026
96% Verified

proto-v0.2.0: vllm-proto 0.2.0

The vllm-proto package has been updated to version 0.2.0. This release is formally validated by pull request #56538 and commit fa2a26f.

Before: vllm-proto was at a version prior to 0.2.0 with unverified or legacy protocol definitions.
After: vllm-proto is now at version 0.2.0, verified against the current vLLM codebase via CI pipeline.
View full change record & proof ➔
featureSep 12, 2026
96% Verified

v0.29.1rc0

vLLM v0.29.1rc0 introduces dual-key Gumbel-max watermarking support specifically for speculative decoding workflows. This implementation enables cryptographically verifiable provenance for model outputs generated through speculative execution pipelines.

Before: Speculative decoding pipelines lacked native support for Gumbel-max watermarking, limiting provenance tracking in high-throughput inference.
After: Native support for dual-key Gumbel-max watermarking is now integrated into the speculative decoding execution path.
View full change record & proof ➔
product launchSep 11, 2026
96% Verified

proto-v0.1.0

The vLLM project has released vllm-proto version 0.1.0. This release introduces the initial implementation of the proto package within the vLLM ecosystem.

Before: The vLLM repository lacked a dedicated proto-specific package for standardized protocol definitions.
After: The vllm-proto package is now available at version 0.1.0.
View full change record & proof ➔
featureSep 10, 2026
96% Verified

v0.29.0

Model Runner V2 (MRV2) is now the default execution engine for all models in vLLM. This release introduces CUDA graph memory profiling for KV cache auto-sizing and batch-sharded sampling to reduce per-step logits memory usage.

Before: Model Runner V1 was the default execution engine, with MRV2 limited to specific pooling models.
After: Model Runner V2 is the default engine for all models, featuring batch-sharded sampling, CUDA graph memory profiling, and enhanced speculative decoding support.
View full change record & proof ➔
featureSep 8, 2026
96% Verified

v0.29.0rc6

vLLM v0.29.0rc6 introduces a default dense prefix cache configuration for hybrid model architectures. This update addresses issue #55 to ensure consistent memory management across mixed-model deployments.

Before: Hybrid models lacked a default dense prefix cache configuration, potentially leading to suboptimal KV cache utilization.
After: Hybrid models now utilize dense prefix caching by default, improving memory efficiency and inference performance.
View full change record & proof ➔
featureSep 8, 2026
96% Verified

v0.29.0rc5

<p>[Core] Default prefix_cache_retention_interval to dense for Mamba + E…</p>...

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with v0.29.0rc5.
View full change record & proof ➔
featureSep 4, 2026
96% Verified

v0.29.0rc4: [Bugfix] Avoid sync in TRT-LLM ragged prefill

<p>Generated-by: Codex <a href="mailto:codex@openai.com">codex@openai.com</a></p> <p>Signed-off-by: Codex <a href="mailto:codex@openai.com">codex@openai.com</a></p>...

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with v0.29.0rc4: [Bugfix] Avoid sync in TRT-LLM ragged prefill.
View full change record & proof ➔

Complete vLLM Change Log Index

DateChange TitleTypeImpactDetails
Oct 1, 2026v0.31.0rc3: [Model Runner V2] Support randomized dummy inputs (#58411)feature8/10View ➔
Sep 30, 2026proto-v0.4.0api5/10View ➔
Sep 30, 2026v0.31.0rc2feature8/10View ➔
Sep 29, 2026v0.31.0rc1: [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)feature4/10View ➔
Sep 23, 2026v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)feature8/10View ➔
Sep 21, 2026v0.30.0feature3/10View ➔
Sep 21, 2026v0.30.0: [Build] Fix DeepGEMM CUDA 12.9 release builds (#57554)feature6/10View ➔
Sep 18, 2026v0.30.0rc2api4/10View ➔
Sep 17, 2026v0.30.0rc1: [Bugfix] Isolate supplemental FlashInfer BF16 autotuning (#57285)feature6/10View ➔
Sep 17, 2026proto-v0.3.0api5/10View ➔
Sep 16, 2026proto-v0.2.0: vllm-proto 0.2.0feature6/10View ➔
Sep 12, 2026v0.29.1rc0feature8/10View ➔
Sep 11, 2026proto-v0.1.0product_launch5/10View ➔
Sep 10, 2026v0.29.0feature9/10View ➔
Sep 8, 2026v0.29.0rc6feature7/10View ➔
Sep 8, 2026v0.29.0rc5feature8/10View ➔
Sep 4, 2026v0.29.0rc4: [Bugfix] Avoid sync in TRT-LLM ragged prefillfeature8/10View ➔