0.7.3
Highlights
๐ 253 commits from 93 contributors, including 29 new contributors!
- Deepseek enhancements:
- Support for DeepSeek Multi-Token Prediction, 1.69x speedup in low QPS scenarios (#12755)
- AMD support: DeepSeek tunings, yielding 17% latency reduction (#13199)
- Using FlashAttention3 for MLA (#12807)
- Align the expert selection code path with official implementation (#13474)
- Optimize moe_align_block_size for deepseek_v3 (#12850)
- Expand MLA to support most types of quantization (#13181)
- V1 Engine:
- LoRA Support (#10957, #12883)
- Logprobs and prompt logprobs support (#9880), min_p sampling support (#13191), logit_bias in v1 Sampler (#13079)
- Use msgpack for core request serialization (#12918)
- Pipeline parallelism support (#12996, #13353, #13472, #13417, #13315)
- Metrics enhancements: GPU prefix cache hit rate % gauge (#12592), iteration_tokens_total histogram (#13288), several request timing histograms (#12644)
- Initial speculative decoding support with ngrams (#12193, #13365)
Model Support
- Enhancement to Qwen2.5-VL: BNB support (#12944), LoRA (#13261), Optimizations (#13155)
- Support GPTQModel Dynamic [2,3,4,8]bit GPTQ quantization (#7086)
- Support Unsloth Dynamic 4bit BnB quantization (#12974)
- IBM/NASA Prithvi Geospatial model (#12830)
- Support Mamba2 (Codestral Mamba) (#9292), Bamba Model (#10909)
- Ultravox Model: Support v0.5 Release (#12912)
transformersbackend- Enable quantization support for
transformersbackend (#12960) - Set
torch_dtypeinTransformersModel(#13088)
- Enable quantization support for
- VLM:
- Implement merged multimodal processor for Mllama (#11427), GLM4V (#12449), Molmo (#12966)
- Separate text-only and vision variants of the same model architecture (#13157)
Hardware Support
- Pluggable platform-specific scheduler (#13161)
- NVIDIA: Support nvfp4 quantization (#12784)
- AMD:
- Per-Token-Activation Per-Channel-Weight FP8 (#12501)
- Tuning for Mixtral on MI325 and Qwen MoE on MI300 (#13503), Mixtral8x7B on MI300 (#13577)
- Add intial ROCm support to V1 (#12790)
- TPU: V1 Support (#13049)
- Neuron: Support Longer Sequences in NKI-based Flash PagedAttention and Improve Efficiency (#12921)
- Gaudi:
- Support Contiguous Cache Fetch (#12139)
- Enable long-contexts + LoRA support (#12812)
Engine Feature
- Add sleep and wake up endpoint and v1 support (#12987)
- Add
/v1/audio/transcriptionsOpenAI API endpoint (#12909)
Performance
- Reduce TTFT with concurrent partial prefills (#10235)
- LoRA - Refactor sgmv kernels (#13110)
Others
- Make vLLM compatible with veRL (#12824)
- Fixes for cases of FA2 illegal memory access error (#12848)
- choice-based structured output with xgrammar (#12632)
- Run v1 benchmark and integrate with PyTorch OSS benchmark database (#13068)
What's Changed
- [Misc] Update w2 scale loading for GPTQMarlinMoE by @dsikka in https://github.com/vllm-project/vllm/pull/12757
- [Docs] Add Google Cloud Slides by @simon-mo in https://github.com/vllm-project/vllm/pull/12814
- [Attention] Use FA3 for MLA on Hopper by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/12807
- [misc] Reduce number of config file requests to HuggingFace by @khluu in https://github.com/vllm-project/vllm/pull/12797
- [Misc] Remove unnecessary decode call by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/12833
- [Kernel] Make rotary_embedding ops more flexible with input shape by @Isotr0py in https://github.com/vllm-project/vllm/pull/12777
- [torch.compile] PyTorch 2.6 and nightly compatibility by @youkaichao in https://github.com/vllm-project/vllm/pull/12393
- [Doc] double quote cmake package in build.inc.md by @jitseklomp in https://github.com/vllm-project/vllm/pull/12840
- [Bugfix] Fix unsupported FA version check for Turing GPU by @Isotr0py in https://github.com/vllm-project/vllm/pull/12828
- [V1] LoRA Support by @varun-sundar-rabindranath in https://github.com/vllm-project/vllm/pull/10957
- Add Bamba Model by @fabianlim in https://github.com/vllm-project/vllm/pull/10909
- [MISC] Check space in the file names in the pre commit checks by @houseroad in https://github.com/vllm-project/vllm/pull/12804
- [misc] Revert # 12833 by @khluu in https://github.com/vllm-project/vllm/pull/12857
- [Bugfix] FA2 illegal memory access by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/12848
- Make vllm compatible with verl by @ZSL98 in https://github.com/vllm-project/vllm/pull/12824
- [Bugfix] Missing quant_config in deepseek embedding layer by @SzymonOzog in https://github.com/vllm-project/vllm/pull/12836
- Prevent unecessary requests to huggingface hub by @maxdebayser in https://github.com/vllm-project/vllm/pull/12837
- [MISC][EASY] Break check file names into entry and args in the pre-commit hooks by @houseroad in https://github.com/vllm-project/vllm/pull/12880
- [Misc] Remove unnecessary detokenization in multimodal processing by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/12868
- [Model] Add support for partial rotary embeddings in Phi3 model by @garg-amit in https://github.com/vllm-project/vllm/pull/12718
- [V1] Logprobs and prompt logprobs support by @afeldman-nm in https://github.com/vllm-project/vllm/pull/9880
- [ROCm] [Feature] [Doc] [Dockerfile] [BugFix] Support Per-Token-Activation Per-Channel-Weight FP8 Quantization Inferencing by @tjtanaa in https://github.com/vllm-project/vllm/pull/12501
- [V1] LM Eval With Streaming Integration Tests by @robertgshaw2-redhat in https://github.com/vllm-project/vllm/pull/11590
- [Bugfix] Fix disagg hang caused by the prefill and decode communication issues by @houseroad in https://github.com/vllm-project/vllm/pull/12723
- [V1][Minor] Remove outdated comment by @WoosukKwon in https://github.com/vllm-project/vllm/pull/12928
- [V1] Move KV block hashes from Request to KVCacheManager by @WoosukKwon in https://github.com/vllm-project/vllm/pull/12922
- [Bugfix] Fix Qwen2_5_VLForConditionalGeneration packed_modules_mapping by @jeejeelee in https://github.com/vllm-project/vllm/pull/12905
- [Misc] Fix typo in the example file by @DK-DARKmatter in https://github.com/vllm-project/vllm/pull/12896
- [Bugfix] Fix multi-round chat error when mistral tokenizer is used by @zifeitong in https://github.com/vllm-project/vllm/pull/12859
- [bugfix] respect distributed_executor_backend in world_size=1 by @youkaichao in https://github.com/vllm-project/vllm/pull/12934
- [Misc] Add offline test for disaggregated prefill by @Shaoting-Feng in https://github.com/vllm-project/vllm/pull/12418
- [V1][Minor] Move cascade attn logic outside _prepare_inputs by @WoosukKwon in https://github.com/vllm-project/vllm/pull/12943
- [Build] Make pypi install work on CPU platform by @wangxiyuan in https://github.com/vllm-project/vllm/pull/12874
- [Hardware][Intel-Gaudi] Enable long-contexts + LoRA support for Intel Gaudi by @SanjuCSudhakaran in https://github.com/vllm-project/vllm/pull/12812
- [misc] Add LoRA to benchmark_serving by @varun-sundar-rabindranath in https://github.com/vllm-project/vllm/pull/12898
- [Misc] Log time consumption on weight downloading by @waltforme in https://github.com/vllm-project/vllm/pull/12926
- [CI] Resolve transformers-neuronx version conflict by @liangfu in https://github.com/vllm-project/vllm/pull/12925
- [Doc] Correct HF repository for TeleChat2 models by @waltforme in https://github.com/vllm-project/vllm/pull/12949
- [Misc] Add qwen2.5-vl BNB support by @Isotr0py in https://github.com/vllm-project/vllm/pull/12944
- [CI/Build] Auto-fix Markdown files by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/12941
- [Bugfix] Remove unused seq_group_metadata_list from ModelInputForGPU by @ShangmingCai in https://github.com/vllm-project/vllm/pull/12935
- [bugfix] fix early import of flash attention by @youkaichao in https://github.com/vllm-project/vllm/pull/12959
- [VLM] Merged multi-modal processor for GLM4V by @jeejeelee in https://github.com/vllm-project/vllm/pull/12449
- [V1][Minor] Remove outdated comment by @WoosukKwon in https://github.com/vllm-project/vllm/pull/12968
- [RFC] [Mistral] FP8 format by @patrickvonplaten in https://github.com/vllm-project/vllm/pull/10130
- [V1] Cache
uses_mropein GPUModelRunner by @WoosukKwon in https://github.com/vllm-project/vllm/pull/12969 - [core] port pynvml into vllm codebase by @youkaichao in https://github.com/vllm-project/vllm/pull/12963
- [MISC] Always import version library first in the vllm package by @houseroad in https://github.com/vllm-project/vllm/pull/12979
- [core] improve error handling when wake up from sleep mode by @youkaichao in https://github.com/vllm-project/vllm/pull/12981
- [core][rlhf] add colocate example for RLHF by @youkaichao in https://github.com/vllm-project/vllm/pull/12984
- [V1] Use msgpack for core request serialization by @njhill in https://github.com/vllm-project/vllm/pull/12918
- [Bugfix][Platform] Check whether selected backend is None in get_attn_backend_cls() by @terrytangyuan in https://github.com/vllm-project/vllm/pull/12975
- [core] fix sleep mode and pytorch checkpoint compatibility by @youkaichao in https://github.com/vllm-project/vllm/pull/13001
- [Doc] Add link to tool_choice tracking issue in tool_calling.md by @terrytangyuan in https://github.com/vllm-project/vllm/pull/13003
- [misc] Add retries with exponential backoff for HF file existence check by @khluu in https://github.com/vllm-project/vllm/pull/13008
- [Bugfix] Clean up and fix multi-modal processors by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13012
- Fix seed parameter behavior in vLLM by @SmartManoj in https://github.com/vllm-project/vllm/pull/13007
- [Model] Ultravox Model: Support v0.5 Release by @farzadab in https://github.com/vllm-project/vllm/pull/12912
- [misc] Fix setup.py condition to avoid AMD from being mistaken with CPU by @khluu in https://github.com/vllm-project/vllm/pull/13022
- [V1][Minor] Move scheduler outputs to a separate file by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13062
- [Docs] Annouce Meta Meetup by @simon-mo in https://github.com/vllm-project/vllm/pull/13065
- [Bugfix] Support missing tool parameters in mistral tokenizer by @fgreinacher in https://github.com/vllm-project/vllm/pull/12884
- [Benchmark] Add BurstGPT to benchmark_serving by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13063
- [Core] Don't do platform detection at import time by @russellb in https://github.com/vllm-project/vllm/pull/12933
- [Misc] LoRA - Refactor Punica ops tests by @varun-sundar-rabindranath in https://github.com/vllm-project/vllm/pull/12970
- [Bugfix]: Reasoning output bug according to the chat template change by @gaocegege in https://github.com/vllm-project/vllm/pull/13025
- [V1][Metrics] Add GPU prefix cache hit rate % gauge by @comaniac in https://github.com/vllm-project/vllm/pull/12592
- [executor] init
local_rankas device index by @MengqingCao in https://github.com/vllm-project/vllm/pull/13027 - [ROCm] Using a more precise memory profiling by @gshtras in https://github.com/vllm-project/vllm/pull/12624
- [Build] Fix cuda link target of cumem_allocator in CPU env by @guoyuhong in https://github.com/vllm-project/vllm/pull/12863
- [Platform] add pre_register_and_update function by @wangxiyuan in https://github.com/vllm-project/vllm/pull/12432
- [Bugfix] fix flaky test by @SmartManoj in https://github.com/vllm-project/vllm/pull/13089
- [V1][Metrics] Add several request timing histograms by @markmc in https://github.com/vllm-project/vllm/pull/12644
- Set
torch_dtypeinTransformersModelby @hmellor in https://github.com/vllm-project/vllm/pull/13088 - [Misc] Fix typo at comments at metrics.py by @je1lee in https://github.com/vllm-project/vllm/pull/13024
- [Bugfix] Do not use resource module on Windows (#12858) by @MoonRide303 in https://github.com/vllm-project/vllm/pull/13029
- [BugFix] Pop instead of del CUDA_VISIBLE_DEVICES by @HollowMan6 in https://github.com/vllm-project/vllm/pull/12962
- Fix initializing GGUF weights for ColumnParallelLinear when using tensor parallel > 1 by @SzymonOzog in https://github.com/vllm-project/vllm/pull/13023
- [CI/Build][Bugfix] Fix CPU backend default threads num by @bigPYJ1151 in https://github.com/vllm-project/vllm/pull/13077
- [Doc] Improve OpenVINO installation doc by @hmellor in https://github.com/vllm-project/vllm/pull/13102
- [Bugfix] Guided decoding falls back to outlines when fails to import xgrammar by @terrytangyuan in https://github.com/vllm-project/vllm/pull/12976
- [Misc] Move pre-commit suggestion back to the end by @russellb in https://github.com/vllm-project/vllm/pull/13114
- [RFC][vllm-API] Support tokenizer registry for customized tokenizer in vLLM by @youngkent in https://github.com/vllm-project/vllm/pull/12518
- [Model] IBM/NASA Prithvi Geospatial model by @christian-pinto in https://github.com/vllm-project/vllm/pull/12830
- [ci] Add more source file dependencies for some tests by @khluu in https://github.com/vllm-project/vllm/pull/13123
- [Neuron][Kernel] Support Longer Sequences in NKI-based Flash PagedAttention and Improve Efficiency by @lingfanyu in https://github.com/vllm-project/vllm/pull/12921
- Bump helm/kind-action from 1.10.0 to 1.12.0 by @dependabot in https://github.com/vllm-project/vllm/pull/11612
- Bump actions/stale from 9.0.0 to 9.1.0 by @dependabot in https://github.com/vllm-project/vllm/pull/12462
- Bump helm/chart-testing-action from 2.6.1 to 2.7.0 by @dependabot in https://github.com/vllm-project/vllm/pull/12463
- Bump actions/setup-python from 5.3.0 to 5.4.0 by @dependabot in https://github.com/vllm-project/vllm/pull/12672
- Further reduce the HTTP calls to huggingface.co by @maxdebayser in https://github.com/vllm-project/vllm/pull/13107
- [Misc] AMD Build Improvements by @842974287 in https://github.com/vllm-project/vllm/pull/12923
- [Bug] [V1] Try fetching stop_reason from EngineOutput before checking the request by @bnellnm in https://github.com/vllm-project/vllm/pull/13108
- [Bugfix] Fix num video tokens calculation for Qwen2-VL by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13148
- [Frontend] Generate valid tool call IDs when using
tokenizer-mode=mistralby @rafvasq in https://github.com/vllm-project/vllm/pull/12332 - [Misc] Delete unused LoRA modules by @jeejeelee in https://github.com/vllm-project/vllm/pull/13151
- Introduce VLLM_CUDART_SO_PATH to allow users specify the .so path by @houseroad in https://github.com/vllm-project/vllm/pull/12998
- [CI/Build] Use mypy matcher for pre-commit CI job by @russellb in https://github.com/vllm-project/vllm/pull/13162
- [CORE] [QUANT] Support for GPTQModel's
dynamicquantization per module override/control by @Qubitium in https://github.com/vllm-project/vllm/pull/7086 - [Bugfix] Allow fallback to AWQ from AWQMarlin at per-layer granularity by @mgoin in https://github.com/vllm-project/vllm/pull/13119
- [CI] Fix failing FP8 cpu offload test by @mgoin in https://github.com/vllm-project/vllm/pull/13170
- [V1][Bugfix] Copy encoder input ids to fix set iteration issue during VLM abort by @andoorve in https://github.com/vllm-project/vllm/pull/13173
- [CI/Build] Ignore ruff warning up007 by @russellb in https://github.com/vllm-project/vllm/pull/13182
- [perf-benchmark] cleanup unused Docker images and volumes in H100 benchmark instance by @khluu in https://github.com/vllm-project/vllm/pull/12706
- [NVIDIA] Support nvfp4 quantization by @kaixih in https://github.com/vllm-project/vllm/pull/12784
- [Bugfix][Example] Fix GCed profiling server for TPU by @mgoin in https://github.com/vllm-project/vllm/pull/12792
- [VLM] Implement merged multimodal processor for Mllama by @Isotr0py in https://github.com/vllm-project/vllm/pull/11427
- Simplify logic of locating CUDART so file path by @houseroad in https://github.com/vllm-project/vllm/pull/13203
- [Build] Automatically use the wheel of the base commit with Python-only build by @comaniac in https://github.com/vllm-project/vllm/pull/13178
- [Bugfix] deepseek_r1_reasoning_parser put reason content in wrong field in certain edge case by @LikeSundayLikeRain in https://github.com/vllm-project/vllm/pull/13097
- [Frontend] Move CLI code into vllm.cmd package by @russellb in https://github.com/vllm-project/vllm/pull/12971
- Allow Unsloth Dynamic 4bit BnB quants to work by @danielhanchen in https://github.com/vllm-project/vllm/pull/12974
- [CI/Build] Allow ruff to auto-fix some issues by @russellb in https://github.com/vllm-project/vllm/pull/13180
- [V1][core] Implement pipeline parallel on Ray by @ruisearch42 in https://github.com/vllm-project/vllm/pull/12996
- [VLM] Remove input processor from clip and siglip by @Isotr0py in https://github.com/vllm-project/vllm/pull/13165
- [Frontend] Pass pre-created socket to uvicorn by @russellb in https://github.com/vllm-project/vllm/pull/13113
- [V1] Clarify input processing and multimodal feature caching logic by @ywang96 in https://github.com/vllm-project/vllm/pull/13211
- [VLM] Merged multi-modal processor for Molmo by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/12966
- [V1][Core] Add worker_base for v1 worker by @AoyuQC in https://github.com/vllm-project/vllm/pull/12816
- [Misc] Qwen2.5-VL Optimization by @wulipc in https://github.com/vllm-project/vllm/pull/13155
- [VLM] Separate text-only and vision variants of the same model architecture by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13157
- [Bugfix] Missing Content Type returns 500 Internal Server Error by @vaibhavjainwiz in https://github.com/vllm-project/vllm/pull/13193
- [Frontend] Add
/v1/audio/transcriptionsOpenAI API endpoint by @NickLucche in https://github.com/vllm-project/vllm/pull/12909 - Add label if pre-commit passes by @hmellor in https://github.com/vllm-project/vllm/pull/12527
- Optimize moe_align_block_size for deepseek_v3 by @mgoin in https://github.com/vllm-project/vllm/pull/12850
- [Kernel][Bugfix] Refactor and Fix CUTLASS 2:4 Sparse Kernels by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/13198
- Revert "Add label if pre-commit passes" by @hmellor in https://github.com/vllm-project/vllm/pull/13242
- [ROCm] Avoid using the default stream on ROCm as it is a performance killer by @gshtras in https://github.com/vllm-project/vllm/pull/13238
- [Kernel] Fix awq error when n is not divisable by 128 by @jinzhen-lin in https://github.com/vllm-project/vllm/pull/13227
- [V1] Consolidate MM cache size to vllm.envs by @ywang96 in https://github.com/vllm-project/vllm/pull/13239
- [Bugfix/CI] Turn test_compressed_tensors_2of4_sparse back on by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/13250
- [Bugfix][CI] Inherit codespell settings from pyproject.toml in the pre-commit-config by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/13237
- [Bugfix] Offline example of disaggregated prefill by @XiaobingSuper in https://github.com/vllm-project/vllm/pull/13214
- [Misc] Remove redundant statements in scheduler.py by @WrRan in https://github.com/vllm-project/vllm/pull/13229
- Consolidate Llama model usage in tests by @hmellor in https://github.com/vllm-project/vllm/pull/13094
- Expand MLA to support most types of quantization by @mgoin in https://github.com/vllm-project/vllm/pull/13181
- [V1] LoRA - Enable Serving Usecase by @varun-sundar-rabindranath in https://github.com/vllm-project/vllm/pull/12883
- [ROCm][V1] Add intial ROCm support to V1 by @SageMoore in https://github.com/vllm-project/vllm/pull/12790
- [Bugfix][V1] GPUModelRunner._update_states should return True when there is a finished request in batch by @imkero in https://github.com/vllm-project/vllm/pull/13126
- [WIP] TPU V1 Support Refactored by @alexm-redhat in https://github.com/vllm-project/vllm/pull/13049
- [Frontend] Optionally remove memory buffer used for uploading to URLs in run_batch by @pooyadavoodi in https://github.com/vllm-project/vllm/pull/12927
- [Bugfix] Fix missing parentheses by @xu-song in https://github.com/vllm-project/vllm/pull/13263
- [Misc] Log time consumption of sleep and wake-up by @waltforme in https://github.com/vllm-project/vllm/pull/13115
- [VLM] Keep track of whether prompt replacements have been applied by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13215
- [V1] Simplify GPUModelRunner._update_states check by @njhill in https://github.com/vllm-project/vllm/pull/13265
- Support logit_bias in v1 Sampler by @houseroad in https://github.com/vllm-project/vllm/pull/13079
- [Core] choice-based structured output with xgrammar by @russellb in https://github.com/vllm-project/vllm/pull/12632
- [Hardware][Gaudi][Bugfix] Fix error for guided decoding by @zhouyu5 in https://github.com/vllm-project/vllm/pull/12317
- [Quant][Perf] Use moe_wna16 kernel by default for MoEs with many experts by @mgoin in https://github.com/vllm-project/vllm/pull/13236
- [Core] Reduce TTFT with concurrent partial prefills by @joerunde in https://github.com/vllm-project/vllm/pull/10235
- [V1][Core] min_p sampling support by @AoyuQC in https://github.com/vllm-project/vllm/pull/13191
- [V1][CI] Fix failed v1-test because of min_p by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13316
- [V1][Sampler] Don't apply temp for greedy-only by @njhill in https://github.com/vllm-project/vllm/pull/13311
- [V1][PP] Fix memory profiling in PP by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13315
- [Bugfix][AMD] Update torch_bindings so that scaled_fp4_quant isn't build on ROCm by @SageMoore in https://github.com/vllm-project/vllm/pull/13235
- [Bugfix][Docs] Fix offline Whisper by @NickLucche in https://github.com/vllm-project/vllm/pull/13274
- [Bugfix] Massage MLA's usage of flash attn for RoCM by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/13310
- [BugFix] Don't scan entire cache dir when loading model by @njhill in https://github.com/vllm-project/vllm/pull/13302
- [Bugfix]Fix search start_index of stop_checker by @xu-song in https://github.com/vllm-project/vllm/pull/13280
- [Bugfix] Fix qwen2.5-vl image processor by @Isotr0py in https://github.com/vllm-project/vllm/pull/13286
- [V1][Metrics] Add iteration_tokens_total histogram from V0 by @markmc in https://github.com/vllm-project/vllm/pull/13288
- [AMD] [Model] DeepSeek tunings by @rasmith in https://github.com/vllm-project/vllm/pull/13199
- [V1][PP] Run engine busy loop with batch queue by @comaniac in https://github.com/vllm-project/vllm/pull/13064
- [ci/build] update flashinfer by @youkaichao in https://github.com/vllm-project/vllm/pull/13323
- [Doc] [2/N] Add Fuyu E2E example for multimodal processor by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13331
- [V1][Spec Decode] Ngram Spec Decode by @LiuXiaoxuanPKU in https://github.com/vllm-project/vllm/pull/12193
- [Quant] Add
SupportsQuantto phi3 and clip by @kylesayrs in https://github.com/vllm-project/vllm/pull/13104 - [Bugfix] Pin xgrammar to 0.1.11 by @mgoin in https://github.com/vllm-project/vllm/pull/13338
- [BugFix] Enhance test_pos_encoding to support execution on multi-devices by @wchen61 in https://github.com/vllm-project/vllm/pull/13187
- [V1] Update doc and examples for H2O-VL by @ywang96 in https://github.com/vllm-project/vllm/pull/13349
- [ci] skip failed tests for flashinfer by @youkaichao in https://github.com/vllm-project/vllm/pull/13352
- [platform] add base class for communicators by @youkaichao in https://github.com/vllm-project/vllm/pull/13208
- [Bugfix] Fix 2 Node and Spec Decode tests by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13341
- [Docs] Change myenv to vllm. Update python_env_setup.inc.md by @arkylin in https://github.com/vllm-project/vllm/pull/13325
- [V1][BugFix] Add init.py to v1/spec_decode/ by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13359
- [V1][PP] Cache Intermediate Tensors by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13353
- [Bugfix][Platform][CPU] Fix cuda platform detection on CPU backend edge case by @Isotr0py in https://github.com/vllm-project/vllm/pull/13358
- [V1][BugFix] Clean up rejection sampler & Fix warning msg by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13362
- [V1][Misc] Avoid unnecessary log output by @jeejeelee in https://github.com/vllm-project/vllm/pull/13289
- [Feature][Spec Decode] Simplify the use of Eagle Spec Decode by @ShangmingCai in https://github.com/vllm-project/vllm/pull/12304
- Fix spelling error in index.md by @yankooo in https://github.com/vllm-project/vllm/pull/13369
- Run v1 benchmark and integrate with PyTorch OSS benchmark database by @huydhn in https://github.com/vllm-project/vllm/pull/13068
- [MISC] tiny fixes by @MengqingCao in https://github.com/vllm-project/vllm/pull/13378
- [VLM] Check required fields before initializing field config in
DictEmbeddingItemsby @DarkLight1337 in https://github.com/vllm-project/vllm/pull/13380 - [Model] Support Mamba2 (Codestral Mamba) by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/9292
- [Bugfix] fix xpu communicator by @yma11 in https://github.com/vllm-project/vllm/pull/13368
- [Bugfix] Fix VLLM_USE_MODELSCOPE issue by @r4ntix in https://github.com/vllm-project/vllm/pull/13384
- [V1] Get input tokens from scheduler by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13339
- [V1][PP] Fix intermediate tensor values by @comaniac in https://github.com/vllm-project/vllm/pull/13417
- [V1][Spec decode] Move drafter to model runner by @WoosukKwon in https://github.com/vllm-project/vllm/pull/13363
- [Bugfix][CI][V1] Work around V1 + CUDA Graph + torch._scaled_mm fallback issue by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/13425
- [Misc] Remove dangling references to
SamplingType.BEAMby @hmellor in https://github.com/vllm-project/vllm/pull/13402 - [Model] Enable quantization support for
transformersbackend by @Isotr0py in https://github.com/vllm-project/vllm/pull/12960 - [ROCm] fix get_device_name for rocm by @divakar-amd in https://github.com/v
These notes run past the length kept in the archive. The rest is on the publisherโs page.