vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes * MAX_NUM_BATCHED_TOKENS fix and CUDA tester * Sys kill worker instead of marking as failed * upgrade to vllm 0.12.0 * Update to vllm 0.15.0 and lora fix * Update for HUB and removal of deprected env variables * reverted docker-bake changes * removed leftovers * Update src/handler.py Co-authored-by: Dj Isaac <contact@dejaydev.com> * Update src/utils.py Co-authored-by: Dj Isaac <contact@dejaydev.com> * Update src/handler.py Co-authored-by: Dj Isaac <contact@dejaydev.com> * Clean up of docs and comments in code * nit: lowercase p * nit: lowercase p --------- Co-authored-by: Dj Isaac <contact@dejaydev.com> Co-authored-by: chrisvela <chris.vela@runpod.io>
This commit is contained in:
co-authored by
Dj Isaac
chrisvela
parent
6d6cbe7095
commit
c45ac42acd
+22
-7
@@ -28,7 +28,6 @@ Complete guide to all environment variables and configuration options for worker
|
||||
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
|
||||
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
|
||||
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
|
||||
| `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. |
|
||||
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
|
||||
| `SEED` | 0 | `int` | Random seed for operations. |
|
||||
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
|
||||
@@ -57,6 +56,8 @@ Complete guide to all environment variables and configuration options for worker
|
||||
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
|
||||
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}]` |
|
||||
|
||||
> **Note (Serverless)**: When LoRA adapters are configured via `LORA_MODULES`, initialization is deferred to the first request to ensure compatibility with RunPod Serverless. This means the first request will include LoRA loading time. Subsequent requests are unaffected. Check logs for "LoRA mode: X adapter(s) will load on first request" at startup.
|
||||
|
||||
## Speculative Decoding Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
@@ -85,7 +86,10 @@ Complete guide to all environment variables and configuration options for worker
|
||||
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
||||
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
|
||||
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
|
||||
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models |
|
||||
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models. |
|
||||
| `ATTENTION_BACKEND` | `None` | `str` | Attention backend to use (e.g., `FLASH_ATTN`, `FLASHINFER`, `TRITON_FLASH_ATTN`). Replaces deprecated `VLLM_ATTENTION_BACKEND`. |
|
||||
| `ASYNC_SCHEDULING` | `None` | `bool` | Enable async scheduling (overlaps engine scheduling with GPU execution). Default: enabled in vLLM 0.14.0+. Set to `false` to disable. |
|
||||
| `STREAM_INTERVAL` | `1` | `int` | Controls how often to yield streaming results. Lower = more frequent updates. |
|
||||
|
||||
## Tokenizer Settings
|
||||
|
||||
@@ -115,6 +119,13 @@ The way this works is that the first request will have a batch size of `DEFAULT_
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
|
||||
| `TRUST_REQUEST_CHAT_TEMPLATE` | `false` | `bool` | Allow clients to send custom chat templates in API requests. **Security consideration:** Only enable if you trust your API clients. |
|
||||
| `RETURN_TOKENS_AS_TOKEN_IDS` | `false` | `bool` | Return token IDs instead of decoded text strings in responses. |
|
||||
| `EXCLUDE_TOOLS_WHEN_TOOL_CHOICE_NONE` | `false` | `bool` | Exclude tool definitions from the prompt when `tool_choice` is set to `none`. |
|
||||
| `ENABLE_PROMPT_TOKENS_DETAILS` | `false` | `bool` | Include detailed prompt token information in API responses. |
|
||||
| `ENABLE_FORCE_INCLUDE_USAGE` | `false` | `bool` | Always include usage statistics in API responses, even when not requested. |
|
||||
| `ENABLE_LOG_OUTPUTS` | `false` | `bool` | Log model outputs for debugging purposes. |
|
||||
| `LOG_ERROR_STACK` | `false` | `bool` | Include full stack traces in error responses for debugging. |
|
||||
|
||||
## Serverless & Concurrency Settings
|
||||
|
||||
@@ -122,7 +133,7 @@ The way this works is that the first request will have a batch size of `DEFAULT_
|
||||
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||
| `ENABLE_LOG_REQUESTS` | False | `bool` | Enables vLLM request logging. (Replaces deprecated `DISABLE_LOG_REQUESTS` in vLLM 0.15.0) |
|
||||
|
||||
## Advanced Settings
|
||||
|
||||
@@ -148,7 +159,11 @@ These variables are used when building custom Docker images with models baked in
|
||||
|
||||
⚠️ **The following variables are deprecated and will be removed in future versions:**
|
||||
|
||||
| Old Variable | New Variable | Note |
|
||||
| ---------------------------- | ------------------------ | --------------------- |
|
||||
| `MAX_CONTEXT_LEN_TO_CAPTURE` | `MAX_SEQ_LEN_TO_CAPTURE` | Use new variable name |
|
||||
| `kv_cache_dtype=fp8_e5m2` | `kv_cache_dtype=fp8` | Simplified fp8 format |
|
||||
| Old Variable | New Variable | Note |
|
||||
| ---------------------------- | ------------------------ | -------------------------------------------------------------------- |
|
||||
| `MAX_CONTEXT_LEN_TO_CAPTURE` | `MAX_SEQ_LEN_TO_CAPTURE` | Use new variable name |
|
||||
| `kv_cache_dtype=fp8_e5m2` | `kv_cache_dtype=fp8` | Simplified fp8 format |
|
||||
| `USE_V2_BLOCK_MANAGER` | *(removed)* | V2 block manager is now the default in vLLM 0.13.0, setting ignored |
|
||||
| `VLLM_ATTENTION_BACKEND` | `ATTENTION_BACKEND` | Use new env var name (old still works with deprecation warning) |
|
||||
| `DISABLE_LOG_REQUESTS` | `ENABLE_LOG_REQUESTS` | Inverted logic in vLLM 0.15.0 (old still works with deprecation warning) |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user