Compare commits

...
13 Commits
Author SHA1 Message Date
chrisvelaandGitHub 17efb0e7d0 Merge pull request #272 from runpod-workers/feat/vllm-0.16.0
Release / release (push) Waiting to run
feat: Update to 0.16.0
2026-03-05 13:06:45 -06:00
velaraptor-runpod 2b5f07df63 feat: Update to 0.16.0, remove NUM_GPU_BLOCKS_OVERRIDE in hub default since 0 will break 2026-03-04 16:38:40 -06:00
chrisvelaandGitHub 13fa71878e Merge pull request #269 from runpod-workers/feat/allow-engine-args-env
Release / release (push) Waiting to run
feat: allow all AsyncEngineArgs as env vars
2026-02-27 15:34:23 -06:00
velaraptor-runpod 8a9365bed4 remove DEFAULT_ARGS that are none, fix MAX_CONTEXT_LEN_TO_CAPTURE 2026-02-27 14:04:15 -06:00
velaraptor-runpod cd485a1af1 update readme 2026-02-25 22:57:17 -06:00
velaraptor-runpod b9043639e9 requested changes/refactor 2026-02-25 16:07:38 -06:00
chrisvelaandGitHub 407dbd7773 Merge pull request #270 from runpod-workers/feat/update-vllm-v0.15.1
feat: update vllm to 0.15.1
2026-02-25 15:35:58 -06:00
velaraptor-runpod f103c142c1 feat: update vllm to 0.15.1 2026-02-24 17:44:49 -06:00
velaraptor-runpod efb093e198 add as VLLM_RUNPOD prefix and update readme 2026-02-24 17:37:57 -06:00
velaraptor-runpod 42443f735e feat: allow engine args through VLLM_ and checks the engine args 2026-02-24 16:05:18 -06:00
chrisvelaandGitHub b7c6d4f9a2 feat: update dockerfile to 12.9.1 (#267)
Release / release (push) Waiting to run
* feat: update dockerfile to 12.9.1

* update readme on VLLM_NIGHTLY build arg
2026-02-19 10:13:14 +01:00
chrisvelaandGitHub d69cc021e8 Merge pull request #268 from runpod-workers/fix/spec-config-0-to-none
Release / release (push) Waiting to run
fix: spec config env vars should be none if zero
2026-02-18 15:51:51 -06:00
velaraptor-runpod 61faa8f137 fix: spec config env vars should be none if zero 2026-02-18 15:41:19 -06:00
6 changed files with 252 additions and 135 deletions
+2
View File
@@ -28,6 +28,8 @@ All behaviour is controlled through environment variables:
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String | | `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer | | `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
**Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options.
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md). For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
## API Usage ## API Usage
-9
View File
@@ -280,15 +280,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "NUM_GPU_BLOCKS_OVERRIDE",
"input": {
"name": "Num GPU Blocks Override",
"type": "number",
"description": "If specified, ignore GPU profiling result and use this number of GPU blocks.",
"advanced": true
}
},
{ {
"key": "MAX_NUM_BATCHED_TOKENS", "key": "MAX_NUM_BATCHED_TOKENS",
"input": { "input": {
+4 -4
View File
@@ -1,13 +1,13 @@
FROM nvidia/cuda:12.8.0-base-ubuntu22.04 FROM nvidia/cuda:12.9.1-base-ubuntu22.04
RUN apt-get update -y \ RUN apt-get update -y \
&& apt-get install -y python3-pip && apt-get install -y python3-pip
RUN ldconfig /usr/local/cuda-12.8/compat/ RUN ldconfig /usr/local/cuda-12.9/compat/
# Install vLLM with FlashInfer - use CUDA 12.8 PyTorch wheels (compatible with vLLM 0.15.0) # Install vLLM with FlashInfer - use CUDA 12.8 PyTorch wheels (compatible with vLLM 0.15.1)
RUN python3 -m pip install --upgrade pip && \ RUN python3 -m pip install --upgrade pip && \
python3 -m pip install "vllm[flashinfer]==0.15.0" --extra-index-url https://download.pytorch.org/whl/cu128 python3 -m pip install "vllm[flashinfer]==0.16.0" --extra-index-url https://download.pytorch.org/whl/cu129
+25
View File
@@ -59,6 +59,16 @@ Configure worker-vllm using environment variables:
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String | | `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer | | `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
**Pass any vLLM engine arg** not listed above by setting an environment variable with the **UPPERCASED** field name (same names vLLM uses). The worker auto-discovers all `AsyncEngineArgs` fields from env. For example:
| Environment Variable | vLLM Engine Arg | Example Value |
| ------------------------- | ------------------------ | ------------- |
| `MAX_MODEL_LEN` | `max_model_len` | `4096` |
| `ENFORCE_EAGER` | `enforce_eager` | `true` |
| `ENABLE_CHUNKED_PREFILL` | `enable_chunked_prefill` | `true` |
Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is applied automatically. Backward-compat aliases: `MODEL_NAME`, `TOKENIZER_NAME`, `MAX_CONTEXT_LEN_TO_CAPTURE`. This lets you configure any vLLM option without waiting for explicit worker support.
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)** For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
## Option 2: Build Docker Image with Model Inside ## Option 2: Build Docker Image with Model Inside
@@ -80,6 +90,7 @@ To build an image with the model baked in, you must specify the following docker
- `WORKER_CUDA_VERSION`: `12.1.0` (`12.1.0` is recommended for optimal performance). - `WORKER_CUDA_VERSION`: `12.1.0` (`12.1.0` is recommended for optimal performance).
- `TOKENIZER_NAME`: Tokenizer repository if you would like to use a different tokenizer than the one that comes with the model. (default: `None`, which uses the model's tokenizer) - `TOKENIZER_NAME`: Tokenizer repository if you would like to use a different tokenizer than the one that comes with the model. (default: `None`, which uses the model's tokenizer)
- `TOKENIZER_REVISION`: Tokenizer revision to load (default: `main`). - `TOKENIZER_REVISION`: Tokenizer revision to load (default: `main`).
- `VLLM_NIGHTLY`: Set to `true` to replace the pinned vLLM release with the latest nightly build and the latest `transformers` from source. Useful for testing unreleased vLLM features. (default: `false`)
For the remaining settings, you may apply them as environment variables when running the container. Supported environment variables are listed in the [Environment Variables](#environment-variables) section. For the remaining settings, you may apply them as environment variables when running the container. Supported environment variables are listed in the [Environment Variables](#environment-variables) section.
@@ -89,6 +100,20 @@ For the remaining settings, you may apply them as environment variables when run
docker build -t username/image:tag --build-arg MODEL_NAME="openchat/openchat_3.5" --build-arg BASE_PATH="/models" . docker build -t username/image:tag --build-arg MODEL_NAME="openchat/openchat_3.5" --build-arg BASE_PATH="/models" .
``` ```
### Example: Building with vLLM Nightly
To use the latest unreleased vLLM build (installs from the nightly wheel index and `transformers` from source):
```bash
docker build -t username/image:tag --build-arg VLLM_NIGHTLY=true .
```
You can combine it with other arguments:
```bash
docker build -t username/image:tag --build-arg VLLM_NIGHTLY=true --build-arg MODEL_NAME="meta-llama/Llama-3.1-8B-Instruct" --build-arg BASE_PATH="/models" .
```
### (Optional) Including Huggingface Token ### (Optional) Including Huggingface Token
If the model you would like to deploy is private or gated, you will need to include it during build time as a Docker secret, which will protect it from being exposed in the image and on DockerHub. If the model you would like to deploy is private or gated, you will need to include it during build time as a Docker secret, which will protect it from being exposed in the image and on DockerHub.
+23
View File
@@ -156,6 +156,29 @@ The way this works is that the first request will have a batch size of `DEFAULT_
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. | | `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. | | `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
## UPPERCASED env vars: Pass any engine arg
Any vLLM `AsyncEngineArgs` field can be set via an environment variable using the **UPPERCASED** field name (the same names vLLM uses). The worker auto-discovers all fields from env — no prefix.
**Format:** `<FIELD_NAME_UPPERCASED>=<value>` (e.g. `MAX_MODEL_LEN=4096`)
**Examples:**
| Environment Variable | vLLM Engine Arg | Value Example |
| ------------------------ | ------------------------ | ------------- |
| `MAX_MODEL_LEN` | `max_model_len` | `4096` |
| `ENFORCE_EAGER` | `enforce_eager` | `true` |
| `ENABLE_CHUNKED_PREFILL` | `enable_chunked_prefill` | `true` |
| `NUM_SCHEDULER_STEPS` | `num_scheduler_steps` | `8` |
| `TOKENIZER_POOL_SIZE` | `tokenizer_pool_size` | `4` |
**Backward-compat aliases:** `MODEL_NAME` → `model`, `TOKENIZER_NAME` → `tokenizer`, `MAX_CONTEXT_LEN_TO_CAPTURE` → `max_seq_len_to_capture`, `MODEL_REVISION` → `revision`.
**Notes:**
- Only valid `AsyncEngineArgs` fields are applied. Unknown keys are silently ignored.
- Values are automatically cast to the correct type (`int`, `float`, `bool`, `str`, or JSON for `dict`/`list`/`tuple`).
- For a full list of available engine args, see the [vLLM AsyncEngineArgs documentation](https://docs.vllm.ai/en/latest/configuration/engine_args/).
## Docker Build Arguments ## Docker Build Arguments
These variables are used when building custom Docker images with models baked in: These variables are used when building custom Docker images with models baked in:
+195 -119
View File
@@ -1,106 +1,169 @@
import os import os
import json import json
import logging import logging
from typing import get_origin, get_args
from torch.cuda import device_count from torch.cuda import device_count
from vllm import AsyncEngineArgs from vllm import AsyncEngineArgs
from vllm.model_executor.model_loader.tensorizer import TensorizerConfig from vllm.model_executor.model_loader.tensorizer import TensorizerConfig
from src.utils import convert_limit_mm_per_prompt from src.utils import convert_limit_mm_per_prompt
RENAME_ARGS_MAP = { # Backward-compat: env var names users already know → engine arg name
ENV_ALIASES = {
"MODEL_NAME": "model", "MODEL_NAME": "model",
"MODEL_REVISION": "revision", "MODEL_REVISION": "revision",
"TOKENIZER_NAME": "tokenizer", "TOKENIZER_NAME": "tokenizer",
"MAX_CONTEXT_LEN_TO_CAPTURE": "max_seq_len_to_capture"
} }
# Literal defaults from original worker (used when env/local do not set a value)
DEFAULT_ARGS = { DEFAULT_ARGS = {
"disable_log_stats": os.getenv('DISABLE_LOG_STATS', 'False').lower() == 'true', "disable_log_stats": False,
# disable_log_requests is deprecated, use enable_log_requests instead "enable_log_requests": False,
"enable_log_requests": os.getenv('ENABLE_LOG_REQUESTS', 'False').lower() == 'true', "gpu_memory_utilization": 0.95,
"gpu_memory_utilization": float(os.getenv('GPU_MEMORY_UTILIZATION', 0.95)), "pipeline_parallel_size": 1,
"pipeline_parallel_size": int(os.getenv('PIPELINE_PARALLEL_SIZE', 1)), "tensor_parallel_size": 1,
"tensor_parallel_size": int(os.getenv('TENSOR_PARALLEL_SIZE', 1)), "skip_tokenizer_init": False,
"served_model_name": os.getenv('SERVED_MODEL_NAME', None), "tokenizer_mode": "auto",
"tokenizer": os.getenv('TOKENIZER', None), "trust_remote_code": False,
"skip_tokenizer_init": os.getenv('SKIP_TOKENIZER_INIT', 'False').lower() == 'true', "load_format": "auto",
"tokenizer_mode": os.getenv('TOKENIZER_MODE', 'auto'), "dtype": "auto",
"trust_remote_code": os.getenv('TRUST_REMOTE_CODE', 'False').lower() == 'true', "kv_cache_dtype": "auto",
"download_dir": os.getenv('DOWNLOAD_DIR', None), "seed": 0,
"load_format": os.getenv('LOAD_FORMAT', 'auto'), "worker_use_ray": False,
"config_format": os.getenv('CONFIG_FORMAT', 'auto'), "block_size": 16,
"dtype": os.getenv('DTYPE', 'auto'), "enable_prefix_caching": False,
"kv_cache_dtype": os.getenv('KV_CACHE_DTYPE', 'auto'), "disable_sliding_window": False,
"quantization_param_path": os.getenv('QUANTIZATION_PARAM_PATH', None), "swap_space": 4,
"seed": int(os.getenv('SEED', 0)), "cpu_offload_gb": 0,
"max_model_len": int(os.getenv('MAX_MODEL_LEN', 0)) or None, "max_num_seqs": 256,
"worker_use_ray": os.getenv('WORKER_USE_RAY', 'False').lower() == 'true', "max_logprobs": 20,
"distributed_executor_backend": os.getenv('DISTRIBUTED_EXECUTOR_BACKEND', None), "enforce_eager": False,
"max_parallel_loading_workers": int(os.getenv('MAX_PARALLEL_LOADING_WORKERS', 0)) or None, "max_seq_len_to_capture": 8192,
"block_size": int(os.getenv('BLOCK_SIZE', 16)), "disable_custom_all_reduce": False,
"enable_prefix_caching": os.getenv('ENABLE_PREFIX_CACHING', 'False').lower() == 'true', "tokenizer_pool_size": 0,
"disable_sliding_window": os.getenv('DISABLE_SLIDING_WINDOW', 'False').lower() == 'true', "tokenizer_pool_type": "ray",
# attention_backend replaces deprecated VLLM_ATTENTION_BACKEND env var "enable_lora": False,
"attention_backend": os.getenv('ATTENTION_BACKEND', None), "max_loras": 1,
# Enabled by default for improved throughput. Set to False to disable if experiencing issues "max_lora_rank": 16,
"async_scheduling": None if os.getenv('ASYNC_SCHEDULING') is None else os.getenv('ASYNC_SCHEDULING', 'True').lower() == 'true', "enable_prompt_adapter": False,
# Controls how often to yield streaming results "max_prompt_adapters": 1,
"stream_interval": int(os.getenv('STREAM_INTERVAL', 1)), "max_prompt_adapter_token": 0,
"swap_space": int(os.getenv('SWAP_SPACE', 4)), # GiB "fully_sharded_loras": False,
"cpu_offload_gb": int(os.getenv('CPU_OFFLOAD_GB', 0)), # GiB "lora_extra_vocab_size": 256,
# vLLM defaults None to 2048; keep 0 as None to let vLLM auto-calculate "lora_dtype": "auto",
"max_num_batched_tokens": int(os.getenv('MAX_NUM_BATCHED_TOKENS', 0)) or None, "device": "auto",
"max_num_seqs": int(os.getenv('MAX_NUM_SEQS', 256)), "ray_workers_use_nsight": False,
"max_logprobs": int(os.getenv('MAX_LOGPROBS', 20)), # Default value for OpenAI Chat Completions API "num_lookahead_slots": 0,
"revision": os.getenv('REVISION', None), "scheduler_delay_factor": 0.0,
"code_revision": os.getenv('CODE_REVISION', None), "guided_decoding_backend": "outlines",
"rope_scaling": os.getenv('ROPE_SCALING', None), "spec_decoding_acceptance_method": "rejection_sampler",
"rope_theta": float(os.getenv('ROPE_THETA', 0)) or None, "stream_interval": 1,
"tokenizer_revision": os.getenv('TOKENIZER_REVISION', None),
"quantization": os.getenv('QUANTIZATION', None),
"enforce_eager": os.getenv('ENFORCE_EAGER', 'False').lower() == 'true',
"max_context_len_to_capture": int(os.getenv('MAX_CONTEXT_LEN_TO_CAPTURE', 0)) or None,
"max_seq_len_to_capture": int(os.getenv('MAX_SEQ_LEN_TO_CAPTURE', 8192)),
"disable_custom_all_reduce": os.getenv('DISABLE_CUSTOM_ALL_REDUCE', 'False').lower() == 'true',
"tokenizer_pool_size": int(os.getenv('TOKENIZER_POOL_SIZE', 0)),
"tokenizer_pool_type": os.getenv('TOKENIZER_POOL_TYPE', 'ray'),
"tokenizer_pool_extra_config": os.getenv('TOKENIZER_POOL_EXTRA_CONFIG', None),
"enable_lora": os.getenv('ENABLE_LORA', 'False').lower() == 'true',
"max_loras": int(os.getenv('MAX_LORAS', 1)),
"max_lora_rank": int(os.getenv('MAX_LORA_RANK', 16)),
"enable_prompt_adapter": os.getenv('ENABLE_PROMPT_ADAPTER', 'False').lower() == 'true',
"max_prompt_adapters": int(os.getenv('MAX_PROMPT_ADAPTERS', 1)),
"max_prompt_adapter_token": int(os.getenv('MAX_PROMPT_ADAPTER_TOKEN', 0)),
"fully_sharded_loras": os.getenv('FULLY_SHARDED_LORAS', 'False').lower() == 'true',
"lora_extra_vocab_size": int(os.getenv('LORA_EXTRA_VOCAB_SIZE', 256)),
"long_lora_scaling_factors": tuple(map(float, os.getenv('LONG_LORA_SCALING_FACTORS', '').split(','))) if os.getenv('LONG_LORA_SCALING_FACTORS') else None,
"lora_dtype": os.getenv('LORA_DTYPE', 'auto'),
"max_cpu_loras": int(os.getenv('MAX_CPU_LORAS', 0)) or None,
"device": os.getenv('DEVICE', 'auto'),
"ray_workers_use_nsight": os.getenv('RAY_WORKERS_USE_NSIGHT', 'False').lower() == 'true',
"num_gpu_blocks_override": int(os.getenv('NUM_GPU_BLOCKS_OVERRIDE', 0)) or None,
"num_lookahead_slots": int(os.getenv('NUM_LOOKAHEAD_SLOTS', 0)),
"model_loader_extra_config": os.getenv('MODEL_LOADER_EXTRA_CONFIG', None),
"ignore_patterns": os.getenv('IGNORE_PATTERNS', None),
"preemption_mode": os.getenv('PREEMPTION_MODE', None),
"scheduler_delay_factor": float(os.getenv('SCHEDULER_DELAY_FACTOR', 0.0)),
"enable_chunked_prefill": os.getenv('ENABLE_CHUNKED_PREFILL', None),
"guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'),
"speculative_model": os.getenv('SPECULATIVE_MODEL', None),
"speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None,
"enable_expert_parallel": bool(os.getenv('ENABLE_EXPERT_PARALLEL', 'False').lower() == 'true'),
"num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None,
"speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None,
"speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None,
"ngram_prompt_lookup_max": int(os.getenv('NGRAM_PROMPT_LOOKUP_MAX', 0)) or None,
"ngram_prompt_lookup_min": int(os.getenv('NGRAM_PROMPT_LOOKUP_MIN', 0)) or None,
"spec_decoding_acceptance_method": os.getenv('SPEC_DECODING_ACCEPTANCE_METHOD', 'rejection_sampler'),
"typical_acceptance_sampler_posterior_threshold": float(os.getenv('TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD', 0)) or None,
"typical_acceptance_sampler_posterior_alpha": float(os.getenv('TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA', 0)) or None,
"qlora_adapter_name_or_path": os.getenv('QLORA_ADAPTER_NAME_OR_PATH', None),
"disable_logprobs_during_spec_decoding": os.getenv('DISABLE_LOGPROBS_DURING_SPEC_DECODING', None),
"otlp_traces_endpoint": os.getenv('OTLP_TRACES_ENDPOINT', None),
} }
def _resolve_field_type(field_type: type) -> type:
"""Resolve Optional/Union to the concrete type for conversion."""
origin = get_origin(field_type)
args = get_args(field_type) if hasattr(field_type, "__args__") else ()
if origin is not None:
# Optional[X] is Union[X, None]; X | None is UnionType
non_none = [a for a in args if a is not type(None)]
if non_none:
return non_none[0]
return field_type
def _convert_env_value_to_field_type(value: str, field_name: str, field_type: type):
"""Convert env var string to the type expected by AsyncEngineArgs for this field."""
val = value.strip() if isinstance(value, str) else value
if val in ("", "None", "none"):
args = get_args(field_type) if hasattr(field_type, "__args__") else ()
if type(None) in (args or ()):
return None
raise ValueError("empty value not allowed for non-optional field")
effective_type = _resolve_field_type(field_type)
# bool
if effective_type is bool:
return str(val).lower() in ("true", "1", "yes", "on")
# int
if effective_type is int:
return int(val)
# float
if effective_type is float:
return float(val)
# str
if effective_type is str:
return str(val)
# dict, list, or complex (try JSON)
origin = get_origin(effective_type)
if effective_type in (dict, list) or origin in (dict, list):
try:
return json.loads(val)
except json.JSONDecodeError:
return val
# tuple (e.g. long_lora_scaling_factors) — comma-separated or JSON array
if effective_type is tuple or origin is tuple:
args = get_args(field_type) if hasattr(field_type, "__args__") else ()
elem_types = [a for a in args if a is not Ellipsis]
elem_type = elem_types[0] if elem_types else str
try:
parsed = json.loads(val)
if isinstance(parsed, list):
return tuple(elem_type(x) for x in parsed)
except (json.JSONDecodeError, TypeError):
pass
return tuple(elem_type(x.strip()) for x in str(val).split(",") if x.strip())
# Fallback: try int, float, then str
try:
return int(val)
except ValueError:
pass
try:
return float(val)
except ValueError:
pass
return str(val)
def _get_args_from_env_auto_discover() -> dict:
"""Auto-discover engine args from env vars using UPPERCASED field names.
For every field in AsyncEngineArgs, check os.getenv(FIELD_NAME).
E.g. MAX_MODEL_LEN=4096 -> max_model_len=4096.
Uses same type conversion as before; supports all vLLM engine args without manual listing.
"""
args = {}
valid_fields = AsyncEngineArgs.__dataclass_fields__
for field_name, field in valid_fields.items():
env_key = field_name.upper()
value = os.environ.get(env_key)
if value is None:
continue
try:
args[field_name] = _convert_env_value_to_field_type(
value, field_name, field.type
)
except (ValueError, TypeError, json.JSONDecodeError) as e:
logging.warning(
"Skip env %s=%r: %s", env_key, value, e
)
return args
def _apply_env_aliases(args: dict) -> None:
"""Apply ENV_ALIASES: if MODEL_NAME etc. are set, set the target engine arg."""
valid_fields = AsyncEngineArgs.__dataclass_fields__
for alias, target in ENV_ALIASES.items():
value = os.environ.get(alias)
if value is None or target not in valid_fields:
continue
try:
args[target] = _convert_env_value_to_field_type(
value, target, valid_fields[target].type
)
except (ValueError, TypeError, json.JSONDecodeError) as e:
logging.warning("Skip env alias %s=%r: %s", alias, value, e)
def get_speculative_config(): def get_speculative_config():
"""Build speculative decoding configuration from environment variables. """Build speculative decoding configuration from environment variables.
@@ -122,9 +185,14 @@ def get_speculative_config():
# Option 2: Build config from individual environment variables # Option 2: Build config from individual environment variables
spec_method = os.getenv('SPECULATIVE_METHOD') spec_method = os.getenv('SPECULATIVE_METHOD')
spec_model = os.getenv('SPECULATIVE_MODEL') spec_model = os.getenv('SPECULATIVE_MODEL')
num_spec_tokens = os.getenv('NUM_SPECULATIVE_TOKENS') _num_spec_tokens = os.getenv('NUM_SPECULATIVE_TOKENS')
ngram_max = os.getenv('NGRAM_PROMPT_LOOKUP_MAX') _ngram_max = os.getenv('NGRAM_PROMPT_LOOKUP_MAX')
ngram_min = os.getenv('NGRAM_PROMPT_LOOKUP_MIN') _ngram_min = os.getenv('NGRAM_PROMPT_LOOKUP_MIN')
# Convert numeric vars to int so '0' (hub.json default) is treated as unset
num_spec_tokens = (int(_num_spec_tokens) or None) if _num_spec_tokens else None
ngram_max = (int(_ngram_max) or None) if _ngram_max else None
ngram_min = (int(_ngram_min) or None) if _ngram_min else None
if not any([spec_method, spec_model, ngram_max]): if not any([spec_method, spec_model, ngram_max]):
return None return None
@@ -150,11 +218,11 @@ def get_speculative_config():
if spec_model: if spec_model:
config['model'] = spec_model config['model'] = spec_model
if num_spec_tokens: if num_spec_tokens:
config['num_speculative_tokens'] = int(num_spec_tokens) config['num_speculative_tokens'] = num_spec_tokens
if ngram_max: if ngram_max:
config['prompt_lookup_max'] = int(ngram_max) config['prompt_lookup_max'] = ngram_max
if ngram_min: if ngram_min:
config['prompt_lookup_min'] = int(ngram_min) config['prompt_lookup_min'] = ngram_min
draft_tp = os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE') draft_tp = os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE')
if draft_tp: if draft_tp:
@@ -186,6 +254,7 @@ def get_speculative_config():
return None return None
def _resolve_max_model_len(model, trust_remote_code=False, revision=None): def _resolve_max_model_len(model, trust_remote_code=False, revision=None):
"""Resolve max_model_len from the model's HuggingFace config.""" """Resolve max_model_len from the model's HuggingFace config."""
try: try:
@@ -204,25 +273,19 @@ def _resolve_max_model_len(model, trust_remote_code=False, revision=None):
logging.warning(f"Could not resolve max_model_len from model config: {e}") logging.warning(f"Could not resolve max_model_len from model config: {e}")
return None return None
limit_mm_env = os.getenv('LIMIT_MM_PER_PROMPT')
if limit_mm_env is not None:
DEFAULT_ARGS["limit_mm_per_prompt"] = convert_limit_mm_per_prompt(limit_mm_env)
def match_vllm_args(args): def _local_args_to_engine_args(local: dict) -> dict:
"""Rename args to match vllm by: """Map local args (e.g. from /local_model_args.json) to engine arg names and filter."""
1. Renaming keys to lower case valid = AsyncEngineArgs.__dataclass_fields__
2. Renaming keys to match vllm out = {}
3. Filtering args to match vllm's AsyncEngineArgs for k, v in local.items():
target = ENV_ALIASES.get(k, k.lower().replace("-", "_"))
if target not in valid or v in (None, "", "None"):
continue
out[target] = v
return out
Args:
args (dict): Dictionary of args
Returns:
dict: Dictionary of args with renamed keys
"""
renamed_args = {RENAME_ARGS_MAP.get(k, k): v for k, v in args.items()}
matched_args = {k: v for k, v in renamed_args.items() if k in AsyncEngineArgs.__dataclass_fields__}
return {k: v for k, v in matched_args.items() if v not in [None, "", "None"]}
def get_local_args(): def get_local_args():
""" """
Retrieve local arguments from a JSON file. Retrieve local arguments from a JSON file.
@@ -245,24 +308,37 @@ def get_local_args():
return local_args return local_args
def get_engine_args(): def get_engine_args():
# Start with default args # Start with worker custom defaults (only where we differ from vLLM)
args = DEFAULT_ARGS args = dict(DEFAULT_ARGS)
# Get env args that match keys in AsyncEngineArgs # Auto-discover: every AsyncEngineArgs field from env UPPERCASED (e.g. MAX_MODEL_LEN)
args.update(os.environ) args.update(_get_args_from_env_auto_discover())
# Get local args if model is baked in and overwrite env args # Backward-compat aliases (MODEL_NAME → model, etc.)
args.update(get_local_args()) _apply_env_aliases(args)
# Local baked-in model overrides
local = get_local_args()
if local:
args.update(_local_args_to_engine_args(local))
# Filter to valid engine args and drop sentinel empty values
valid_fields = AsyncEngineArgs.__dataclass_fields__
args = {
k: v for k, v in args.items()
if k in valid_fields and v not in (None, "", "None")
}
# Special conversion for limit_mm_per_prompt (e.g. "image=1,video=0")
limit_mm_env = os.getenv("LIMIT_MM_PER_PROMPT")
if limit_mm_env is not None:
args["limit_mm_per_prompt"] = convert_limit_mm_per_prompt(limit_mm_env)
# if args.get("TENSORIZER_URI"): TODO: add back once tensorizer is ready # if args.get("TENSORIZER_URI"): TODO: add back once tensorizer is ready
# args["load_format"] = "tensorizer" # args["load_format"] = "tensorizer"
# args["model_loader_extra_config"] = TensorizerConfig(tensorizer_uri=args["TENSORIZER_URI"], num_readers=None) # args["model_loader_extra_config"] = TensorizerConfig(tensorizer_uri=args["TENSORIZER_URI"], num_readers=None)
# logging.info(f"Using tensorized model from {args['TENSORIZER_URI']}") # logging.info(f"Using tensorized model from {args['TENSORIZER_URI']}")
# Rename and match to vllm args
args = match_vllm_args(args)
if args.get("load_format") == "bitsandbytes": if args.get("load_format") == "bitsandbytes":
args["quantization"] = args["load_format"] args["quantization"] = args["load_format"]