Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d3a962c33b | ||
|
|
105c125698 | ||
|
|
0a0ccfcb60 | ||
|
|
c8ce53c72c | ||
|
|
9618e799ba | ||
|
|
cb3f077dba | ||
|
|
69646b9e99 | ||
|
|
8b991a7ad7 | ||
|
|
d356c31675 | ||
|
|
80072047ab |
@@ -33,6 +33,17 @@ All behaviour is controlled through environment variables:
|
||||
|
||||
**Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options.
|
||||
|
||||
**Configuration file:** You can also supply a `config.yaml` instead of (or alongside) env vars. Mount it at `/vllm_config.yaml` in the container, or set `VLLM_CONFIG_FILE` to a custom path. Use the same key names as `vllm serve` — hyphens and underscores both work:
|
||||
|
||||
```yaml
|
||||
model: meta-llama/Llama-3.1-8B-Instruct
|
||||
max-model-len: 8192
|
||||
gpu-memory-utilization: 0.90
|
||||
quantization: awq
|
||||
```
|
||||
|
||||
Environment variables always override config file values.
|
||||
|
||||
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
|
||||
|
||||
### Specify Transformers Version
|
||||
|
||||
+12
-1
@@ -8,9 +8,20 @@ ENV PATH="/root/.local/bin:$PATH"
|
||||
|
||||
RUN ldconfig /usr/local/cuda-13.0/compat/
|
||||
|
||||
# nixl_ep PyPI wheels are compiled against CUDA 12.x and require libcudart.so.12.
|
||||
# CUDA 13 runtime is ABI-compatible with CUDA 12, so symlinking is safe.
|
||||
# Symlink into /usr/local/cuda/lib64 (already in LD_LIBRARY_PATH) so the linker
|
||||
# finds it by filename scan rather than relying on ldcache SONAME lookup.
|
||||
RUN ln -sf /usr/local/cuda/lib64/libcudart.so.13 /usr/local/cuda/lib64/libcudart.so.12 && ldconfig
|
||||
|
||||
# CUDA 13.0 containers return libs to /usr/local/nvidia/lib64 so container
|
||||
# providers (RunPod, Lambda, etc.) can mount host drivers there consistently.
|
||||
# See: https://github.com/vllm-project/vllm/issues/18859
|
||||
ENV LD_LIBRARY_PATH=/usr/local/nvidia/lib64:/usr/local/cuda/lib64:$LD_LIBRARY_PATH
|
||||
|
||||
# Install vLLM with FlashInfer - use CUDA 130 PyTorch wheels
|
||||
RUN uv pip install --system "packaging>=24.2" && \
|
||||
uv pip install --system "vllm[flashinfer]==0.20.2" && \
|
||||
uv pip install --system "vllm[flashinfer]==0.21.0" && \
|
||||
uv pip install --system git+https://github.com/deepseek-ai/DeepGEMM.git@714dd1a4a980f7937a74343d19a8eba4fe321480 --no-build-isolation
|
||||
|
||||
# Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts)
|
||||
|
||||
@@ -78,6 +78,20 @@ Configure worker-vllm using environment variables:
|
||||
|
||||
Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is applied automatically. Backward-compat aliases: `MODEL_NAME`, `TOKENIZER_NAME`, `MAX_CONTEXT_LEN_TO_CAPTURE`. This lets you configure any vLLM option without waiting for explicit worker support.
|
||||
|
||||
### Configuration File (config.yaml)
|
||||
|
||||
As an alternative to environment variables, you can supply a `config.yaml` file using the same key names as `vllm serve` (hyphens or underscores both work):
|
||||
|
||||
```yaml
|
||||
model: meta-llama/Llama-3.1-8B-Instruct
|
||||
max-model-len: 8192
|
||||
gpu-memory-utilization: 0.90
|
||||
quantization: awq
|
||||
tensor-parallel-size: 2
|
||||
```
|
||||
|
||||
Mount the file into the container at `/vllm_config.yaml`, or point to a custom path with the `VLLM_CONFIG_FILE` env var. Environment variables always take precedence over config file values.
|
||||
|
||||
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
||||
|
||||
### Specify Transformers Version
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
ray
|
||||
pandas
|
||||
pyarrow
|
||||
runpod==1.9.0
|
||||
runpod==1.9.1
|
||||
huggingface-hub
|
||||
lmcache==0.4.5
|
||||
lmcache==0.4.6
|
||||
packaging>=24.2
|
||||
typing-extensions>=4.8.0
|
||||
pydantic
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
model: meta-llama/Llama-3.1-8B-Instruct
|
||||
gpu-memory-utilization: 0.95
|
||||
max-model-len: 8192
|
||||
dtype: auto
|
||||
trust-remote-code: true
|
||||
quantization: fp8
|
||||
kv-cache-dtype: fp8
|
||||
enforce-eager: false
|
||||
enable-prefix-caching: true
|
||||
speculative-config: '{"model":"RedHatAI/Llama-3.1-8B-Instruct-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
|
||||
@@ -0,0 +1,10 @@
|
||||
model: Qwen/Qwen3-8B
|
||||
gpu-memory-utilization: 0.95
|
||||
max-model-len: 8192
|
||||
dtype: auto
|
||||
trust-remote-code: true
|
||||
quantization: fp8
|
||||
kv-cache-dtype: fp8
|
||||
enforce-eager: false
|
||||
enable-prefix-caching: true
|
||||
speculative-config: '{"model":"RedHatAI/Qwen3-8B-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
|
||||
@@ -404,6 +404,23 @@ def _resolve_cached_model_path(model_name: str) -> str:
|
||||
return resolved
|
||||
|
||||
|
||||
def _get_args_from_config_file() -> dict:
|
||||
"""Load engine args from a vLLM-style config.yaml.
|
||||
|
||||
Checks VLLM_CONFIG_FILE env var, then falls back to /vllm_config.yaml.
|
||||
Keys use the same long-form names as vllm serve (hyphens converted to underscores).
|
||||
"""
|
||||
import yaml
|
||||
path = os.getenv("VLLM_CONFIG_FILE", "/vllm_config.yaml")
|
||||
if not os.path.exists(path):
|
||||
return {}
|
||||
with open(path) as f:
|
||||
raw = yaml.safe_load(f) or {}
|
||||
normalized = {k.replace("-", "_"): v for k, v in raw.items()}
|
||||
logging.info("Loaded engine args from config file %s: %s", path, list(normalized.keys()))
|
||||
return normalized
|
||||
|
||||
|
||||
def get_local_args():
|
||||
"""
|
||||
Retrieve local arguments from a JSON file.
|
||||
@@ -429,6 +446,9 @@ def get_engine_args():
|
||||
# Start with worker custom defaults (only where we differ from vLLM)
|
||||
args = dict(DEFAULT_ARGS)
|
||||
|
||||
# Config file values sit above defaults but below env vars
|
||||
args.update(_get_args_from_config_file())
|
||||
|
||||
# Auto-discover: every AsyncEngineArgs field from env UPPERCASED (e.g. MAX_MODEL_LEN)
|
||||
args.update(_get_args_from_env_auto_discover())
|
||||
|
||||
|
||||
Reference in New Issue
Block a user