Compare commits

...
5 Commits
Author SHA1 Message Date
chrisvelaandGitHub 69646b9e99 Merge pull request #294 from runpod-workers/feat/allow-config
Release / release (push) Waiting to run
feat: allow config.yaml like vllm serve
2026-06-02 20:28:02 -05:00
Jacob CiparandGitHub 8b991a7ad7 Merge pull request #297 from runpod-workers/jhcipar/bump-runpod-python-version
feat: bump runpod-python version
2026-06-02 11:17:48 -04:00
Tim PietruskyandGitHub dac05b62b3 fix: make .runpod/tests.json hub tests pass on CUDA 13.0 (#299)
Release / release (push) Waiting to run
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (https://github.com/huggingface/kernels/pull/544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
2026-06-02 17:08:54 +02:00
jhcipar d356c31675 feat: bump runpod-python version 2026-06-01 20:28:40 -04:00
velaraptor-runpod 80072047ab feat: allow config.yaml like vllm serve 2026-05-29 15:50:22 -05:00
5 changed files with 50 additions and 5 deletions
+11
View File
@@ -33,6 +33,17 @@ All behaviour is controlled through environment variables:
**Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options.
**Configuration file:** You can also supply a `config.yaml` instead of (or alongside) env vars. Mount it at `/vllm_config.yaml` in the container, or set `VLLM_CONFIG_FILE` to a custom path. Use the same key names as `vllm serve` — hyphens and underscores both work:
```yaml
model: meta-llama/Llama-3.1-8B-Instruct
max-model-len: 8192
gpu-memory-utilization: 0.90
quantization: awq
```
Environment variables always override config file values.
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
### Specify Transformers Version
+3 -3
View File
@@ -5,7 +5,7 @@
"input": {
"prompt": "Write a short poem about artificial intelligence."
},
"timeout": 30000
"timeout": 300000
},
{
"name": "openai_messages_test",
@@ -26,7 +26,7 @@
"temperature": 0.1
}
},
"timeout": 30000
"timeout": 300000
}
],
"config": {
@@ -38,6 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
}
],
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
"allowedCudaVersions": ["13.0"]
}
}
+14
View File
@@ -78,6 +78,20 @@ Configure worker-vllm using environment variables:
Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is applied automatically. Backward-compat aliases: `MODEL_NAME`, `TOKENIZER_NAME`, `MAX_CONTEXT_LEN_TO_CAPTURE`. This lets you configure any vLLM option without waiting for explicit worker support.
### Configuration File (config.yaml)
As an alternative to environment variables, you can supply a `config.yaml` file using the same key names as `vllm serve` (hyphens or underscores both work):
```yaml
model: meta-llama/Llama-3.1-8B-Instruct
max-model-len: 8192
gpu-memory-utilization: 0.90
quantization: awq
tensor-parallel-size: 2
```
Mount the file into the container at `/vllm_config.yaml`, or point to a custom path with the `VLLM_CONFIG_FILE` env var. Environment variables always take precedence over config file values.
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
### Specify Transformers Version
+2 -2
View File
@@ -1,7 +1,7 @@
ray
pandas
pyarrow
runpod==1.9.0
runpod==1.9.1
huggingface-hub
lmcache==0.4.5
packaging>=24.2
@@ -11,5 +11,5 @@ pydantic-settings
hf-transfer
transformers>=5
bitsandbytes>=0.45.0
kernels
kernels<0.15
torch-c-dlpack-ext
+20
View File
@@ -404,6 +404,23 @@ def _resolve_cached_model_path(model_name: str) -> str:
return resolved
def _get_args_from_config_file() -> dict:
"""Load engine args from a vLLM-style config.yaml.
Checks VLLM_CONFIG_FILE env var, then falls back to /vllm_config.yaml.
Keys use the same long-form names as vllm serve (hyphens converted to underscores).
"""
import yaml
path = os.getenv("VLLM_CONFIG_FILE", "/vllm_config.yaml")
if not os.path.exists(path):
return {}
with open(path) as f:
raw = yaml.safe_load(f) or {}
normalized = {k.replace("-", "_"): v for k, v in raw.items()}
logging.info("Loaded engine args from config file %s: %s", path, list(normalized.keys()))
return normalized
def get_local_args():
"""
Retrieve local arguments from a JSON file.
@@ -429,6 +446,9 @@ def get_engine_args():
# Start with worker custom defaults (only where we differ from vLLM)
args = dict(DEFAULT_ARGS)
# Config file values sit above defaults but below env vars
args.update(_get_args_from_config_file())
# Auto-discover: every AsyncEngineArgs field from env UPPERCASED (e.g. MAX_MODEL_LEN)
args.update(_get_args_from_env_auto_discover())