Compare commits

..
14 Commits
Author SHA1 Message Date
chrisvelaandGitHub 5d9a48fc70 Merge pull request #302 from runpod-workers/feat/0.22.1
Release / release (push) Waiting to run
feat: upgrade vllm to 0.22.1
2026-06-11 17:27:38 -05:00
velaraptor-runpod 5d1579e361 fix: update readme links 2026-06-11 17:00:31 -05:00
velaraptor-runpod 0488b77d89 feat: upgrade vllm to 0.22.1 2026-06-11 16:29:02 -05:00
chrisvelaandGitHub d3a962c33b Merge pull request #301 from adithyaJRunpod/feature/tuned-configs
Release / release (push) Waiting to run
Add tuned configs,  CON-239
2026-06-11 14:26:00 -05:00
chrisvelaandGitHub 105c125698 Merge pull request #300 from runpod-workers/feat/0.21.0
feat: upgrade vllm to 0.21.0
2026-06-11 12:47:08 -05:00
AdithyaJob 0a0ccfcb60 Add tuned configs for Llama 3.1 8B and Qwen3 8B 2026-06-10 21:07:41 -07:00
velaraptor-runpod c8ce53c72c fix: add kenels, and fix for cuda 2026-06-10 16:01:53 -05:00
velaraptor-runpod 9618e799ba chore: fix cuda libraries 2026-06-04 15:35:58 -05:00
velaraptor-runpod cb3f077dba feat: upgrade vllm to 0.21.0 2026-06-03 14:43:56 -05:00
chrisvelaandGitHub 69646b9e99 Merge pull request #294 from runpod-workers/feat/allow-config
Release / release (push) Waiting to run
feat: allow config.yaml like vllm serve
2026-06-02 20:28:02 -05:00
Jacob CiparandGitHub 8b991a7ad7 Merge pull request #297 from runpod-workers/jhcipar/bump-runpod-python-version
feat: bump runpod-python version
2026-06-02 11:17:48 -04:00
Tim PietruskyandGitHub dac05b62b3 fix: make .runpod/tests.json hub tests pass on CUDA 13.0 (#299)
Release / release (push) Waiting to run
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (https://github.com/huggingface/kernels/pull/544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
2026-06-02 17:08:54 +02:00
jhcipar d356c31675 feat: bump runpod-python version 2026-06-01 20:28:40 -04:00
velaraptor-runpod 80072047ab feat: allow config.yaml like vllm serve 2026-05-29 15:50:22 -05:00
8 changed files with 85 additions and 9 deletions
+12 -1
View File
@@ -6,7 +6,7 @@ Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
[![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm) [![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
Current vLLM version: [0.20.2](https://github.com/vllm-project/vllm/releases/tag/v0.20.2) Current vLLM version: [0.22.1](https://github.com/vllm-project/vllm/releases/tag/v0.22.1)
--- ---
@@ -33,6 +33,17 @@ All behaviour is controlled through environment variables:
**Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options. **Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options.
**Configuration file:** You can also supply a `config.yaml` instead of (or alongside) env vars. Mount it at `/vllm_config.yaml` in the container, or set `VLLM_CONFIG_FILE` to a custom path. Use the same key names as `vllm serve` — hyphens and underscores both work:
```yaml
model: meta-llama/Llama-3.1-8B-Instruct
max-model-len: 8192
gpu-memory-utilization: 0.90
quantization: awq
```
Environment variables always override config file values.
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md). For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
### Specify Transformers Version ### Specify Transformers Version
+3 -3
View File
@@ -5,7 +5,7 @@
"input": { "input": {
"prompt": "Write a short poem about artificial intelligence." "prompt": "Write a short poem about artificial intelligence."
}, },
"timeout": 30000 "timeout": 300000
}, },
{ {
"name": "openai_messages_test", "name": "openai_messages_test",
@@ -26,7 +26,7 @@
"temperature": 0.1 "temperature": 0.1
} }
}, },
"timeout": 30000 "timeout": 300000
} }
], ],
"config": { "config": {
@@ -38,6 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct" "value": "HuggingFaceTB/SmolLM2-135M-Instruct"
} }
], ],
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"] "allowedCudaVersions": ["13.0"]
} }
} }
+12 -1
View File
@@ -8,9 +8,20 @@ ENV PATH="/root/.local/bin:$PATH"
RUN ldconfig /usr/local/cuda-13.0/compat/ RUN ldconfig /usr/local/cuda-13.0/compat/
# nixl_ep PyPI wheels are compiled against CUDA 12.x and require libcudart.so.12.
# CUDA 13 runtime is ABI-compatible with CUDA 12, so symlinking is safe.
# Symlink into /usr/local/cuda/lib64 (already in LD_LIBRARY_PATH) so the linker
# finds it by filename scan rather than relying on ldcache SONAME lookup.
RUN ln -sf /usr/local/cuda/lib64/libcudart.so.13 /usr/local/cuda/lib64/libcudart.so.12 && ldconfig
# CUDA 13.0 containers return libs to /usr/local/nvidia/lib64 so container
# providers (RunPod, Lambda, etc.) can mount host drivers there consistently.
# See: https://github.com/vllm-project/vllm/issues/18859
ENV LD_LIBRARY_PATH=/usr/local/nvidia/lib64:/usr/local/cuda/lib64:$LD_LIBRARY_PATH
# Install vLLM with FlashInfer - use CUDA 130 PyTorch wheels # Install vLLM with FlashInfer - use CUDA 130 PyTorch wheels
RUN uv pip install --system "packaging>=24.2" && \ RUN uv pip install --system "packaging>=24.2" && \
uv pip install --system "vllm[flashinfer]==0.20.2" && \ uv pip install --system "vllm[flashinfer]==0.22.1" && \
uv pip install --system git+https://github.com/deepseek-ai/DeepGEMM.git@714dd1a4a980f7937a74343d19a8eba4fe321480 --no-build-isolation uv pip install --system git+https://github.com/deepseek-ai/DeepGEMM.git@714dd1a4a980f7937a74343d19a8eba4fe321480 --no-build-isolation
# Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts) # Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts)
+15 -1
View File
@@ -8,7 +8,7 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png) ![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png)
Current vLLM version: [0.20.2](https://github.com/vllm-project/vllm/releases/tag/v0.20.2) Current vLLM version: [0.22.1](https://github.com/vllm-project/vllm/releases/tag/v0.22.1)
> Check out our Load Balancer implementation here: [vLLM Load Balancer](https://github.com/runpod-workers/vllm-loadbalancer-ep) > Check out our Load Balancer implementation here: [vLLM Load Balancer](https://github.com/runpod-workers/vllm-loadbalancer-ep)
@@ -78,6 +78,20 @@ Configure worker-vllm using environment variables:
Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is applied automatically. Backward-compat aliases: `MODEL_NAME`, `TOKENIZER_NAME`, `MAX_CONTEXT_LEN_TO_CAPTURE`. This lets you configure any vLLM option without waiting for explicit worker support. Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is applied automatically. Backward-compat aliases: `MODEL_NAME`, `TOKENIZER_NAME`, `MAX_CONTEXT_LEN_TO_CAPTURE`. This lets you configure any vLLM option without waiting for explicit worker support.
### Configuration File (config.yaml)
As an alternative to environment variables, you can supply a `config.yaml` file using the same key names as `vllm serve` (hyphens or underscores both work):
```yaml
model: meta-llama/Llama-3.1-8B-Instruct
max-model-len: 8192
gpu-memory-utilization: 0.90
quantization: awq
tensor-parallel-size: 2
```
Mount the file into the container at `/vllm_config.yaml`, or point to a custom path with the `VLLM_CONFIG_FILE` env var. Environment variables always take precedence over config file values.
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)** For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
### Specify Transformers Version ### Specify Transformers Version
+3 -3
View File
@@ -1,9 +1,9 @@
ray ray
pandas pandas
pyarrow pyarrow
runpod==1.9.0 runpod==1.9.1
huggingface-hub huggingface-hub
lmcache==0.4.5 lmcache==0.4.6
packaging>=24.2 packaging>=24.2
typing-extensions>=4.8.0 typing-extensions>=4.8.0
pydantic pydantic
@@ -11,5 +11,5 @@ pydantic-settings
hf-transfer hf-transfer
transformers>=5 transformers>=5
bitsandbytes>=0.45.0 bitsandbytes>=0.45.0
kernels kernels<0.15
torch-c-dlpack-ext torch-c-dlpack-ext
+10
View File
@@ -0,0 +1,10 @@
model: meta-llama/Llama-3.1-8B-Instruct
gpu-memory-utilization: 0.95
max-model-len: 8192
dtype: auto
trust-remote-code: true
quantization: fp8
kv-cache-dtype: fp8
enforce-eager: false
enable-prefix-caching: true
speculative-config: '{"model":"RedHatAI/Llama-3.1-8B-Instruct-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
+10
View File
@@ -0,0 +1,10 @@
model: Qwen/Qwen3-8B
gpu-memory-utilization: 0.95
max-model-len: 8192
dtype: auto
trust-remote-code: true
quantization: fp8
kv-cache-dtype: fp8
enforce-eager: false
enable-prefix-caching: true
speculative-config: '{"model":"RedHatAI/Qwen3-8B-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
+20
View File
@@ -404,6 +404,23 @@ def _resolve_cached_model_path(model_name: str) -> str:
return resolved return resolved
def _get_args_from_config_file() -> dict:
"""Load engine args from a vLLM-style config.yaml.
Checks VLLM_CONFIG_FILE env var, then falls back to /vllm_config.yaml.
Keys use the same long-form names as vllm serve (hyphens converted to underscores).
"""
import yaml
path = os.getenv("VLLM_CONFIG_FILE", "/vllm_config.yaml")
if not os.path.exists(path):
return {}
with open(path) as f:
raw = yaml.safe_load(f) or {}
normalized = {k.replace("-", "_"): v for k, v in raw.items()}
logging.info("Loaded engine args from config file %s: %s", path, list(normalized.keys()))
return normalized
def get_local_args(): def get_local_args():
""" """
Retrieve local arguments from a JSON file. Retrieve local arguments from a JSON file.
@@ -429,6 +446,9 @@ def get_engine_args():
# Start with worker custom defaults (only where we differ from vLLM) # Start with worker custom defaults (only where we differ from vLLM)
args = dict(DEFAULT_ARGS) args = dict(DEFAULT_ARGS)
# Config file values sit above defaults but below env vars
args.update(_get_args_from_config_file())
# Auto-discover: every AsyncEngineArgs field from env UPPERCASED (e.g. MAX_MODEL_LEN) # Auto-discover: every AsyncEngineArgs field from env UPPERCASED (e.g. MAX_MODEL_LEN)
args.update(_get_args_from_env_auto_discover()) args.update(_get_args_from_env_auto_discover())