Compare commits

..
18 Commits
Author SHA1 Message Date
90c16b472d fix: update CUDA to 12.4.1 for Blackwell GPU support (#251)
Release / release (push) Waiting to run
* fix: update CUDA to 12.4.1 for Blackwell GPU support

- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json

This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.

Fixes: DR-1118

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* revert: remove NVIDIA B200 from default gpuIds

The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove FlashInfer to avoid JIT compilation errors

FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.

vLLM will use its built-in fallback sampling methods instead.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-13 22:01:36 +01:00
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
chrisvelaandGitHub 3851d53f93 add ENABLE_EXPERT_PARALLEL engine arg for MoE models (#239)
Release / release (push) Waiting to run
* enable expert parallel arg for moe models

* add ENABLE_EXPERT_PARALLEL to hub config
2025-11-17 19:25:19 +01:00
Witold WydmańskiandGitHub c896438f21 feat: bump transformers to allow Qwen3-VL (#225)
Release / release (push) Waiting to run
2025-11-14 17:23:34 +01:00
Tim PietruskyandGitHub 912892f94e fix: remove space from gpuIds (#234) 2025-11-14 17:23:09 +01:00
Tim PietruskyandGitHub f8bf82469c fix(config): update allowed cuda versions in hub and tests config (#236)
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment

- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions

refs: AE-1452
2025-11-14 17:22:43 +01:00
Hailong YangandGitHub ec1664902b Merge pull request #230 from runpod-workers/feat/cse-853-vllm-template-params
Feat/cse 853 vllm template params
2025-10-31 13:32:34 -04:00
Eugene Klitenik d09122de4a remove un-needed 2025-10-29 13:06:33 -04:00
Eugene Klitenik e27dc68dea remove uneeded 2025-10-29 13:05:26 -04:00
Eugene Klitenik 1ee18d06a9 determine num gpus in python 2025-10-29 11:31:42 -04:00
Eugene Klitenik 5c4edd15cc update entrypoint command 2025-10-28 18:07:18 -04:00
Eugene Klitenik b074d3a23b auto detect num GPUs 2025-10-28 14:39:55 -04:00
Eugene Klitenik 205847471c reduce default container disk size to 150GB 2025-10-28 13:38:16 -04:00
Tim PietruskyandGitHub 6337a6673a fix: allow also CUDA 12.8 & 12.9 (#228)
Release / release (push) Waiting to run
2025-10-24 18:48:26 +02:00
Tim PietruskyandGitHub 66e1b1605b Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
2025-10-22 22:54:21 +02:00
Tim PietruskyandGitHub 60c8f257a8 Merge pull request #227 from runpod-workers/chore/vllm-0.11.0
chore: update vllm to 0.11.0
2025-10-22 22:53:51 +02:00
Tim Pietrusky fae16e7ee1 chore: update vllm to 0.11.0 2025-10-22 13:40:53 -07:00
max4c 2becd35345 Revert "fix: added back the HF_TOKEN (#219)"
Release / release (push) Waiting to run
This reverts commit 33d88df6c0.
2025-09-23 12:38:24 -07:00
10 changed files with 29 additions and 1559 deletions
+14 -24
View File
@@ -6,20 +6,10 @@
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
"config": {
"runsOn": "GPU",
"containerDiskInGb": 200,
"gpuIds": "ADA_80_PRO, AMPERE_80",
"containerDiskInGb": 150,
"gpuIds": "ADA_80_PRO,AMPERE_80",
"gpuCount": 1,
"allowedCudaVersions": [
"12.9",
"12.8",
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
],
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5", "12.4"],
"presets": [
{
"name": "deepseek-ai/deepseek-r1-distill-llama-8b",
@@ -38,16 +28,6 @@
"required": true
}
},
{
"key": "HF_TOKEN",
"input": {
"name": "Access Token",
"type": "string",
"description": "Hugging Face access token for gated & private models",
"default": "",
"required": false
}
},
{
"key": "TOKENIZER",
"input": {
@@ -945,7 +925,17 @@
"name": "Max Concurrency",
"type": "number",
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
"default": 300,
"default": 30,
"advanced": true
}
},
{
"key": "ENABLE_EXPERT_PARALLEL",
"input": {
"name": "Enable Expert Parallel",
"type": "boolean",
"description": "Enable Expert Parallel for MoE models",
"default": false,
"advanced": true
}
},
+1 -9
View File
@@ -38,14 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
}
],
"allowedCudaVersions": [
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
]
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
}
}
+4 -5
View File
@@ -1,9 +1,9 @@
FROM nvidia/cuda:12.1.0-base-ubuntu22.04
FROM nvidia/cuda:12.4.1-base-ubuntu22.04
RUN apt-get update -y \
&& apt-get install -y python3-pip
RUN ldconfig /usr/local/cuda-12.1/compat/
RUN ldconfig /usr/local/cuda-12.4/compat/
# Install Python dependencies
COPY builder/requirements.txt /requirements.txt
@@ -11,9 +11,8 @@ RUN --mount=type=cache,target=/root/.cache/pip \
python3 -m pip install --upgrade pip && \
python3 -m pip install --upgrade -r /requirements.txt
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
RUN python3 -m pip install vllm==0.10.0 && \
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
# Install vLLM
RUN python3 -m pip install vllm==0.11.0
# Setup for Option 2: Building the Image with the Model included
ARG MODEL_NAME=""
+1 -1
View File
@@ -57,7 +57,7 @@ Configure worker-vllm using environment variables:
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
+2 -2
View File
@@ -1,14 +1,14 @@
ray
pandas
pyarrow
runpod~=1.7.7
runpod>=1.8,<2.0
huggingface-hub
packaging
typing-extensions>=4.8.0
pydantic
pydantic-settings
hf-transfer
transformers>=4.55.0
transformers>=4.57.0
bitsandbytes>=0.45.0
kernels
torch==2.6.0
+2 -1
View File
@@ -85,6 +85,7 @@ Complete guide to all environment variables and configuration options for worker
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models |
## Tokenizer Settings
@@ -119,7 +120,7 @@ The way this works is that the first request will have a batch size of `DEFAULT_
| Variable | Default | Type/Choices | Description |
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
+3 -2
View File
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
- `src/engine_args.py`: Centralized configuration management
- `src/constants.py`: Default values for core settings
- `worker-config.json`: UI form generation for RunPod console
- `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
- `worker-config.json`: UI form generation for RunPod console (if exists)
## Core Development Concepts
@@ -222,7 +223,7 @@ src/
### 2. **Concurrency Patterns**
- **Max Concurrency**: 300 concurrent requests by default
- **Max Concurrency**: 30 concurrent requests by default
- **vLLM Queuing**: Internal request batching and scheduling
- **RunPod Integration**: Concurrency modifier for auto-scaling
+1 -1
View File
@@ -1,4 +1,4 @@
DEFAULT_BATCH_SIZE = 50
DEFAULT_MAX_CONCURRENCY = 300
DEFAULT_MAX_CONCURRENCY = 30
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
DEFAULT_MIN_BATCH_SIZE = 1
+1
View File
@@ -80,6 +80,7 @@ DEFAULT_ARGS = {
"guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'),
"speculative_model": os.getenv('SPECULATIVE_MODEL', None),
"speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None,
"enable_expert_parallel": bool(os.getenv('ENABLE_EXPERT_PARALLEL', 'False').lower() == 'true'),
"num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None,
"speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None,
"speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None,
-1514
View File
File diff suppressed because it is too large Load Diff