Compare commits

...
16 Commits
Author SHA1 Message Date
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
chrisvelaandGitHub 3851d53f93 add ENABLE_EXPERT_PARALLEL engine arg for MoE models (#239)
Release / release (push) Waiting to run
* enable expert parallel arg for moe models

* add ENABLE_EXPERT_PARALLEL to hub config
2025-11-17 19:25:19 +01:00
Witold WydmańskiandGitHub c896438f21 feat: bump transformers to allow Qwen3-VL (#225)
Release / release (push) Waiting to run
2025-11-14 17:23:34 +01:00
Tim PietruskyandGitHub 912892f94e fix: remove space from gpuIds (#234) 2025-11-14 17:23:09 +01:00
Tim PietruskyandGitHub f8bf82469c fix(config): update allowed cuda versions in hub and tests config (#236)
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment

- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions

refs: AE-1452
2025-11-14 17:22:43 +01:00
Hailong YangandGitHub ec1664902b Merge pull request #230 from runpod-workers/feat/cse-853-vllm-template-params
Feat/cse 853 vllm template params
2025-10-31 13:32:34 -04:00
Eugene Klitenik d09122de4a remove un-needed 2025-10-29 13:06:33 -04:00
Eugene Klitenik e27dc68dea remove uneeded 2025-10-29 13:05:26 -04:00
Eugene Klitenik 1ee18d06a9 determine num gpus in python 2025-10-29 11:31:42 -04:00
Eugene Klitenik 5c4edd15cc update entrypoint command 2025-10-28 18:07:18 -04:00
Eugene Klitenik b074d3a23b auto detect num GPUs 2025-10-28 14:39:55 -04:00
Eugene Klitenik 205847471c reduce default container disk size to 150GB 2025-10-28 13:38:16 -04:00
Tim PietruskyandGitHub 6337a6673a fix: allow also CUDA 12.8 & 12.9 (#228)
Release / release (push) Waiting to run
2025-10-24 18:48:26 +02:00
Tim PietruskyandGitHub 66e1b1605b Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
2025-10-22 22:54:21 +02:00
Tim PietruskyandGitHub 60c8f257a8 Merge pull request #227 from runpod-workers/chore/vllm-0.11.0
chore: update vllm to 0.11.0
2025-10-22 22:53:51 +02:00
Tim Pietrusky fae16e7ee1 chore: update vllm to 0.11.0 2025-10-22 13:40:53 -07:00
10 changed files with 26 additions and 1545 deletions
+14 -14
View File
@@ -6,20 +6,10 @@
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png", "iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
"config": { "config": {
"runsOn": "GPU", "runsOn": "GPU",
"containerDiskInGb": 200, "containerDiskInGb": 150,
"gpuIds": "ADA_80_PRO, AMPERE_80", "gpuIds": "ADA_80_PRO,AMPERE_80",
"gpuCount": 1, "gpuCount": 1,
"allowedCudaVersions": [ "allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5", "12.4"],
"12.9",
"12.8",
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
],
"presets": [ "presets": [
{ {
"name": "deepseek-ai/deepseek-r1-distill-llama-8b", "name": "deepseek-ai/deepseek-r1-distill-llama-8b",
@@ -935,7 +925,17 @@
"name": "Max Concurrency", "name": "Max Concurrency",
"type": "number", "type": "number",
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency", "description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
"default": 300, "default": 30,
"advanced": true
}
},
{
"key": "ENABLE_EXPERT_PARALLEL",
"input": {
"name": "Enable Expert Parallel",
"type": "boolean",
"description": "Enable Expert Parallel for MoE models",
"default": false,
"advanced": true "advanced": true
} }
}, },
+1 -9
View File
@@ -38,14 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct" "value": "HuggingFaceTB/SmolLM2-135M-Instruct"
} }
], ],
"allowedCudaVersions": [ "allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
]
} }
} }
+1 -1
View File
@@ -12,7 +12,7 @@ RUN --mount=type=cache,target=/root/.cache/pip \
python3 -m pip install --upgrade -r /requirements.txt python3 -m pip install --upgrade -r /requirements.txt
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer # Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
RUN python3 -m pip install vllm==0.10.0 && \ RUN python3 -m pip install vllm==0.11.0 && \
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3 python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
# Setup for Option 2: Building the Image with the Model included # Setup for Option 2: Building the Image with the Model included
+1 -1
View File
@@ -57,7 +57,7 @@ Configure worker-vllm using environment variables:
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) | | `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. | | `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String | | `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer | | `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)** For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
+2 -2
View File
@@ -1,14 +1,14 @@
ray ray
pandas pandas
pyarrow pyarrow
runpod~=1.7.7 runpod>=1.8,<2.0
huggingface-hub huggingface-hub
packaging packaging
typing-extensions>=4.8.0 typing-extensions>=4.8.0
pydantic pydantic
pydantic-settings pydantic-settings
hf-transfer hf-transfer
transformers>=4.55.0 transformers>=4.57.0
bitsandbytes>=0.45.0 bitsandbytes>=0.45.0
kernels kernels
torch==2.6.0 torch==2.6.0
+2 -1
View File
@@ -85,6 +85,7 @@ Complete guide to all environment variables and configuration options for worker
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. | | `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. | | `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. | | `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models |
## Tokenizer Settings ## Tokenizer Settings
@@ -119,7 +120,7 @@ The way this works is that the first request will have a batch size of `DEFAULT_
| Variable | Default | Type/Choices | Description | | Variable | Default | Type/Choices | Description |
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency | | `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. | | `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. | | `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
+3 -2
View File
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
- `src/engine_args.py`: Centralized configuration management - `src/engine_args.py`: Centralized configuration management
- `src/constants.py`: Default values for core settings - `src/constants.py`: Default values for core settings
- `worker-config.json`: UI form generation for RunPod console - `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
- `worker-config.json`: UI form generation for RunPod console (if exists)
## Core Development Concepts ## Core Development Concepts
@@ -222,7 +223,7 @@ src/
### 2. **Concurrency Patterns** ### 2. **Concurrency Patterns**
- **Max Concurrency**: 300 concurrent requests by default - **Max Concurrency**: 30 concurrent requests by default
- **vLLM Queuing**: Internal request batching and scheduling - **vLLM Queuing**: Internal request batching and scheduling
- **RunPod Integration**: Concurrency modifier for auto-scaling - **RunPod Integration**: Concurrency modifier for auto-scaling
+1 -1
View File
@@ -1,4 +1,4 @@
DEFAULT_BATCH_SIZE = 50 DEFAULT_BATCH_SIZE = 50
DEFAULT_MAX_CONCURRENCY = 300 DEFAULT_MAX_CONCURRENCY = 30
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3 DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
DEFAULT_MIN_BATCH_SIZE = 1 DEFAULT_MIN_BATCH_SIZE = 1
+1
View File
@@ -80,6 +80,7 @@ DEFAULT_ARGS = {
"guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'), "guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'),
"speculative_model": os.getenv('SPECULATIVE_MODEL', None), "speculative_model": os.getenv('SPECULATIVE_MODEL', None),
"speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None, "speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None,
"enable_expert_parallel": bool(os.getenv('ENABLE_EXPERT_PARALLEL', 'False').lower() == 'true'),
"num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None, "num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None,
"speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None, "speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None,
"speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None, "speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None,
-1514
View File
File diff suppressed because it is too large Load Diff