Compare commits

...
22 Commits
Author SHA1 Message Date
Tim PietruskyandGitHub 6d6cbe7095 fix: deactivate RunPod tests to fix hub release (#253)
Release / release (push) Waiting to run
Rename tests.json to tests_json to temporarily disable automated
tests while fixing the release on the hub.
2026-01-22 18:06:36 +01:00
90c16b472d fix: update CUDA to 12.4.1 for Blackwell GPU support (#251)
Release / release (push) Waiting to run
* fix: update CUDA to 12.4.1 for Blackwell GPU support

- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json

This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.

Fixes: DR-1118

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* revert: remove NVIDIA B200 from default gpuIds

The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove FlashInfer to avoid JIT compilation errors

FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.

vLLM will use its built-in fallback sampling methods instead.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-13 22:01:36 +01:00
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
chrisvelaandGitHub 3851d53f93 add ENABLE_EXPERT_PARALLEL engine arg for MoE models (#239)
Release / release (push) Waiting to run
* enable expert parallel arg for moe models

* add ENABLE_EXPERT_PARALLEL to hub config
2025-11-17 19:25:19 +01:00
Witold WydmańskiandGitHub c896438f21 feat: bump transformers to allow Qwen3-VL (#225)
Release / release (push) Waiting to run
2025-11-14 17:23:34 +01:00
Tim PietruskyandGitHub 912892f94e fix: remove space from gpuIds (#234) 2025-11-14 17:23:09 +01:00
Tim PietruskyandGitHub f8bf82469c fix(config): update allowed cuda versions in hub and tests config (#236)
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment

- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions

refs: AE-1452
2025-11-14 17:22:43 +01:00
Hailong YangandGitHub ec1664902b Merge pull request #230 from runpod-workers/feat/cse-853-vllm-template-params
Feat/cse 853 vllm template params
2025-10-31 13:32:34 -04:00
Eugene Klitenik d09122de4a remove un-needed 2025-10-29 13:06:33 -04:00
Eugene Klitenik e27dc68dea remove uneeded 2025-10-29 13:05:26 -04:00
Eugene Klitenik 1ee18d06a9 determine num gpus in python 2025-10-29 11:31:42 -04:00
Eugene Klitenik 5c4edd15cc update entrypoint command 2025-10-28 18:07:18 -04:00
Eugene Klitenik b074d3a23b auto detect num GPUs 2025-10-28 14:39:55 -04:00
Eugene Klitenik 205847471c reduce default container disk size to 150GB 2025-10-28 13:38:16 -04:00
Tim PietruskyandGitHub 6337a6673a fix: allow also CUDA 12.8 & 12.9 (#228)
Release / release (push) Waiting to run
2025-10-24 18:48:26 +02:00
Tim PietruskyandGitHub 66e1b1605b Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
2025-10-22 22:54:21 +02:00
Tim PietruskyandGitHub 60c8f257a8 Merge pull request #227 from runpod-workers/chore/vllm-0.11.0
chore: update vllm to 0.11.0
2025-10-22 22:53:51 +02:00
Tim Pietrusky fae16e7ee1 chore: update vllm to 0.11.0 2025-10-22 13:40:53 -07:00
max4c 2becd35345 Revert "fix: added back the HF_TOKEN (#219)"
Release / release (push) Waiting to run
This reverts commit 33d88df6c0.
2025-09-23 12:38:24 -07:00
33d88df6c0 fix: added back the HF_TOKEN (#219)
Release / release (push) Waiting to run
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-23 19:25:12 +02:00
Tim Pietrusky ecd562e112 fix: remove "access token" as this is handled by the platform
Release / release (push) Waiting to run
2025-09-19 21:00:47 +02:00
5cffaab8e8 docs: how to use the reasoning parser (#218)
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-17 07:49:37 +02:00
11 changed files with 309 additions and 1820 deletions
+261 -260
View File
@@ -1,260 +1,261 @@
![vLLM worker banner](https://cpjrphpz3t5wbwfe.public.blob.vercel-storage.com/worker-vllm_banner.jpeg) ![vLLM worker banner](https://cpjrphpz3t5wbwfe.public.blob.vercel-storage.com/worker-vllm_banner.jpeg)
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
--- ---
[![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm) [![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
--- ---
## Endpoint Configuration ## Endpoint Configuration
All behaviour is controlled through environment variables: All behaviour is controlled through environment variables:
| Environment Variable | Description | Default | Options | | Environment Variable | Description | Default | Options |
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ | | ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID | | `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token | | `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) | | `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" | | `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer | | `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 | | `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer | | `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string | | `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) | | `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. | | `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String | | `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer | | `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
## API Usage
## API Usage
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
### RunPod Native API
### RunPod Native API
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
#### Chat Completions
#### Chat Completions
```json
{ ```json
"input": { {
"messages": [ "input": {
{ "role": "system", "content": "You are a helpful assistant." }, "messages": [
{ "role": "user", "content": "What is the capital of France?" } { "role": "system", "content": "You are a helpful assistant." },
], { "role": "user", "content": "What is the capital of France?" }
"sampling_params": { ],
"max_tokens": 100, "sampling_params": {
"temperature": 0.7 "max_tokens": 100,
} "temperature": 0.7
} }
} }
``` }
```
#### Chat Completions (Streaming)
#### Chat Completions (Streaming)
```json
{ ```json
"input": { {
"messages": [ "input": {
{ "role": "user", "content": "Write a short story about a robot." } "messages": [
], { "role": "user", "content": "Write a short story about a robot." }
"sampling_params": { ],
"max_tokens": 500, "sampling_params": {
"temperature": 0.8 "max_tokens": 500,
}, "temperature": 0.8
"stream": true },
} "stream": true
} }
``` }
```
#### Text Generation
#### Text Generation
For direct text generation without chat format:
For direct text generation without chat format:
```json
{ ```json
"input": { {
"prompt": "The capital of France is", "input": {
"sampling_params": { "prompt": "The capital of France is",
"max_tokens": 64, "sampling_params": {
"temperature": 0.0 "max_tokens": 64,
} "temperature": 0.0
} }
} }
``` }
```
#### List Models
#### List Models
```json
{ ```json
"input": { {
"openai_route": "/v1/models" "input": {
} "openai_route": "/v1/models"
} }
``` }
```
---
---
### OpenAI-Compatible API
### OpenAI-Compatible API
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
#### Chat Completions
#### Chat Completions
**Path:** `/openai/v1/chat/completions`
**Path:** `/openai/v1/chat/completions`
```json
{ ```json
"model": "meta-llama/Llama-2-7b-chat-hf", {
"messages": [ "model": "meta-llama/Llama-2-7b-chat-hf",
{ "role": "system", "content": "You are a helpful assistant." }, "messages": [
{ "role": "user", "content": "What is the capital of France?" } { "role": "system", "content": "You are a helpful assistant." },
], { "role": "user", "content": "What is the capital of France?" }
"max_tokens": 100, ],
"temperature": 0.7 "max_tokens": 100,
} "temperature": 0.7
``` }
```
#### Chat Completions (Streaming)
#### Chat Completions (Streaming)
```json
{ ```json
"model": "meta-llama/Llama-2-7b-chat-hf", {
"messages": [ "model": "meta-llama/Llama-2-7b-chat-hf",
{ "role": "user", "content": "Write a short story about a robot." } "messages": [
], { "role": "user", "content": "Write a short story about a robot." }
"max_tokens": 500, ],
"temperature": 0.8, "max_tokens": 500,
"stream": true "temperature": 0.8,
} "stream": true
``` }
```
#### Text Completions
#### Text Completions
**Path:** `/openai/v1/completions`
**Path:** `/openai/v1/completions`
```json
{ ```json
"model": "meta-llama/Llama-2-7b-chat-hf", {
"prompt": "The capital of France is", "model": "meta-llama/Llama-2-7b-chat-hf",
"max_tokens": 100, "prompt": "The capital of France is",
"temperature": 0.7 "max_tokens": 100,
} "temperature": 0.7
``` }
```
#### List Models
#### List Models
**Path:** `/openai/v1/models`
**Path:** `/openai/v1/models`
```json
{} ```json
``` {}
```
#### Response Format
#### Response Format
Both APIs return the same response format:
Both APIs return the same response format:
```json
{ ```json
"choices": [ {
{ "choices": [
"index": 0, {
"message": { "role": "assistant", "content": "Paris." }, "index": 0,
"finish_reason": "stop" "message": { "role": "assistant", "content": "Paris." },
} "finish_reason": "stop"
], }
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 } ],
} "usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
``` }
```
---
---
## Usage
## Usage
Below are minimal `python` snippets so you can copy-paste to get started quickly.
Below are minimal `python` snippets so you can copy-paste to get started quickly.
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
### OpenAI compatible API
### OpenAI compatible API
Minimal Python example using the official `openai` SDK:
Minimal Python example using the official `openai` SDK:
```python
from openai import OpenAI ```python
import os from openai import OpenAI
import os
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
client = OpenAI( # Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
api_key=os.getenv("RUNPOD_API_KEY"), client = OpenAI(
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1", api_key=os.getenv("RUNPOD_API_KEY"),
) base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
``` )
```
`Chat Completions (Non-Streaming)`
`Chat Completions (Non-Streaming)`
```python
response = client.chat.completions.create( ```python
model="meta-llama/Llama-2-7b-chat-hf", response = client.chat.completions.create(
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}], model="meta-llama/Llama-2-7b-chat-hf",
temperature=0, messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
max_tokens=100, temperature=0,
) max_tokens=100,
print(f"Response: {response.choices[0].message.content}") )
``` print(f"Response: {response.choices[0].message.content}")
```
`Chat Completions (Streaming)`
`Chat Completions (Streaming)`
```python
response_stream = client.chat.completions.create( ```python
model="meta-llama/Llama-2-7b-chat-hf", response_stream = client.chat.completions.create(
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}], model="meta-llama/Llama-2-7b-chat-hf",
temperature=0, messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
max_tokens=100, temperature=0,
stream=True max_tokens=100,
) stream=True
for response in response_stream: )
print(response.choices[0].delta.content or "", end="", flush=True) for response in response_stream:
``` print(response.choices[0].delta.content or "", end="", flush=True)
```
### RunPod Native API
### RunPod Native API
```python
import requests ```python
import requests
response = requests.post(
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run", response = requests.post(
headers={"Authorization": "Bearer <API_KEY>"}, "https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
json={ headers={"Authorization": "Bearer <API_KEY>"},
"input": { json={
"messages": [ "input": {
{"role": "system", "content": "You are a helpful assistant."}, "messages": [
{"role": "user", "content": "Explain quantum computing in simple terms"} {"role": "system", "content": "You are a helpful assistant."},
], {"role": "user", "content": "Explain quantum computing in simple terms"}
"sampling_params": { ],
"temperature": 0.7, "sampling_params": {
"max_tokens": 150 "temperature": 0.7,
} "max_tokens": 150
} }
} }
) }
)
result = response.json()
print(result["output"]) result = response.json()
``` print(result["output"])
```
## Compatibility
## Compatibility
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
## Documentation
## Documentation
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables - **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies - **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns - **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
+32 -25
View File
@@ -1,25 +1,15 @@
{ {
"title": "vLLM", "title": "vLLM",
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the vLLM Inference Engine on RunPod Serverless", "description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by vLLM",
"type": "serverless", "type": "serverless",
"category": "language", "category": "language",
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png", "iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
"config": { "config": {
"runsOn": "GPU", "runsOn": "GPU",
"containerDiskInGb": 200, "containerDiskInGb": 150,
"gpuIds": "ADA_80_PRO, AMPERE_80", "gpuIds": "ADA_80_PRO,AMPERE_80",
"gpuCount": 1, "gpuCount": 1,
"allowedCudaVersions": [ "allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5", "12.4"],
"12.9",
"12.8",
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
],
"presets": [ "presets": [
{ {
"name": "deepseek-ai/deepseek-r1-distill-llama-8b", "name": "deepseek-ai/deepseek-r1-distill-llama-8b",
@@ -38,16 +28,6 @@
"required": true "required": true
} }
}, },
{
"key": "HF_TOKEN",
"input": {
"name": "Access Token",
"type": "string",
"description": "Hugging Face access token for gated & private models",
"default": "",
"required": false
}
},
{ {
"key": "TOKENIZER", "key": "TOKENIZER",
"input": { "input": {
@@ -945,7 +925,17 @@
"name": "Max Concurrency", "name": "Max Concurrency",
"type": "number", "type": "number",
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency", "description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
"default": 300, "default": 30,
"advanced": true
}
},
{
"key": "ENABLE_EXPERT_PARALLEL",
"input": {
"name": "Enable Expert Parallel",
"type": "boolean",
"description": "Enable Expert Parallel for MoE models",
"default": false,
"advanced": true "advanced": true
} }
}, },
@@ -1023,6 +1013,23 @@
"default": "", "default": "",
"advanced": true "advanced": true
} }
},
{
"key": "REASONING_PARSER",
"input": {
"name": "Reasoning Parser",
"type": "string",
"description": "Parser for reasoning-capable models (enables reasoning mode)",
"options": [
{ "label": "None", "value": "" },
{ "label": "DeepSeek R1", "value": "deepseek_r1" },
{ "label": "Qwen3", "value": "qwen3" },
{ "label": "Granite", "value": "granite" },
{ "label": "Hunyuan A13B", "value": "hunyuan_a13b" }
],
"default": "",
"advanced": true
}
} }
] ]
} }
+1 -9
View File
@@ -38,14 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct" "value": "HuggingFaceTB/SmolLM2-135M-Instruct"
} }
], ],
"allowedCudaVersions": [ "allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
]
} }
} }
+4 -5
View File
@@ -1,9 +1,9 @@
FROM nvidia/cuda:12.1.0-base-ubuntu22.04 FROM nvidia/cuda:12.4.1-base-ubuntu22.04
RUN apt-get update -y \ RUN apt-get update -y \
&& apt-get install -y python3-pip && apt-get install -y python3-pip
RUN ldconfig /usr/local/cuda-12.1/compat/ RUN ldconfig /usr/local/cuda-12.4/compat/
# Install Python dependencies # Install Python dependencies
COPY builder/requirements.txt /requirements.txt COPY builder/requirements.txt /requirements.txt
@@ -11,9 +11,8 @@ RUN --mount=type=cache,target=/root/.cache/pip \
python3 -m pip install --upgrade pip && \ python3 -m pip install --upgrade pip && \
python3 -m pip install --upgrade -r /requirements.txt python3 -m pip install --upgrade -r /requirements.txt
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer # Install vLLM
RUN python3 -m pip install vllm==0.10.0 && \ RUN python3 -m pip install vllm==0.11.0
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
# Setup for Option 2: Building the Image with the Model included # Setup for Option 2: Building the Image with the Model included
ARG MODEL_NAME="" ARG MODEL_NAME=""
+1 -1
View File
@@ -57,7 +57,7 @@ Configure worker-vllm using environment variables:
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) | | `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. | | `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String | | `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer | | `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)** For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
+2 -2
View File
@@ -1,14 +1,14 @@
ray ray
pandas pandas
pyarrow pyarrow
runpod~=1.7.7 runpod>=1.8,<2.0
huggingface-hub huggingface-hub
packaging packaging
typing-extensions>=4.8.0 typing-extensions>=4.8.0
pydantic pydantic
pydantic-settings pydantic-settings
hf-transfer hf-transfer
transformers>=4.55.0 transformers>=4.57.0
bitsandbytes>=0.45.0 bitsandbytes>=0.45.0
kernels kernels
torch==2.6.0 torch==2.6.0
+3 -1
View File
@@ -85,6 +85,7 @@ Complete guide to all environment variables and configuration options for worker
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. | | `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. | | `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. | | `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models |
## Tokenizer Settings ## Tokenizer Settings
@@ -113,12 +114,13 @@ The way this works is that the first request will have a batch size of `DEFAULT_
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. | | `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. | | `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` | | `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
## Serverless & Concurrency Settings ## Serverless & Concurrency Settings
| Variable | Default | Type/Choices | Description | | Variable | Default | Type/Choices | Description |
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency | | `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. | | `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. | | `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
+3 -2
View File
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
- `src/engine_args.py`: Centralized configuration management - `src/engine_args.py`: Centralized configuration management
- `src/constants.py`: Default values for core settings - `src/constants.py`: Default values for core settings
- `worker-config.json`: UI form generation for RunPod console - `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
- `worker-config.json`: UI form generation for RunPod console (if exists)
## Core Development Concepts ## Core Development Concepts
@@ -222,7 +223,7 @@ src/
### 2. **Concurrency Patterns** ### 2. **Concurrency Patterns**
- **Max Concurrency**: 300 concurrent requests by default - **Max Concurrency**: 30 concurrent requests by default
- **vLLM Queuing**: Internal request batching and scheduling - **vLLM Queuing**: Internal request batching and scheduling
- **RunPod Integration**: Concurrency modifier for auto-scaling - **RunPod Integration**: Concurrency modifier for auto-scaling
+1 -1
View File
@@ -1,4 +1,4 @@
DEFAULT_BATCH_SIZE = 50 DEFAULT_BATCH_SIZE = 50
DEFAULT_MAX_CONCURRENCY = 300 DEFAULT_MAX_CONCURRENCY = 30
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3 DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
DEFAULT_MIN_BATCH_SIZE = 1 DEFAULT_MIN_BATCH_SIZE = 1
+1
View File
@@ -80,6 +80,7 @@ DEFAULT_ARGS = {
"guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'), "guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'),
"speculative_model": os.getenv('SPECULATIVE_MODEL', None), "speculative_model": os.getenv('SPECULATIVE_MODEL', None),
"speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None, "speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None,
"enable_expert_parallel": bool(os.getenv('ENABLE_EXPERT_PARALLEL', 'False').lower() == 'true'),
"num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None, "num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None,
"speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None, "speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None,
"speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None, "speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None,
-1514
View File
File diff suppressed because it is too large Load Diff