Compare commits

...
26 Commits
Author SHA1 Message Date
c45ac42acd vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes

* MAX_NUM_BATCHED_TOKENS fix and CUDA tester

* Sys kill worker instead of marking as failed

* upgrade to vllm 0.12.0

* Update to vllm 0.15.0 and lora fix

* Update for HUB and removal of deprected env variables

* reverted docker-bake changes

* removed leftovers

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/utils.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Clean up of docs and comments in code

* nit: lowercase p

* nit: lowercase p

---------

Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
2026-02-12 21:50:34 +01:00
Tim PietruskyandGitHub 6d6cbe7095 fix: deactivate RunPod tests to fix hub release (#253)
Release / release (push) Waiting to run
Rename tests.json to tests_json to temporarily disable automated
tests while fixing the release on the hub.
2026-01-22 18:06:36 +01:00
90c16b472d fix: update CUDA to 12.4.1 for Blackwell GPU support (#251)
Release / release (push) Waiting to run
* fix: update CUDA to 12.4.1 for Blackwell GPU support

- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json

This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.

Fixes: DR-1118

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* revert: remove NVIDIA B200 from default gpuIds

The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove FlashInfer to avoid JIT compilation errors

FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.

vLLM will use its built-in fallback sampling methods instead.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-13 22:01:36 +01:00
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
chrisvelaandGitHub 3851d53f93 add ENABLE_EXPERT_PARALLEL engine arg for MoE models (#239)
Release / release (push) Waiting to run
* enable expert parallel arg for moe models

* add ENABLE_EXPERT_PARALLEL to hub config
2025-11-17 19:25:19 +01:00
Witold WydmańskiandGitHub c896438f21 feat: bump transformers to allow Qwen3-VL (#225)
Release / release (push) Waiting to run
2025-11-14 17:23:34 +01:00
Tim PietruskyandGitHub 912892f94e fix: remove space from gpuIds (#234) 2025-11-14 17:23:09 +01:00
Tim PietruskyandGitHub f8bf82469c fix(config): update allowed cuda versions in hub and tests config (#236)
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment

- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions

refs: AE-1452
2025-11-14 17:22:43 +01:00
Hailong YangandGitHub ec1664902b Merge pull request #230 from runpod-workers/feat/cse-853-vllm-template-params
Feat/cse 853 vllm template params
2025-10-31 13:32:34 -04:00
Eugene Klitenik d09122de4a remove un-needed 2025-10-29 13:06:33 -04:00
Eugene Klitenik e27dc68dea remove uneeded 2025-10-29 13:05:26 -04:00
Eugene Klitenik 1ee18d06a9 determine num gpus in python 2025-10-29 11:31:42 -04:00
Eugene Klitenik 5c4edd15cc update entrypoint command 2025-10-28 18:07:18 -04:00
Eugene Klitenik b074d3a23b auto detect num GPUs 2025-10-28 14:39:55 -04:00
Eugene Klitenik 205847471c reduce default container disk size to 150GB 2025-10-28 13:38:16 -04:00
Tim PietruskyandGitHub 6337a6673a fix: allow also CUDA 12.8 & 12.9 (#228)
Release / release (push) Waiting to run
2025-10-24 18:48:26 +02:00
Tim PietruskyandGitHub 66e1b1605b Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
2025-10-22 22:54:21 +02:00
Tim PietruskyandGitHub 60c8f257a8 Merge pull request #227 from runpod-workers/chore/vllm-0.11.0
chore: update vllm to 0.11.0
2025-10-22 22:53:51 +02:00
Tim Pietrusky fae16e7ee1 chore: update vllm to 0.11.0 2025-10-22 13:40:53 -07:00
max4c 2becd35345 Revert "fix: added back the HF_TOKEN (#219)"
Release / release (push) Waiting to run
This reverts commit 33d88df6c0.
2025-09-23 12:38:24 -07:00
33d88df6c0 fix: added back the HF_TOKEN (#219)
Release / release (push) Waiting to run
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-23 19:25:12 +02:00
Tim Pietrusky ecd562e112 fix: remove "access token" as this is handled by the platform
Release / release (push) Waiting to run
2025-09-19 21:00:47 +02:00
5cffaab8e8 docs: how to use the reasoning parser (#218)
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-17 07:49:37 +02:00
a0fe1dfdad feat: better hub support & concise README for the main repo (#215)
Release / release (push) Waiting to run
* feat: moved config into docs; added banner; auto detect "messages" in input

* docs: moved config into docs

* chore: added .DS_Store

* chore: get the original stuff working again

* chore: remove all changes

* docs: reduced toc and added small config table

---------

Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-01 16:47:48 +02:00
Tim Pietrusky 0e0d6df859 docs: updated example-tag for dev and release 2025-08-28 14:28:19 +02:00
Tim Pietrusky d1718aec00 ci: removed "github release" step as that is not needed 2025-08-28 14:27:50 +02:00
19 changed files with 681 additions and 2056 deletions
+2 -2
View File
@@ -12,7 +12,7 @@
git push origin feature/your-feature-name git push origin feature/your-feature-name
``` ```
- Creates pull request → triggers dev build: `runpod/worker-v1-vllm:dev-feature-your-feature-name` - Creates pull request → triggers dev build: `runpod/worker-v1-vllm:dev-refs-pull-214-merge`
2. **Main Branch** 2. **Main Branch**
```bash ```bash
@@ -74,6 +74,6 @@ See [README.md](../README.md) for full list of supported environment variables.
## 🔧 CI/CD Workflows ## 🔧 CI/CD Workflows
- **Dev builds**: All pull requests → `dev-<branch-name>` images - **Dev builds**: All pull requests → `dev-refs-pull-<PR#>-merge` images
- **Release builds**: Git tags → versioned images + GitHub releases - **Release builds**: Git tags → versioned images + GitHub releases
- **Manual triggers**: Available in GitHub Actions for emergency releases - **Manual triggers**: Available in GitHub Actions for emergency releases
+7 -23
View File
@@ -13,7 +13,6 @@ on:
permissions: permissions:
contents: write # Required for creating GitHub releases contents: write # Required for creating GitHub releases
packages: write # Required for pushing Docker images (if using GitHub packages)
jobs: jobs:
release: release:
@@ -75,28 +74,13 @@ jobs:
*.args.RELEASE_VERSION=${{ env.RELEASE_VERSION }} *.args.RELEASE_VERSION=${{ env.RELEASE_VERSION }}
*.args.HUGGINGFACE_ACCESS_TOKEN=${{ env.HUGGINGFACE_ACCESS_TOKEN }} *.args.HUGGINGFACE_ACCESS_TOKEN=${{ env.HUGGINGFACE_ACCESS_TOKEN }}
- name: Create GitHub Release - name: Release Summary
if: env.IS_MANUAL_RELEASE == 'false'
uses: actions/create-release@v1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
with:
tag_name: ${{ github.ref_name }}
release_name: Release ${{ github.ref_name }}
body: |
Release ${{ github.ref_name }}
Docker Image: `${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}`
## Changes
See [commit history](https://github.com/${{ github.repository }}/commits/${{ github.ref_name }}) for detailed changes.
draft: false
prerelease: false
- name: Manual Release Summary
if: env.IS_MANUAL_RELEASE == 'true'
run: | run: |
echo "🚀 Manual release completed!" echo "🚀 Release completed!"
echo "Version: ${{ env.RELEASE_VERSION }}" echo "Version: ${{ env.RELEASE_VERSION }}"
echo "Docker Image: ${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}" echo "Docker Image: ${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}"
echo "Note: No GitHub release created for manual triggers" if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
echo "Trigger: Manual workflow dispatch"
else
echo "Trigger: GitHub release (tag: ${{ github.ref_name }})"
fi
+1
View File
@@ -4,3 +4,4 @@ runpod.toml
.env .env
test/* test/*
vllm-base/vllm-* vllm-base/vllm-*
.DS_Store
+261
View File
@@ -0,0 +1,261 @@
![vLLM worker banner](https://cpjrphpz3t5wbwfe.public.blob.vercel-storage.com/worker-vllm_banner.jpeg)
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
---
[![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
---
## Endpoint Configuration
All behaviour is controlled through environment variables:
| Environment Variable | Description | Default | Options |
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
| `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
## API Usage
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
### RunPod Native API
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
#### Chat Completions
```json
{
"input": {
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "What is the capital of France?" }
],
"sampling_params": {
"max_tokens": 100,
"temperature": 0.7
}
}
}
```
#### Chat Completions (Streaming)
```json
{
"input": {
"messages": [
{ "role": "user", "content": "Write a short story about a robot." }
],
"sampling_params": {
"max_tokens": 500,
"temperature": 0.8
},
"stream": true
}
}
```
#### Text Generation
For direct text generation without chat format:
```json
{
"input": {
"prompt": "The capital of France is",
"sampling_params": {
"max_tokens": 64,
"temperature": 0.0
}
}
}
```
#### List Models
```json
{
"input": {
"openai_route": "/v1/models"
}
}
```
---
### OpenAI-Compatible API
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
#### Chat Completions
**Path:** `/openai/v1/chat/completions`
```json
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "What is the capital of France?" }
],
"max_tokens": 100,
"temperature": 0.7
}
```
#### Chat Completions (Streaming)
```json
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"messages": [
{ "role": "user", "content": "Write a short story about a robot." }
],
"max_tokens": 500,
"temperature": 0.8,
"stream": true
}
```
#### Text Completions
**Path:** `/openai/v1/completions`
```json
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"prompt": "The capital of France is",
"max_tokens": 100,
"temperature": 0.7
}
```
#### List Models
**Path:** `/openai/v1/models`
```json
{}
```
#### Response Format
Both APIs return the same response format:
```json
{
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Paris." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
}
```
---
## Usage
Below are minimal `python` snippets so you can copy-paste to get started quickly.
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
### OpenAI compatible API
Minimal Python example using the official `openai` SDK:
```python
from openai import OpenAI
import os
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
client = OpenAI(
api_key=os.getenv("RUNPOD_API_KEY"),
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
)
```
`Chat Completions (Non-Streaming)`
```python
response = client.chat.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
temperature=0,
max_tokens=100,
)
print(f"Response: {response.choices[0].message.content}")
```
`Chat Completions (Streaming)`
```python
response_stream = client.chat.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
temperature=0,
max_tokens=100,
stream=True
)
for response in response_stream:
print(response.choices[0].delta.content or "", end="", flush=True)
```
### RunPod Native API
```python
import requests
response = requests.post(
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
headers={"Authorization": "Bearer <API_KEY>"},
json={
"input": {
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in simple terms"}
],
"sampling_params": {
"temperature": 0.7,
"max_tokens": 150
}
}
}
)
result = response.json()
print(result["output"])
```
## Compatibility
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
## Documentation
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
+35 -312
View File
@@ -1,25 +1,15 @@
{ {
"title": "vLLM", "title": "vLLM",
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the vLLM Inference Engine on RunPod Serverless", "description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by vLLM",
"type": "serverless", "type": "serverless",
"category": "language", "category": "language",
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png", "iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
"config": { "config": {
"runsOn": "GPU", "runsOn": "GPU",
"containerDiskInGb": 200, "containerDiskInGb": 150,
"gpuIds": "ADA_80_PRO, AMPERE_80", "gpuIds": "ADA_80_PRO,AMPERE_80",
"gpuCount": 1, "gpuCount": 1,
"allowedCudaVersions": [ "allowedCudaVersions": ["12.9", "12.8"],
"12.9",
"12.8",
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
],
"presets": [ "presets": [
{ {
"name": "deepseek-ai/deepseek-r1-distill-llama-8b", "name": "deepseek-ai/deepseek-r1-distill-llama-8b",
@@ -38,16 +28,6 @@
"required": true "required": true
} }
}, },
{
"key": "HF_TOKEN",
"input": {
"name": "Access Token",
"type": "string",
"description": "Hugging Face access token for gated & private models",
"default": "",
"required": false
}
},
{ {
"key": "TOKENIZER", "key": "TOKENIZER",
"input": { "input": {
@@ -201,15 +181,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "QUANTIZATION_PARAM_PATH",
"input": {
"name": "Quantization Param Path",
"type": "string",
"description": "Path to the JSON file containing the KV cache scaling factors.",
"advanced": true
}
},
{ {
"key": "MAX_MODEL_LEN", "key": "MAX_MODEL_LEN",
"input": { "input": {
@@ -219,26 +190,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "GUIDED_DECODING_BACKEND",
"input": {
"name": "Guided Decoding Backend",
"type": "string",
"description": "Which engine will be used for guided decoding by default.",
"options": [
{
"label": "outlines",
"value": "outlines"
},
{
"label": "lm-format-enforcer",
"value": "lm-format-enforcer"
}
],
"default": "outlines",
"advanced": true
}
},
{ {
"key": "DISTRIBUTED_EXECUTOR_BACKEND", "key": "DISTRIBUTED_EXECUTOR_BACKEND",
"input": { "input": {
@@ -258,16 +209,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "WORKER_USE_RAY",
"input": {
"name": "Worker Use Ray",
"type": "boolean",
"description": "Deprecated, use --distributed-executor-backend=ray.",
"default": false,
"advanced": true
}
},
{ {
"key": "RAY_WORKERS_USE_NSIGHT", "key": "RAY_WORKERS_USE_NSIGHT",
"input": { "input": {
@@ -327,26 +268,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "USE_V2_BLOCK_MANAGER",
"input": {
"name": "Use V2 Block Manager",
"type": "boolean",
"description": "Use BlockSpaceMangerV2.",
"default": false,
"advanced": true
}
},
{
"key": "NUM_LOOKAHEAD_SLOTS",
"input": {
"name": "Num Lookahead Slots",
"type": "number",
"description": "Experimental scheduling config necessary for speculative decoding.",
"default": 0,
"advanced": true
}
},
{ {
"key": "SEED", "key": "SEED",
"input": { "input": {
@@ -432,53 +353,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "ROPE_SCALING",
"input": {
"name": "RoPE Scaling",
"type": "string",
"description": "RoPE scaling configuration in JSON format.",
"advanced": true
}
},
{
"key": "ROPE_THETA",
"input": {
"name": "RoPE Theta",
"type": "number",
"description": "RoPE theta. Use with rope_scaling.",
"advanced": true
}
},
{
"key": "TOKENIZER_POOL_SIZE",
"input": {
"name": "Tokenizer Pool Size",
"type": "number",
"description": "Size of tokenizer pool to use for asynchronous tokenization.",
"default": 0,
"advanced": true
}
},
{
"key": "TOKENIZER_POOL_TYPE",
"input": {
"name": "Tokenizer Pool Type",
"type": "string",
"description": "Type of tokenizer pool to use for asynchronous tokenization.",
"default": "ray",
"advanced": true
}
},
{
"key": "TOKENIZER_POOL_EXTRA_CONFIG",
"input": {
"name": "Tokenizer Pool Extra Config",
"type": "string",
"description": "Extra config for tokenizer pool.",
"advanced": true
}
},
{ {
"key": "ENABLE_LORA", "key": "ENABLE_LORA",
"input": { "input": {
@@ -509,16 +383,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "LORA_EXTRA_VOCAB_SIZE",
"input": {
"name": "LoRA Extra Vocab Size",
"type": "number",
"description": "Maximum size of extra vocabulary for LoRA adapters.",
"default": 256,
"advanced": true
}
},
{ {
"key": "LORA_DTYPE", "key": "LORA_DTYPE",
"input": { "input": {
@@ -547,15 +411,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "LONG_LORA_SCALING_FACTORS",
"input": {
"name": "Long LoRA Scaling Factors",
"type": "string",
"description": "Specify multiple scaling factors for LoRA adapters.",
"advanced": true
}
},
{ {
"key": "MAX_CPU_LORAS", "key": "MAX_CPU_LORAS",
"input": { "input": {
@@ -635,107 +490,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "SPECULATIVE_MODEL",
"input": {
"name": "Speculative Model",
"type": "string",
"description": "The name of the draft model to be used in speculative decoding.",
"advanced": true
}
},
{
"key": "NUM_SPECULATIVE_TOKENS",
"input": {
"name": "Num Speculative Tokens",
"type": "number",
"description": "The number of speculative tokens to sample from the draft model.",
"advanced": true
}
},
{
"key": "SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE",
"input": {
"name": "Speculative Draft Tensor Parallel Size",
"type": "number",
"description": "Number of tensor parallel replicas for the draft model.",
"advanced": true
}
},
{
"key": "SPECULATIVE_MAX_MODEL_LEN",
"input": {
"name": "Speculative Max Model Length",
"type": "number",
"description": "The maximum sequence length supported by the draft model.",
"advanced": true
}
},
{
"key": "SPECULATIVE_DISABLE_BY_BATCH_SIZE",
"input": {
"name": "Speculative Disable by Batch Size",
"type": "number",
"description": "Disable speculative decoding if the number of enqueue requests is larger than this value.",
"advanced": true
}
},
{
"key": "NGRAM_PROMPT_LOOKUP_MAX",
"input": {
"name": "Ngram Prompt Lookup Max",
"type": "number",
"description": "Max size of window for ngram prompt lookup in speculative decoding.",
"advanced": true
}
},
{
"key": "NGRAM_PROMPT_LOOKUP_MIN",
"input": {
"name": "Ngram Prompt Lookup Min",
"type": "number",
"description": "Min size of window for ngram prompt lookup in speculative decoding.",
"advanced": true
}
},
{
"key": "SPEC_DECODING_ACCEPTANCE_METHOD",
"input": {
"name": "Speculative Decoding Acceptance Method",
"type": "string",
"description": "Specify the acceptance method for draft token verification in speculative decoding.",
"options": [
{
"label": "rejection_sampler",
"value": "rejection_sampler"
},
{
"label": "typical_acceptance_sampler",
"value": "typical_acceptance_sampler"
}
],
"default": "rejection_sampler",
"advanced": true
}
},
{
"key": "TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD",
"input": {
"name": "Typical Acceptance Sampler Posterior Threshold",
"type": "number",
"description": "Set the lower bound threshold for the posterior probability of a token to be accepted.",
"advanced": true
}
},
{
"key": "TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA",
"input": {
"name": "Typical Acceptance Sampler Posterior Alpha",
"type": "number",
"description": "A scaling factor for the entropy-based threshold for token acceptance.",
"advanced": true
}
},
{ {
"key": "MODEL_LOADER_EXTRA_CONFIG", "key": "MODEL_LOADER_EXTRA_CONFIG",
"input": { "input": {
@@ -746,49 +500,11 @@
} }
}, },
{ {
"key": "PREEMPTION_MODE", "key": "ENABLE_LOG_REQUESTS",
"input": { "input": {
"name": "Preemption Mode", "name": "Enable Log Requests",
"type": "string",
"description": "If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens.",
"advanced": true
}
},
{
"key": "PREEMPTION_CHECK_PERIOD",
"input": {
"name": "Preemption Check Period",
"type": "number",
"description": "How frequently the engine checks if a preemption happens.",
"default": 1,
"advanced": true
}
},
{
"key": "PREEMPTION_CPU_CAPACITY",
"input": {
"name": "Preemption CPU Capacity",
"type": "number",
"description": "The percentage of CPU memory used for the saved activations.",
"default": 2,
"advanced": true
}
},
{
"key": "MAX_LOG_LEN",
"input": {
"name": "Max Log Length",
"type": "number",
"description": "Max number of characters or ID numbers being printed in log.",
"advanced": true
}
},
{
"key": "DISABLE_LOGGING_REQUEST",
"input": {
"name": "Disable Logging Request",
"type": "boolean", "type": "boolean",
"description": "Disable logging requests.", "description": "Enable vLLM request logging.",
"default": false, "default": false,
"advanced": true "advanced": true
} }
@@ -860,16 +576,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "MAX_SEQ_LEN_TO_CAPTURE",
"input": {
"name": "CUDA Graph Max Content Length",
"type": "number",
"description": "Maximum context length covered by CUDA graphs. If a sequence has context length larger than this, we fall back to eager mode",
"default": 8192,
"advanced": true
}
},
{ {
"key": "DISABLE_CUSTOM_ALL_REDUCE", "key": "DISABLE_CUSTOM_ALL_REDUCE",
"input": { "input": {
@@ -945,7 +651,17 @@
"name": "Max Concurrency", "name": "Max Concurrency",
"type": "number", "type": "number",
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency", "description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
"default": 300, "default": 30,
"advanced": true
}
},
{
"key": "ENABLE_EXPERT_PARALLEL",
"input": {
"name": "Enable Expert Parallel",
"type": "boolean",
"description": "Enable Expert Parallel for MoE models",
"default": false,
"advanced": true "advanced": true
} }
}, },
@@ -968,16 +684,6 @@
"advanced": true "advanced": true
} }
}, },
{
"key": "DISABLE_LOG_REQUESTS",
"input": {
"name": "Disable Log Requests",
"type": "boolean",
"description": "Enables or disables vLLM request logging",
"default": true,
"advanced": true
}
},
{ {
"key": "ENABLE_AUTO_TOOL_CHOICE", "key": "ENABLE_AUTO_TOOL_CHOICE",
"input": { "input": {
@@ -1023,6 +729,23 @@
"default": "", "default": "",
"advanced": true "advanced": true
} }
},
{
"key": "REASONING_PARSER",
"input": {
"name": "Reasoning Parser",
"type": "string",
"description": "Parser for reasoning-capable models (enables reasoning mode)",
"options": [
{ "label": "None", "value": "" },
{ "label": "DeepSeek R1", "value": "deepseek_r1" },
{ "label": "Qwen3", "value": "qwen3" },
{ "label": "Granite", "value": "granite" },
{ "label": "Hunyuan A13B", "value": "hunyuan_a13b" }
],
"default": "",
"advanced": true
}
} }
] ]
} }
+3 -12
View File
@@ -12,7 +12,6 @@
"input": { "input": {
"openai_route": "/v1/chat/completions", "openai_route": "/v1/chat/completions",
"openai_input": { "openai_input": {
"model": "HuggingFaceTB/SmolLM2-135M-Instruct",
"messages": [ "messages": [
{ {
"role": "system", "role": "system",
@@ -23,8 +22,8 @@
"content": "Explain what a neural network is in one sentence." "content": "Explain what a neural network is in one sentence."
} }
], ],
"max_tokens": 50, "max_tokens": 200,
"temperature": 0.7 "temperature": 0.1
} }
}, },
"timeout": 30000 "timeout": 30000
@@ -39,14 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct" "value": "HuggingFaceTB/SmolLM2-135M-Instruct"
} }
], ],
"allowedCudaVersions": [ "allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
"12.7",
"12.6",
"12.5",
"12.4",
"12.3",
"12.2",
"12.1"
]
} }
} }
+17 -9
View File
@@ -1,20 +1,21 @@
FROM nvidia/cuda:12.1.0-base-ubuntu22.04 FROM nvidia/cuda:12.8.0-base-ubuntu22.04
RUN apt-get update -y \ RUN apt-get update -y \
&& apt-get install -y python3-pip && apt-get install -y python3-pip
RUN ldconfig /usr/local/cuda-12.1/compat/ RUN ldconfig /usr/local/cuda-12.8/compat/
# Install Python dependencies # Install vLLM with FlashInfer - use CUDA 12.8 PyTorch wheels (compatible with vLLM 0.15.0)
RUN python3 -m pip install --upgrade pip && \
python3 -m pip install "vllm[flashinfer]==0.15.0" --extra-index-url https://download.pytorch.org/whl/cu128
# Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts)
COPY builder/requirements.txt /requirements.txt COPY builder/requirements.txt /requirements.txt
RUN --mount=type=cache,target=/root/.cache/pip \ RUN --mount=type=cache,target=/root/.cache/pip \
python3 -m pip install --upgrade pip && \
python3 -m pip install --upgrade -r /requirements.txt python3 -m pip install --upgrade -r /requirements.txt
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
RUN python3 -m pip install vllm==0.10.0 && \
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
# Setup for Option 2: Building the Image with the Model included # Setup for Option 2: Building the Image with the Model included
ARG MODEL_NAME="" ARG MODEL_NAME=""
ARG TOKENIZER_NAME="" ARG TOKENIZER_NAME=""
@@ -32,7 +33,14 @@ ENV MODEL_NAME=$MODEL_NAME \
HF_DATASETS_CACHE="${BASE_PATH}/huggingface-cache/datasets" \ HF_DATASETS_CACHE="${BASE_PATH}/huggingface-cache/datasets" \
HUGGINGFACE_HUB_CACHE="${BASE_PATH}/huggingface-cache/hub" \ HUGGINGFACE_HUB_CACHE="${BASE_PATH}/huggingface-cache/hub" \
HF_HOME="${BASE_PATH}/huggingface-cache/hub" \ HF_HOME="${BASE_PATH}/huggingface-cache/hub" \
HF_HUB_ENABLE_HF_TRANSFER=0 HF_HUB_ENABLE_HF_TRANSFER=0 \
# Suppress Ray metrics agent warnings (not needed in containerized environments)
RAY_METRICS_EXPORT_ENABLED=0 \
RAY_DISABLE_USAGE_STATS=1 \
# Prevent rayon thread pool panic in containers where ulimit -u < nproc
# (tokenizers uses Rust's rayon which tries to spawn threads = CPU cores)
TOKENIZERS_PARALLELISM=false \
RAYON_NUM_THREADS=4
ENV PYTHONPATH="/:/vllm-workspace" ENV PYTHONPATH="/:/vllm-workspace"
+27 -129
View File
@@ -9,16 +9,10 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
## Table of Contents ## Table of Contents
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker) - [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
- [Option 1: Deploy Any Model Using Pre-Built Docker Image **[RECOMMENDED]**](#option-1-deploy-any-model-using-pre-built-docker-image-recommended) - [Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended]](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
- [Environment Variables](#environment-variables) - [Configuration](#configuration)
- [LLM Settings](#llm-settings)
- [Tokenizer Settings](#tokenizer-settings)
- [System and Parallelism Settings](#system-and-parallelism-settings)
- [Streaming Batch Size Settings](#streaming-batch-size-settings)
- [OpenAI Settings](#openai-settings)
- [Serverless Settings](#serverless-settings)
- [Option 2: Build Docker Image with Model Inside](#option-2-build-docker-image-with-model-inside) - [Option 2: Build Docker Image with Model Inside](#option-2-build-docker-image-with-model-inside)
- [Prerequisites](#prerequisites-1) - [Prerequisites](#prerequisites)
- [Arguments](#arguments) - [Arguments](#arguments)
- [Example: Building an image with OpenChat-3.5](#example-building-an-image-with-openchat-35) - [Example: Building an image with OpenChat-3.5](#example-building-an-image-with-openchat-35)
- [(Optional) Including Huggingface Token](#optional-including-huggingface-token) - [(Optional) Including Huggingface Token](#optional-including-huggingface-token)
@@ -26,12 +20,14 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
- [Usage: OpenAI Compatibility](#usage-openai-compatibility) - [Usage: OpenAI Compatibility](#usage-openai-compatibility)
- [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker) - [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker)
- [OpenAI Request Input Parameters](#openai-request-input-parameters) - [OpenAI Request Input Parameters](#openai-request-input-parameters)
- [Chat Completions](#chat-completions) - [Chat Completions [RECOMMENDED]](#chat-completions-recommended)
- [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai) - [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai)
- [Usage: standard](#non-openai-usage) - [Chat Completions](#chat-completions)
- [Input Request Parameters](#input-request-parameters) - [Getting a list of names for available models](#getting-a-list-of-names-for-available-models)
- [Text Input Formats](#text-input-formats) - [Usage: Standard (Non-OpenAI)](#usage-standard-non-openai)
- [Request Input Parameters](#request-input-parameters)
- [Sampling Parameters](#sampling-parameters) - [Sampling Parameters](#sampling-parameters)
- [Text Input Formats](#text-input-formats)
# Setting up the Serverless Worker # Setting up the Serverless Worker
@@ -39,129 +35,31 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
**🚀 Deploy Guide**: Follow our [step-by-step deployment guide](https://docs.runpod.io/serverless/vllm/get-started) to deploy using the RunPod Console. **🚀 Deploy Guide**: Follow our [step-by-step deployment guide](https://docs.runpod.io/serverless/vllm/get-started) to deploy using the RunPod Console.
**📦 Docker Image**: `runpod/worker-v1-vllm:<version>stable-cuda12.1.0` **📦 Docker Image**: `runpod/worker-v1-vllm:<version>`
- **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases) - **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases)
- **CUDA Compatibility**: Requires CUDA >= 12.1 - **CUDA Compatibility**: Requires CUDA >= 12.1
### Environment Variables ### Configuration
Use these to configure worker-vllm so it works for your use case / model. Configure worker-vllm using environment variables:
#### LLM Settings | Environment Variable | Description | Default | Options |
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
| `Name` | `Default` | `Type/Choices` | `Description` | For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
| ------------------------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
| `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. |
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
| `SEED` | 0 | `int` | Random seed for operations. |
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}` |
| `SCHEDULER_DELAY_FACTOR` | 0.0 | `float` | Apply a delay before scheduling next prompt. |
| `ENABLE_CHUNKED_PREFILL` | False | `bool` | Enable chunked prefill requests. |
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
| `SPEC_DECODING_ACCEPTANCE_METHOD` | 'rejection_sampler' | ['rejection_sampler', 'typical_acceptance_sampler'] | Specify the acceptance method for draft token verification in speculative decoding. |
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD` | None | `float` | Set the lower bound threshold for the posterior probability of a token to be accepted. |
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA` | None | `float` | A scaling factor for the entropy-based threshold for token acceptance. |
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
#### Tokenizer Settings
| `Name` | `Default` | `Type/Choices` | `Description` |
| ---------------------- | --------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
#### System and Parallelism Settings
| `Name` | `Default` | `Type/Choices` | `Description` |
| ------------------------------ | --------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
#### Streaming Batch Size Settings
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker
| `Name` | `Default` | `Type/Choices` | `Description` |
| ---------------------------------- | --------- | -------------- | --------------------------------------------------------------------------------------------------------- |
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
#### OpenAI Settings
| `Name` | `Default` | `Type/Choices` | `Description` |
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
#### Serverless Settings
| `Name` | `Default` | `Type/Choices` | `Description` |
| ---------------------- | --------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
## Option 2: Build Docker Image with Model Inside ## Option 2: Build Docker Image with Model Inside
+3 -3
View File
@@ -1,14 +1,14 @@
ray ray
pandas pandas
pyarrow pyarrow
runpod~=1.7.7 runpod
huggingface-hub huggingface-hub
packaging packaging
typing-extensions>=4.8.0 typing-extensions>=4.8.0
pydantic pydantic
pydantic-settings pydantic-settings
hf-transfer hf-transfer
transformers>=4.55.0 transformers>=4.57.0
bitsandbytes>=0.45.0 bitsandbytes>=0.45.0
kernels kernels
torch==2.6.0 torch-c-dlpack-ext
+169
View File
@@ -0,0 +1,169 @@
# Configuration Reference
Complete guide to all environment variables and configuration options for worker-vllm.
## LLM Settings
| Variable | Default | Type/Choices | Description |
| ------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------- |
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
| `MODEL_REVISION` | 'main' | `str` | Model revision to load (default: main). |
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
| `SEED` | 0 | `int` | Random seed for operations. |
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
## LoRA (Low-Rank Adaptation) Settings
| Variable | Default | Type | Description |
| --------------------------- | ------- | ------------------------------------------ | --------------------------------------------------------------------------------------------------------- |
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}]` |
> **Note (Serverless)**: When LoRA adapters are configured via `LORA_MODULES`, initialization is deferred to the first request to ensure compatibility with RunPod Serverless. This means the first request will include LoRA loading time. Subsequent requests are unaffected. Check logs for "LoRA mode: X adapter(s) will load on first request" at startup.
## Speculative Decoding Settings
| Variable | Default | Type/Choices | Description |
| ------------------------------------------------ | ------------------- | --------------------------------------------------- | ----------------------------------------------------------------------------------------- |
| `SCHEDULER_DELAY_FACTOR` | 0.0 | `float` | Apply a delay before scheduling next prompt. |
| `ENABLE_CHUNKED_PREFILL` | False | `bool` | Enable chunked prefill requests. |
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
| `SPEC_DECODING_ACCEPTANCE_METHOD` | 'rejection_sampler' | ['rejection_sampler', 'typical_acceptance_sampler'] | Specify the acceptance method for draft token verification in speculative decoding. |
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD` | None | `float` | Set the lower bound threshold for the posterior probability of a token to be accepted. |
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA` | None | `float` | A scaling factor for the entropy-based threshold for token acceptance. |
## System Performance Settings
| Variable | Default | Type/Choices | Description |
| ------------------------------ | ------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models. |
| `ATTENTION_BACKEND` | `None` | `str` | Attention backend to use (e.g., `FLASH_ATTN`, `FLASHINFER`, `TRITON_FLASH_ATTN`). Replaces deprecated `VLLM_ATTENTION_BACKEND`. |
| `ASYNC_SCHEDULING` | `None` | `bool` | Enable async scheduling (overlaps engine scheduling with GPU execution). Default: enabled in vLLM 0.14.0+. Set to `false` to disable. |
| `STREAM_INTERVAL` | `1` | `int` | Controls how often to yield streaming results. Lower = more frequent updates. |
## Tokenizer Settings
| Variable | Default | Type/Choices | Description |
| ---------------------- | ------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
## Streaming & Batch Settings
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker.
| Variable | Default | Type/Choices | Description |
| ---------------------------------- | ------- | ------------ | --------------------------------------------------------------------------------------------------------- |
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
## OpenAI Compatibility Settings
| Variable | Default | Type/Choices | Description |
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
| `TRUST_REQUEST_CHAT_TEMPLATE` | `false` | `bool` | Allow clients to send custom chat templates in API requests. **Security consideration:** Only enable if you trust your API clients. |
| `RETURN_TOKENS_AS_TOKEN_IDS` | `false` | `bool` | Return token IDs instead of decoded text strings in responses. |
| `EXCLUDE_TOOLS_WHEN_TOOL_CHOICE_NONE` | `false` | `bool` | Exclude tool definitions from the prompt when `tool_choice` is set to `none`. |
| `ENABLE_PROMPT_TOKENS_DETAILS` | `false` | `bool` | Include detailed prompt token information in API responses. |
| `ENABLE_FORCE_INCLUDE_USAGE` | `false` | `bool` | Always include usage statistics in API responses, even when not requested. |
| `ENABLE_LOG_OUTPUTS` | `false` | `bool` | Log model outputs for debugging purposes. |
| `LOG_ERROR_STACK` | `false` | `bool` | Include full stack traces in error responses for debugging. |
## Serverless & Concurrency Settings
| Variable | Default | Type/Choices | Description |
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
| `ENABLE_LOG_REQUESTS` | False | `bool` | Enables vLLM request logging. (Replaces deprecated `DISABLE_LOG_REQUESTS` in vLLM 0.15.0) |
## Advanced Settings
| Variable | Default | Type | Description |
| --------------------------- | ------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
## Docker Build Arguments
These variables are used when building custom Docker images with models baked in:
| Variable | Default | Type | Description |
| --------------------- | ---------------- | ----- | ------------------------------------------------- |
| `BASE_PATH` | `/runpod-volume` | `str` | Storage directory for huggingface cache and model |
| `WORKER_CUDA_VERSION` | `12.1.0` | `str` | CUDA version for the worker image |
## Deprecated Variables
⚠️ **The following variables are deprecated and will be removed in future versions:**
| Old Variable | New Variable | Note |
| ---------------------------- | ------------------------ | -------------------------------------------------------------------- |
| `MAX_CONTEXT_LEN_TO_CAPTURE` | `MAX_SEQ_LEN_TO_CAPTURE` | Use new variable name |
| `kv_cache_dtype=fp8_e5m2` | `kv_cache_dtype=fp8` | Simplified fp8 format |
| `USE_V2_BLOCK_MANAGER` | *(removed)* | V2 block manager is now the default in vLLM 0.13.0, setting ignored |
| `VLLM_ATTENTION_BACKEND` | `ATTENTION_BACKEND` | Use new env var name (old still works with deprecation warning) |
| `DISABLE_LOG_REQUESTS` | `ENABLE_LOG_REQUESTS` | Inverted logic in vLLM 0.15.0 (old still works with deprecation warning) |
+3 -2
View File
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
- `src/engine_args.py`: Centralized configuration management - `src/engine_args.py`: Centralized configuration management
- `src/constants.py`: Default values for core settings - `src/constants.py`: Default values for core settings
- `worker-config.json`: UI form generation for RunPod console - `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
- `worker-config.json`: UI form generation for RunPod console (if exists)
## Core Development Concepts ## Core Development Concepts
@@ -222,7 +223,7 @@ src/
### 2. **Concurrency Patterns** ### 2. **Concurrency Patterns**
- **Max Concurrency**: 300 concurrent requests by default - **Max Concurrency**: 30 concurrent requests by default
- **vLLM Queuing**: Internal request batching and scheduling - **vLLM Queuing**: Internal request batching and scheduling
- **RunPod Integration**: Concurrency modifier for auto-scaling - **RunPod Integration**: Concurrency modifier for auto-scaling
Binary file not shown.

After

Width:  |  Height:  |  Size: 86 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 MiB

+1 -1
View File
@@ -1,4 +1,4 @@
DEFAULT_BATCH_SIZE = 50 DEFAULT_BATCH_SIZE = 50
DEFAULT_MAX_CONCURRENCY = 300 DEFAULT_MAX_CONCURRENCY = 30
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3 DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
DEFAULT_MIN_BATCH_SIZE = 1 DEFAULT_MIN_BATCH_SIZE = 1
+63 -24
View File
@@ -1,24 +1,25 @@
import os
import logging
import json
import asyncio import asyncio
import json
import logging
import os
import time
from typing import AsyncGenerator, Optional
from dotenv import load_dotenv from dotenv import load_dotenv
from typing import AsyncGenerator, Optional
import time
from vllm import AsyncLLMEngine from vllm import AsyncLLMEngine
from vllm.entrypoints.logger import RequestLogger from vllm.entrypoints.logger import RequestLogger
from vllm.entrypoints.openai.serving_chat import OpenAIServingChat from vllm.entrypoints.openai.chat_completion.protocol import ChatCompletionRequest
from vllm.entrypoints.openai.serving_completion import OpenAIServingCompletion from vllm.entrypoints.openai.chat_completion.serving import OpenAIServingChat
from vllm.entrypoints.openai.protocol import ChatCompletionRequest, CompletionRequest, ErrorResponse from vllm.entrypoints.openai.completion.protocol import CompletionRequest
from vllm.entrypoints.openai.serving_models import BaseModelPath, LoRAModulePath, OpenAIServingModels from vllm.entrypoints.openai.completion.serving import OpenAIServingCompletion
from vllm.entrypoints.openai.engine.protocol import ErrorResponse
from vllm.entrypoints.openai.models.protocol import BaseModelPath, LoRAModulePath
from vllm.entrypoints.openai.models.serving import OpenAIServingModels
from constants import DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MAX_CONCURRENCY, DEFAULT_MIN_BATCH_SIZE
from utils import DummyRequest, JobInput, BatchSize, create_error_response
from constants import DEFAULT_MAX_CONCURRENCY, DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MIN_BATCH_SIZE
from tokenizer import TokenizerWrapper
from engine_args import get_engine_args from engine_args import get_engine_args
from tokenizer import TokenizerWrapper
from utils import BatchSize, DummyRequest, JobInput, create_error_response
class vLLMEngine: class vLLMEngine:
def __init__(self, engine = None): def __init__(self, engine = None):
@@ -177,7 +178,21 @@ class OpenAIvLLMEngine(vLLMEngine):
self.served_model_name = os.getenv("OPENAI_SERVED_MODEL_NAME_OVERRIDE") or self.engine_args.model self.served_model_name = os.getenv("OPENAI_SERVED_MODEL_NAME_OVERRIDE") or self.engine_args.model
self.response_role = os.getenv("OPENAI_RESPONSE_ROLE") or "assistant" self.response_role = os.getenv("OPENAI_RESPONSE_ROLE") or "assistant"
self.lora_adapters = self._load_lora_adapters() self.lora_adapters = self._load_lora_adapters()
asyncio.run(self._initialize_engines())
# Always defer OpenAI engine initialization to the first request.
# asyncio.run() creates a temporary event loop that gets closed, but async
# components (tokenizer pool, serving engines) bind futures to that loop.
# When Runpod's serverless handler runs in its own event loop, those futures
# are "attached to a different loop" causing RuntimeError.
# This affects all configurations, not just LoRA.
self._engines_initialized = False
if self.lora_adapters:
logging.info(f"LoRA mode: {len(self.lora_adapters)} adapter(s) will load on first request")
for adapter in self.lora_adapters:
logging.info(f" - {adapter.name}: {adapter.path}")
else:
logging.info("OpenAI engines will initialize on first request")
# Handle both integer and boolean string values for RAW_OPENAI_OUTPUT # Handle both integer and boolean string values for RAW_OPENAI_OUTPUT
raw_output_env = os.getenv("RAW_OPENAI_OUTPUT", "1") raw_output_env = os.getenv("RAW_OPENAI_OUTPUT", "1")
if raw_output_env.lower() in ('true', 'false'): if raw_output_env.lower() in ('true', 'false'):
@@ -201,15 +216,28 @@ class OpenAIvLLMEngine(vLLMEngine):
continue continue
return adapters return adapters
async def _ensure_engines_initialized(self):
"""Initialize engines on first request to avoid event loop mismatch.
In Runpod Serverless, the startup code runs outside the handler's event
loop. Deferring initialization to the first request ensures all async
components (tokenizer pool, serving engines, LoRA state) are created in
the correct event loop context.
"""
if not self._engines_initialized:
logging.info("Initializing OpenAI serving engines...")
await self._initialize_engines()
self._engines_initialized = True
logging.info("OpenAI serving engines initialized successfully")
async def _initialize_engines(self): async def _initialize_engines(self):
self.model_config = await self.llm.get_model_config() self.model_config = self.llm.model_config
self.base_model_paths = [ self.base_model_paths = [
BaseModelPath(name=self.engine_args.model, model_path=self.engine_args.model) BaseModelPath(name=self.engine_args.model, model_path=self.engine_args.model)
] ]
self.serving_models = OpenAIServingModels( self.serving_models = OpenAIServingModels(
engine_client=self.llm, engine_client=self.llm,
model_config=self.model_config,
base_model_paths=self.base_model_paths, base_model_paths=self.base_model_paths,
lora_modules=self.lora_adapters, lora_modules=self.lora_adapters,
) )
@@ -222,28 +250,39 @@ class OpenAIvLLMEngine(vLLMEngine):
self.chat_engine = OpenAIServingChat( self.chat_engine = OpenAIServingChat(
engine_client=self.llm, engine_client=self.llm,
model_config=self.model_config,
models=self.serving_models, models=self.serving_models,
response_role=self.response_role, response_role=self.response_role,
request_logger=None, request_logger=None,
chat_template=chat_template, chat_template=chat_template,
chat_template_content_format="auto", chat_template_content_format="auto",
# enable_reasoning=os.getenv('ENABLE_REASONING', 'false').lower() == 'true', trust_request_chat_template=os.getenv('TRUST_REQUEST_CHAT_TEMPLATE', 'false').lower() == 'true',
reasoning_parser= os.getenv('REASONING_PARSER', "") or None, return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
# return_token_as_token_ids=False, reasoning_parser=os.getenv('REASONING_PARSER', "") or "",
enable_auto_tools=os.getenv('ENABLE_AUTO_TOOL_CHOICE', 'false').lower() == 'true', enable_auto_tools=os.getenv('ENABLE_AUTO_TOOL_CHOICE', 'false').lower() == 'true',
exclude_tools_when_tool_choice_none=os.getenv('EXCLUDE_TOOLS_WHEN_TOOL_CHOICE_NONE', 'false').lower() == 'true',
tool_parser=os.getenv('TOOL_CALL_PARSER', "") or None, tool_parser=os.getenv('TOOL_CALL_PARSER', "") or None,
enable_prompt_tokens_details=False enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
enable_log_outputs=os.getenv('ENABLE_LOG_OUTPUTS', 'false').lower() == 'true',
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
) )
self.completion_engine = OpenAIServingCompletion( self.completion_engine = OpenAIServingCompletion(
engine_client=self.llm, engine_client=self.llm,
model_config=self.model_config,
models=self.serving_models, models=self.serving_models,
request_logger=None, request_logger=None,
# return_token_as_token_ids=False, return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
) )
if hasattr(self.chat_engine, 'warmup'):
await self.chat_engine.warmup()
async def generate(self, openai_request: JobInput): async def generate(self, openai_request: JobInput):
# Ensure engines are ready (no-op if already initialized at startup)
await self._ensure_engines_initialized()
if openai_request.openai_route == "/v1/models": if openai_request.openai_route == "/v1/models":
yield await self._handle_model_request() yield await self._handle_model_request()
elif openai_request.openai_route in ["/v1/chat/completions", "/v1/completions"]: elif openai_request.openai_route in ["/v1/chat/completions", "/v1/completions"]:
+35 -3
View File
@@ -15,7 +15,8 @@ RENAME_ARGS_MAP = {
DEFAULT_ARGS = { DEFAULT_ARGS = {
"disable_log_stats": os.getenv('DISABLE_LOG_STATS', 'False').lower() == 'true', "disable_log_stats": os.getenv('DISABLE_LOG_STATS', 'False').lower() == 'true',
"disable_log_requests": os.getenv('DISABLE_LOG_REQUESTS', 'False').lower() == 'true', # disable_log_requests is deprecated, use enable_log_requests instead
"enable_log_requests": os.getenv('ENABLE_LOG_REQUESTS', 'False').lower() == 'true',
"gpu_memory_utilization": float(os.getenv('GPU_MEMORY_UTILIZATION', 0.95)), "gpu_memory_utilization": float(os.getenv('GPU_MEMORY_UTILIZATION', 0.95)),
"pipeline_parallel_size": int(os.getenv('PIPELINE_PARALLEL_SIZE', 1)), "pipeline_parallel_size": int(os.getenv('PIPELINE_PARALLEL_SIZE', 1)),
"tensor_parallel_size": int(os.getenv('TENSOR_PARALLEL_SIZE', 1)), "tensor_parallel_size": int(os.getenv('TENSOR_PARALLEL_SIZE', 1)),
@@ -38,9 +39,15 @@ DEFAULT_ARGS = {
"block_size": int(os.getenv('BLOCK_SIZE', 16)), "block_size": int(os.getenv('BLOCK_SIZE', 16)),
"enable_prefix_caching": os.getenv('ENABLE_PREFIX_CACHING', 'False').lower() == 'true', "enable_prefix_caching": os.getenv('ENABLE_PREFIX_CACHING', 'False').lower() == 'true',
"disable_sliding_window": os.getenv('DISABLE_SLIDING_WINDOW', 'False').lower() == 'true', "disable_sliding_window": os.getenv('DISABLE_SLIDING_WINDOW', 'False').lower() == 'true',
"use_v2_block_manager": os.getenv('USE_V2_BLOCK_MANAGER', 'False').lower() == 'true', # attention_backend replaces deprecated VLLM_ATTENTION_BACKEND env var
"attention_backend": os.getenv('ATTENTION_BACKEND', None),
# Enabled by default for improved throughput. Set to False to disable if experiencing issues
"async_scheduling": None if os.getenv('ASYNC_SCHEDULING') is None else os.getenv('ASYNC_SCHEDULING', 'True').lower() == 'true',
# Controls how often to yield streaming results
"stream_interval": int(os.getenv('STREAM_INTERVAL', 1)),
"swap_space": int(os.getenv('SWAP_SPACE', 4)), # GiB "swap_space": int(os.getenv('SWAP_SPACE', 4)), # GiB
"cpu_offload_gb": int(os.getenv('CPU_OFFLOAD_GB', 0)), # GiB "cpu_offload_gb": int(os.getenv('CPU_OFFLOAD_GB', 0)), # GiB
# vLLM defaults None to 2048; keep 0 as None to let vLLM auto-calculate
"max_num_batched_tokens": int(os.getenv('MAX_NUM_BATCHED_TOKENS', 0)) or None, "max_num_batched_tokens": int(os.getenv('MAX_NUM_BATCHED_TOKENS', 0)) or None,
"max_num_seqs": int(os.getenv('MAX_NUM_SEQS', 256)), "max_num_seqs": int(os.getenv('MAX_NUM_SEQS', 256)),
"max_logprobs": int(os.getenv('MAX_LOGPROBS', 20)), # Default value for OpenAI Chat Completions API "max_logprobs": int(os.getenv('MAX_LOGPROBS', 20)), # Default value for OpenAI Chat Completions API
@@ -80,6 +87,7 @@ DEFAULT_ARGS = {
"guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'), "guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'),
"speculative_model": os.getenv('SPECULATIVE_MODEL', None), "speculative_model": os.getenv('SPECULATIVE_MODEL', None),
"speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None, "speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None,
"enable_expert_parallel": bool(os.getenv('ENABLE_EXPERT_PARALLEL', 'False').lower() == 'true'),
"num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None, "num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None,
"speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None, "speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None,
"speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None, "speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None,
@@ -91,7 +99,6 @@ DEFAULT_ARGS = {
"qlora_adapter_name_or_path": os.getenv('QLORA_ADAPTER_NAME_OR_PATH', None), "qlora_adapter_name_or_path": os.getenv('QLORA_ADAPTER_NAME_OR_PATH', None),
"disable_logprobs_during_spec_decoding": os.getenv('DISABLE_LOGPROBS_DURING_SPEC_DECODING', None), "disable_logprobs_during_spec_decoding": os.getenv('DISABLE_LOGPROBS_DURING_SPEC_DECODING', None),
"otlp_traces_endpoint": os.getenv('OTLP_TRACES_ENDPOINT', None), "otlp_traces_endpoint": os.getenv('OTLP_TRACES_ENDPOINT', None),
"use_v2_block_manager": os.getenv('USE_V2_BLOCK_MANAGER', 'true'),
} }
limit_mm_env = os.getenv('LIMIT_MM_PER_PROMPT') limit_mm_env = os.getenv('LIMIT_MM_PER_PROMPT')
if limit_mm_env is not None: if limit_mm_env is not None:
@@ -175,4 +182,29 @@ def get_engine_args():
# os.environ["VLLM_ATTENTION_BACKEND"] = "FLASHINFER" # os.environ["VLLM_ATTENTION_BACKEND"] = "FLASHINFER"
# logging.info("Using FLASHINFER for gemma-2 model.") # logging.info("Using FLASHINFER for gemma-2 model.")
# When max_num_batched_tokens is None (env var was 0), set to max_model_len
# to preserve "unlimited" behavior. vLLM defaults None to 2048.
if args.get("max_num_batched_tokens") is None and args.get("max_model_len") is not None:
args["max_num_batched_tokens"] = args["max_model_len"]
logging.info(f"Setting max_num_batched_tokens to max_model_len ({args['max_model_len']}) for unlimited batching.")
# VLLM_ATTENTION_BACKEND is deprecated, migrate to attention_backend
if os.getenv('VLLM_ATTENTION_BACKEND'):
logging.warning(
"VLLM_ATTENTION_BACKEND env var is deprecated. "
"Use ATTENTION_BACKEND instead (maps to --attention-backend CLI arg)."
)
if not args.get('attention_backend'):
args['attention_backend'] = os.getenv('VLLM_ATTENTION_BACKEND')
# DISABLE_LOG_REQUESTS is deprecated, use ENABLE_LOG_REQUESTS instead
if os.getenv('DISABLE_LOG_REQUESTS'):
logging.warning(
"DISABLE_LOG_REQUESTS env var is deprecated. "
"Use ENABLE_LOG_REQUESTS instead (default: False)."
)
# Honor old behavior: if DISABLE_LOG_REQUESTS=true, don't enable logging
if os.getenv('DISABLE_LOG_REQUESTS', 'False').lower() == 'true':
args['enable_log_requests'] = False
return AsyncEngineArgs(**args) return AsyncEngineArgs(**args)
+42 -9
View File
@@ -1,22 +1,55 @@
import os import sys
import multiprocessing
import traceback
import runpod import runpod
from utils import JobInput from runpod import RunPodLogger
from engine import vLLMEngine, OpenAIvLLMEngine
log = RunPodLogger()
vllm_engine = None
openai_engine = None
vllm_engine = vLLMEngine()
OpenAIvLLMEngine = OpenAIvLLMEngine(vllm_engine)
async def handler(job): async def handler(job):
try:
from utils import JobInput
job_input = JobInput(job["input"]) job_input = JobInput(job["input"])
engine = OpenAIvLLMEngine if job_input.openai_route else vllm_engine engine = openai_engine if job_input.openai_route else vllm_engine
results_generator = engine.generate(job_input) results_generator = engine.generate(job_input)
async for batch in results_generator: async for batch in results_generator:
yield batch yield batch
except Exception as e:
error_str = str(e)
full_traceback = traceback.format_exc()
runpod.serverless.start( log.error(f"Error during inference: {error_str}")
log.error(f"Full traceback:\n{full_traceback}")
# CUDA errors = worker is broken, exit to let RunPod spin up a healthy one
if "CUDA" in error_str or "cuda" in error_str:
log.error("Terminating worker due to CUDA/GPU error")
sys.exit(1)
yield {"error": error_str}
# Only run in main process to prevent re-initialization when vLLM spawns worker subprocesses
if __name__ == "__main__" or multiprocessing.current_process().name == "MainProcess":
try:
from engine import vLLMEngine, OpenAIvLLMEngine
vllm_engine = vLLMEngine()
openai_engine = OpenAIvLLMEngine(vllm_engine)
log.info("vLLM engines initialized successfully")
except Exception as e:
log.error(f"Worker startup failed: {e}\n{traceback.format_exc()}")
sys.exit(1)
runpod.serverless.start(
{ {
"handler": handler, "handler": handler,
"concurrency_modifier": lambda x: vllm_engine.max_concurrency, "concurrency_modifier": lambda x: vllm_engine.max_concurrency if vllm_engine else 1,
"return_aggregate_stream": True, "return_aggregate_stream": True,
} }
) )
+1 -2
View File
@@ -3,11 +3,10 @@ import logging
from http import HTTPStatus from http import HTTPStatus
from functools import wraps from functools import wraps
from time import time from time import time
from vllm.entrypoints.openai.protocol import RequestResponseMetadata
try: try:
from vllm.utils import random_uuid from vllm.utils import random_uuid
from vllm.entrypoints.openai.protocol import ErrorResponse from vllm.entrypoints.openai.engine.protocol import ErrorResponse, RequestResponseMetadata
from vllm import SamplingParams from vllm import SamplingParams
except ImportError: except ImportError:
logging.warning("Error importing vllm, skipping related imports. This is ONLY expected when baking model into docker image from a machine without GPUs") logging.warning("Error importing vllm, skipping related imports. This is ONLY expected when baking model into docker image from a machine without GPUs")
-1514
View File
File diff suppressed because it is too large Load Diff