Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
90c16b472d | ||
|
|
6f2381a9a1 | ||
|
|
3851d53f93 | ||
|
|
c896438f21 | ||
|
|
912892f94e | ||
|
|
f8bf82469c | ||
|
|
ec1664902b | ||
|
|
d09122de4a | ||
|
|
e27dc68dea | ||
|
|
1ee18d06a9 | ||
|
|
5c4edd15cc | ||
|
|
b074d3a23b | ||
|
|
205847471c | ||
|
|
6337a6673a | ||
|
|
66e1b1605b | ||
|
|
60c8f257a8 | ||
|
|
fae16e7ee1 | ||
|
|
2becd35345 | ||
|
|
33d88df6c0 | ||
|
|
ecd562e112 | ||
|
|
5cffaab8e8 | ||
|
|
a0fe1dfdad | ||
|
|
0e0d6df859 | ||
|
|
d1718aec00 |
@@ -12,7 +12,7 @@
|
||||
git push origin feature/your-feature-name
|
||||
```
|
||||
|
||||
- Creates pull request → triggers dev build: `runpod/worker-v1-vllm:dev-feature-your-feature-name`
|
||||
- Creates pull request → triggers dev build: `runpod/worker-v1-vllm:dev-refs-pull-214-merge`
|
||||
|
||||
2. **Main Branch**
|
||||
```bash
|
||||
@@ -74,6 +74,6 @@ See [README.md](../README.md) for full list of supported environment variables.
|
||||
|
||||
## 🔧 CI/CD Workflows
|
||||
|
||||
- **Dev builds**: All pull requests → `dev-<branch-name>` images
|
||||
- **Dev builds**: All pull requests → `dev-refs-pull-<PR#>-merge` images
|
||||
- **Release builds**: Git tags → versioned images + GitHub releases
|
||||
- **Manual triggers**: Available in GitHub Actions for emergency releases
|
||||
|
||||
@@ -13,7 +13,6 @@ on:
|
||||
|
||||
permissions:
|
||||
contents: write # Required for creating GitHub releases
|
||||
packages: write # Required for pushing Docker images (if using GitHub packages)
|
||||
|
||||
jobs:
|
||||
release:
|
||||
@@ -75,28 +74,13 @@ jobs:
|
||||
*.args.RELEASE_VERSION=${{ env.RELEASE_VERSION }}
|
||||
*.args.HUGGINGFACE_ACCESS_TOKEN=${{ env.HUGGINGFACE_ACCESS_TOKEN }}
|
||||
|
||||
- name: Create GitHub Release
|
||||
if: env.IS_MANUAL_RELEASE == 'false'
|
||||
uses: actions/create-release@v1
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
with:
|
||||
tag_name: ${{ github.ref_name }}
|
||||
release_name: Release ${{ github.ref_name }}
|
||||
body: |
|
||||
Release ${{ github.ref_name }}
|
||||
|
||||
Docker Image: `${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}`
|
||||
|
||||
## Changes
|
||||
See [commit history](https://github.com/${{ github.repository }}/commits/${{ github.ref_name }}) for detailed changes.
|
||||
draft: false
|
||||
prerelease: false
|
||||
|
||||
- name: Manual Release Summary
|
||||
if: env.IS_MANUAL_RELEASE == 'true'
|
||||
- name: Release Summary
|
||||
run: |
|
||||
echo "🚀 Manual release completed!"
|
||||
echo "🚀 Release completed!"
|
||||
echo "Version: ${{ env.RELEASE_VERSION }}"
|
||||
echo "Docker Image: ${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}"
|
||||
echo "Note: No GitHub release created for manual triggers"
|
||||
if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
|
||||
echo "Trigger: Manual workflow dispatch"
|
||||
else
|
||||
echo "Trigger: GitHub release (tag: ${{ github.ref_name }})"
|
||||
fi
|
||||
|
||||
@@ -4,3 +4,4 @@ runpod.toml
|
||||
.env
|
||||
test/*
|
||||
vllm-base/vllm-*
|
||||
.DS_Store
|
||||
@@ -0,0 +1,261 @@
|
||||

|
||||
|
||||
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
|
||||
|
||||
---
|
||||
|
||||
[](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
|
||||
|
||||
---
|
||||
|
||||
## Endpoint Configuration
|
||||
|
||||
All behaviour is controlled through environment variables:
|
||||
|
||||
| Environment Variable | Description | Default | Options |
|
||||
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
||||
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
||||
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
||||
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
||||
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
||||
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
||||
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
||||
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||
| `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
||||
|
||||
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
|
||||
|
||||
## API Usage
|
||||
|
||||
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
|
||||
|
||||
### RunPod Native API
|
||||
|
||||
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
|
||||
|
||||
#### Chat Completions
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"messages": [
|
||||
{ "role": "system", "content": "You are a helpful assistant." },
|
||||
{ "role": "user", "content": "What is the capital of France?" }
|
||||
],
|
||||
"sampling_params": {
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Chat Completions (Streaming)
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"messages": [
|
||||
{ "role": "user", "content": "Write a short story about a robot." }
|
||||
],
|
||||
"sampling_params": {
|
||||
"max_tokens": 500,
|
||||
"temperature": 0.8
|
||||
},
|
||||
"stream": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Text Generation
|
||||
|
||||
For direct text generation without chat format:
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"prompt": "The capital of France is",
|
||||
"sampling_params": {
|
||||
"max_tokens": 64,
|
||||
"temperature": 0.0
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### List Models
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"openai_route": "/v1/models"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### OpenAI-Compatible API
|
||||
|
||||
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
|
||||
|
||||
#### Chat Completions
|
||||
|
||||
**Path:** `/openai/v1/chat/completions`
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"messages": [
|
||||
{ "role": "system", "content": "You are a helpful assistant." },
|
||||
{ "role": "user", "content": "What is the capital of France?" }
|
||||
],
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}
|
||||
```
|
||||
|
||||
#### Chat Completions (Streaming)
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"messages": [
|
||||
{ "role": "user", "content": "Write a short story about a robot." }
|
||||
],
|
||||
"max_tokens": 500,
|
||||
"temperature": 0.8,
|
||||
"stream": true
|
||||
}
|
||||
```
|
||||
|
||||
#### Text Completions
|
||||
|
||||
**Path:** `/openai/v1/completions`
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"prompt": "The capital of France is",
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}
|
||||
```
|
||||
|
||||
#### List Models
|
||||
|
||||
**Path:** `/openai/v1/models`
|
||||
|
||||
```json
|
||||
{}
|
||||
```
|
||||
|
||||
#### Response Format
|
||||
|
||||
Both APIs return the same response format:
|
||||
|
||||
```json
|
||||
{
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": { "role": "assistant", "content": "Paris." },
|
||||
"finish_reason": "stop"
|
||||
}
|
||||
],
|
||||
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
Below are minimal `python` snippets so you can copy-paste to get started quickly.
|
||||
|
||||
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
|
||||
|
||||
### OpenAI compatible API
|
||||
|
||||
Minimal Python example using the official `openai` SDK:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
import os
|
||||
|
||||
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
|
||||
client = OpenAI(
|
||||
api_key=os.getenv("RUNPOD_API_KEY"),
|
||||
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
|
||||
)
|
||||
```
|
||||
|
||||
`Chat Completions (Non-Streaming)`
|
||||
|
||||
```python
|
||||
response = client.chat.completions.create(
|
||||
model="meta-llama/Llama-2-7b-chat-hf",
|
||||
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
)
|
||||
print(f"Response: {response.choices[0].message.content}")
|
||||
```
|
||||
|
||||
`Chat Completions (Streaming)`
|
||||
|
||||
```python
|
||||
response_stream = client.chat.completions.create(
|
||||
model="meta-llama/Llama-2-7b-chat-hf",
|
||||
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
stream=True
|
||||
)
|
||||
for response in response_stream:
|
||||
print(response.choices[0].delta.content or "", end="", flush=True)
|
||||
```
|
||||
|
||||
### RunPod Native API
|
||||
|
||||
```python
|
||||
import requests
|
||||
|
||||
response = requests.post(
|
||||
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
|
||||
headers={"Authorization": "Bearer <API_KEY>"},
|
||||
json={
|
||||
"input": {
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": "Explain quantum computing in simple terms"}
|
||||
],
|
||||
"sampling_params": {
|
||||
"temperature": 0.7,
|
||||
"max_tokens": 150
|
||||
}
|
||||
}
|
||||
}
|
||||
)
|
||||
|
||||
result = response.json()
|
||||
print(result["output"])
|
||||
```
|
||||
|
||||
## Compatibility
|
||||
|
||||
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
|
||||
|
||||
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
|
||||
|
||||
## Documentation
|
||||
|
||||
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
|
||||
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
|
||||
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
|
||||
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
|
||||
+32
-25
@@ -1,25 +1,15 @@
|
||||
{
|
||||
"title": "vLLM",
|
||||
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the vLLM Inference Engine on RunPod Serverless",
|
||||
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by vLLM",
|
||||
"type": "serverless",
|
||||
"category": "language",
|
||||
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
|
||||
"config": {
|
||||
"runsOn": "GPU",
|
||||
"containerDiskInGb": 200,
|
||||
"gpuIds": "ADA_80_PRO, AMPERE_80",
|
||||
"containerDiskInGb": 150,
|
||||
"gpuIds": "ADA_80_PRO,AMPERE_80",
|
||||
"gpuCount": 1,
|
||||
"allowedCudaVersions": [
|
||||
"12.9",
|
||||
"12.8",
|
||||
"12.7",
|
||||
"12.6",
|
||||
"12.5",
|
||||
"12.4",
|
||||
"12.3",
|
||||
"12.2",
|
||||
"12.1"
|
||||
],
|
||||
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5", "12.4"],
|
||||
"presets": [
|
||||
{
|
||||
"name": "deepseek-ai/deepseek-r1-distill-llama-8b",
|
||||
@@ -38,16 +28,6 @@
|
||||
"required": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "HF_TOKEN",
|
||||
"input": {
|
||||
"name": "Access Token",
|
||||
"type": "string",
|
||||
"description": "Hugging Face access token for gated & private models",
|
||||
"default": "",
|
||||
"required": false
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "TOKENIZER",
|
||||
"input": {
|
||||
@@ -945,7 +925,17 @@
|
||||
"name": "Max Concurrency",
|
||||
"type": "number",
|
||||
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
|
||||
"default": 300,
|
||||
"default": 30,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "ENABLE_EXPERT_PARALLEL",
|
||||
"input": {
|
||||
"name": "Enable Expert Parallel",
|
||||
"type": "boolean",
|
||||
"description": "Enable Expert Parallel for MoE models",
|
||||
"default": false,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
@@ -1023,6 +1013,23 @@
|
||||
"default": "",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "REASONING_PARSER",
|
||||
"input": {
|
||||
"name": "Reasoning Parser",
|
||||
"type": "string",
|
||||
"description": "Parser for reasoning-capable models (enables reasoning mode)",
|
||||
"options": [
|
||||
{ "label": "None", "value": "" },
|
||||
{ "label": "DeepSeek R1", "value": "deepseek_r1" },
|
||||
{ "label": "Qwen3", "value": "qwen3" },
|
||||
{ "label": "Granite", "value": "granite" },
|
||||
{ "label": "Hunyuan A13B", "value": "hunyuan_a13b" }
|
||||
],
|
||||
"default": "",
|
||||
"advanced": true
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
+3
-12
@@ -12,7 +12,6 @@
|
||||
"input": {
|
||||
"openai_route": "/v1/chat/completions",
|
||||
"openai_input": {
|
||||
"model": "HuggingFaceTB/SmolLM2-135M-Instruct",
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
@@ -23,8 +22,8 @@
|
||||
"content": "Explain what a neural network is in one sentence."
|
||||
}
|
||||
],
|
||||
"max_tokens": 50,
|
||||
"temperature": 0.7
|
||||
"max_tokens": 200,
|
||||
"temperature": 0.1
|
||||
}
|
||||
},
|
||||
"timeout": 30000
|
||||
@@ -39,14 +38,6 @@
|
||||
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
|
||||
}
|
||||
],
|
||||
"allowedCudaVersions": [
|
||||
"12.7",
|
||||
"12.6",
|
||||
"12.5",
|
||||
"12.4",
|
||||
"12.3",
|
||||
"12.2",
|
||||
"12.1"
|
||||
]
|
||||
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
|
||||
}
|
||||
}
|
||||
|
||||
+4
-5
@@ -1,9 +1,9 @@
|
||||
FROM nvidia/cuda:12.1.0-base-ubuntu22.04
|
||||
FROM nvidia/cuda:12.4.1-base-ubuntu22.04
|
||||
|
||||
RUN apt-get update -y \
|
||||
&& apt-get install -y python3-pip
|
||||
|
||||
RUN ldconfig /usr/local/cuda-12.1/compat/
|
||||
RUN ldconfig /usr/local/cuda-12.4/compat/
|
||||
|
||||
# Install Python dependencies
|
||||
COPY builder/requirements.txt /requirements.txt
|
||||
@@ -11,9 +11,8 @@ RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
python3 -m pip install --upgrade pip && \
|
||||
python3 -m pip install --upgrade -r /requirements.txt
|
||||
|
||||
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
|
||||
RUN python3 -m pip install vllm==0.10.0 && \
|
||||
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
|
||||
# Install vLLM
|
||||
RUN python3 -m pip install vllm==0.11.0
|
||||
|
||||
# Setup for Option 2: Building the Image with the Model included
|
||||
ARG MODEL_NAME=""
|
||||
|
||||
@@ -9,16 +9,10 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
||||
## Table of Contents
|
||||
|
||||
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
|
||||
- [Option 1: Deploy Any Model Using Pre-Built Docker Image **[RECOMMENDED]**](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
||||
- [Environment Variables](#environment-variables)
|
||||
- [LLM Settings](#llm-settings)
|
||||
- [Tokenizer Settings](#tokenizer-settings)
|
||||
- [System and Parallelism Settings](#system-and-parallelism-settings)
|
||||
- [Streaming Batch Size Settings](#streaming-batch-size-settings)
|
||||
- [OpenAI Settings](#openai-settings)
|
||||
- [Serverless Settings](#serverless-settings)
|
||||
- [Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended]](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
||||
- [Configuration](#configuration)
|
||||
- [Option 2: Build Docker Image with Model Inside](#option-2-build-docker-image-with-model-inside)
|
||||
- [Prerequisites](#prerequisites-1)
|
||||
- [Prerequisites](#prerequisites)
|
||||
- [Arguments](#arguments)
|
||||
- [Example: Building an image with OpenChat-3.5](#example-building-an-image-with-openchat-35)
|
||||
- [(Optional) Including Huggingface Token](#optional-including-huggingface-token)
|
||||
@@ -26,12 +20,14 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
||||
- [Usage: OpenAI Compatibility](#usage-openai-compatibility)
|
||||
- [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker)
|
||||
- [OpenAI Request Input Parameters](#openai-request-input-parameters)
|
||||
- [Chat Completions](#chat-completions)
|
||||
- [Chat Completions [RECOMMENDED]](#chat-completions-recommended)
|
||||
- [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai)
|
||||
- [Usage: standard](#non-openai-usage)
|
||||
- [Input Request Parameters](#input-request-parameters)
|
||||
- [Chat Completions](#chat-completions)
|
||||
- [Getting a list of names for available models](#getting-a-list-of-names-for-available-models)
|
||||
- [Usage: Standard (Non-OpenAI)](#usage-standard-non-openai)
|
||||
- [Request Input Parameters](#request-input-parameters)
|
||||
- [Sampling Parameters](#sampling-parameters)
|
||||
- [Text Input Formats](#text-input-formats)
|
||||
- [Sampling Parameters](#sampling-parameters)
|
||||
|
||||
# Setting up the Serverless Worker
|
||||
|
||||
@@ -39,129 +35,31 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
||||
|
||||
**🚀 Deploy Guide**: Follow our [step-by-step deployment guide](https://docs.runpod.io/serverless/vllm/get-started) to deploy using the RunPod Console.
|
||||
|
||||
**📦 Docker Image**: `runpod/worker-v1-vllm:<version>stable-cuda12.1.0`
|
||||
**📦 Docker Image**: `runpod/worker-v1-vllm:<version>`
|
||||
|
||||
- **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases)
|
||||
- **CUDA Compatibility**: Requires CUDA >= 12.1
|
||||
|
||||
### Environment Variables
|
||||
### Configuration
|
||||
|
||||
Use these to configure worker-vllm so it works for your use case / model.
|
||||
Configure worker-vllm using environment variables:
|
||||
|
||||
#### LLM Settings
|
||||
| Environment Variable | Description | Default | Options |
|
||||
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
||||
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
||||
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
||||
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
||||
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
||||
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
||||
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
||||
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ------------------------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
|
||||
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
|
||||
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
|
||||
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
|
||||
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
|
||||
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
|
||||
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
|
||||
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
|
||||
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
|
||||
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
|
||||
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
|
||||
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
|
||||
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
|
||||
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
|
||||
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
|
||||
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
|
||||
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
|
||||
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
|
||||
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
|
||||
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
|
||||
| `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. |
|
||||
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
|
||||
| `SEED` | 0 | `int` | Random seed for operations. |
|
||||
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
|
||||
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
|
||||
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
|
||||
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
|
||||
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
|
||||
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
|
||||
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
|
||||
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
|
||||
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
|
||||
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
|
||||
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
|
||||
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
|
||||
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
|
||||
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
|
||||
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
|
||||
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
|
||||
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}` |
|
||||
| `SCHEDULER_DELAY_FACTOR` | 0.0 | `float` | Apply a delay before scheduling next prompt. |
|
||||
| `ENABLE_CHUNKED_PREFILL` | False | `bool` | Enable chunked prefill requests. |
|
||||
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
|
||||
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
|
||||
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
|
||||
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
|
||||
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
|
||||
| `SPEC_DECODING_ACCEPTANCE_METHOD` | 'rejection_sampler' | ['rejection_sampler', 'typical_acceptance_sampler'] | Specify the acceptance method for draft token verification in speculative decoding. |
|
||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD` | None | `float` | Set the lower bound threshold for the posterior probability of a token to be accepted. |
|
||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA` | None | `float` | A scaling factor for the entropy-based threshold for token acceptance. |
|
||||
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
|
||||
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
|
||||
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
|
||||
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
|
||||
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
|
||||
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
|
||||
|
||||
#### Tokenizer Settings
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ---------------------- | --------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
|
||||
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
|
||||
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
|
||||
|
||||
#### System and Parallelism Settings
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ------------------------------ | --------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
|
||||
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
|
||||
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
|
||||
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
||||
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
|
||||
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
|
||||
|
||||
#### Streaming Batch Size Settings
|
||||
|
||||
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ---------------------------------- | --------- | -------------- | --------------------------------------------------------------------------------------------------------- |
|
||||
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
|
||||
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
|
||||
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
|
||||
|
||||
#### OpenAI Settings
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
|
||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||
|
||||
#### Serverless Settings
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ---------------------- | --------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
||||
|
||||
## Option 2: Build Docker Image with Model Inside
|
||||
|
||||
|
||||
@@ -1,14 +1,14 @@
|
||||
ray
|
||||
pandas
|
||||
pyarrow
|
||||
runpod~=1.7.7
|
||||
runpod>=1.8,<2.0
|
||||
huggingface-hub
|
||||
packaging
|
||||
typing-extensions>=4.8.0
|
||||
pydantic
|
||||
pydantic-settings
|
||||
hf-transfer
|
||||
transformers>=4.55.0
|
||||
transformers>=4.57.0
|
||||
bitsandbytes>=0.45.0
|
||||
kernels
|
||||
torch==2.6.0
|
||||
|
||||
@@ -0,0 +1,154 @@
|
||||
# Configuration Reference
|
||||
|
||||
Complete guide to all environment variables and configuration options for worker-vllm.
|
||||
|
||||
## LLM Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------- |
|
||||
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
|
||||
| `MODEL_REVISION` | 'main' | `str` | Model revision to load (default: main). |
|
||||
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
|
||||
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
|
||||
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
|
||||
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
|
||||
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
|
||||
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
|
||||
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
|
||||
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
|
||||
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
|
||||
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
|
||||
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
|
||||
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
|
||||
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
|
||||
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
|
||||
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
|
||||
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
|
||||
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
|
||||
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
|
||||
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
|
||||
| `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. |
|
||||
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
|
||||
| `SEED` | 0 | `int` | Random seed for operations. |
|
||||
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
|
||||
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
|
||||
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
|
||||
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
|
||||
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
|
||||
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
|
||||
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
|
||||
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
|
||||
|
||||
## LoRA (Low-Rank Adaptation) Settings
|
||||
|
||||
| Variable | Default | Type | Description |
|
||||
| --------------------------- | ------- | ------------------------------------------ | --------------------------------------------------------------------------------------------------------- |
|
||||
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
|
||||
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
|
||||
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
|
||||
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
|
||||
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
|
||||
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
|
||||
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
|
||||
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
|
||||
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}]` |
|
||||
|
||||
## Speculative Decoding Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ------------------------------------------------ | ------------------- | --------------------------------------------------- | ----------------------------------------------------------------------------------------- |
|
||||
| `SCHEDULER_DELAY_FACTOR` | 0.0 | `float` | Apply a delay before scheduling next prompt. |
|
||||
| `ENABLE_CHUNKED_PREFILL` | False | `bool` | Enable chunked prefill requests. |
|
||||
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
|
||||
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
|
||||
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
|
||||
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
|
||||
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
|
||||
| `SPEC_DECODING_ACCEPTANCE_METHOD` | 'rejection_sampler' | ['rejection_sampler', 'typical_acceptance_sampler'] | Specify the acceptance method for draft token verification in speculative decoding. |
|
||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD` | None | `float` | Set the lower bound threshold for the posterior probability of a token to be accepted. |
|
||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA` | None | `float` | A scaling factor for the entropy-based threshold for token acceptance. |
|
||||
|
||||
## System Performance Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ------------------------------ | ------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
|
||||
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
|
||||
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
|
||||
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
||||
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
|
||||
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
|
||||
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models |
|
||||
|
||||
## Tokenizer Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------- | ------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
|
||||
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
|
||||
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
|
||||
|
||||
## Streaming & Batch Settings
|
||||
|
||||
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker.
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------------------- | ------- | ------------ | --------------------------------------------------------------------------------------------------------- |
|
||||
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
|
||||
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
|
||||
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
|
||||
|
||||
## OpenAI Compatibility Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
|
||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
|
||||
|
||||
## Serverless & Concurrency Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||
|
||||
## Advanced Settings
|
||||
|
||||
| Variable | Default | Type | Description |
|
||||
| --------------------------- | ------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
|
||||
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
|
||||
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
|
||||
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
|
||||
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
|
||||
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
|
||||
|
||||
## Docker Build Arguments
|
||||
|
||||
These variables are used when building custom Docker images with models baked in:
|
||||
|
||||
| Variable | Default | Type | Description |
|
||||
| --------------------- | ---------------- | ----- | ------------------------------------------------- |
|
||||
| `BASE_PATH` | `/runpod-volume` | `str` | Storage directory for huggingface cache and model |
|
||||
| `WORKER_CUDA_VERSION` | `12.1.0` | `str` | CUDA version for the worker image |
|
||||
|
||||
## Deprecated Variables
|
||||
|
||||
⚠️ **The following variables are deprecated and will be removed in future versions:**
|
||||
|
||||
| Old Variable | New Variable | Note |
|
||||
| ---------------------------- | ------------------------ | --------------------- |
|
||||
| `MAX_CONTEXT_LEN_TO_CAPTURE` | `MAX_SEQ_LEN_TO_CAPTURE` | Use new variable name |
|
||||
| `kv_cache_dtype=fp8_e5m2` | `kv_cache_dtype=fp8` | Simplified fp8 format |
|
||||
+3
-2
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
|
||||
|
||||
- `src/engine_args.py`: Centralized configuration management
|
||||
- `src/constants.py`: Default values for core settings
|
||||
- `worker-config.json`: UI form generation for RunPod console
|
||||
- `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
|
||||
- `worker-config.json`: UI form generation for RunPod console (if exists)
|
||||
|
||||
## Core Development Concepts
|
||||
|
||||
@@ -222,7 +223,7 @@ src/
|
||||
|
||||
### 2. **Concurrency Patterns**
|
||||
|
||||
- **Max Concurrency**: 300 concurrent requests by default
|
||||
- **Max Concurrency**: 30 concurrent requests by default
|
||||
- **vLLM Queuing**: Internal request batching and scheduling
|
||||
- **RunPod Integration**: Concurrency modifier for auto-scaling
|
||||
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 86 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 1.1 MiB |
+1
-1
@@ -1,4 +1,4 @@
|
||||
DEFAULT_BATCH_SIZE = 50
|
||||
DEFAULT_MAX_CONCURRENCY = 300
|
||||
DEFAULT_MAX_CONCURRENCY = 30
|
||||
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
|
||||
DEFAULT_MIN_BATCH_SIZE = 1
|
||||
@@ -80,6 +80,7 @@ DEFAULT_ARGS = {
|
||||
"guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'),
|
||||
"speculative_model": os.getenv('SPECULATIVE_MODEL', None),
|
||||
"speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None,
|
||||
"enable_expert_parallel": bool(os.getenv('ENABLE_EXPERT_PARALLEL', 'False').lower() == 'true'),
|
||||
"num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None,
|
||||
"speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None,
|
||||
"speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None,
|
||||
|
||||
-1514
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user