Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
08c6ff7490 | ||
|
|
9d1686960d | ||
|
|
45d1eeee47 | ||
|
|
17efb0e7d0 | ||
|
|
2b5f07df63 | ||
|
|
13fa71878e | ||
|
|
8a9365bed4 | ||
|
|
cd485a1af1 | ||
|
|
b9043639e9 | ||
|
|
407dbd7773 | ||
|
|
f103c142c1 | ||
|
|
efb093e198 | ||
|
|
42443f735e | ||
|
|
b7c6d4f9a2 | ||
|
|
d69cc021e8 | ||
|
|
61faa8f137 | ||
|
|
1606cff557 | ||
|
|
e705c9494b | ||
|
|
b749aa5718 | ||
|
|
4705ba8a7c | ||
|
|
767c66c301 | ||
|
|
fefdbe21a9 | ||
|
|
ee961ad28d | ||
|
|
2e8c251447 | ||
|
|
c3cf43b228 | ||
|
|
7ec10b98cd | ||
|
|
c45ac42acd | ||
|
|
340bc0b3c6 | ||
|
|
e1e9ef74ad | ||
|
|
461f89cea6 | ||
|
|
8eb55b90c1 | ||
|
|
6d6cbe7095 | ||
|
|
90c16b472d | ||
|
|
6f2381a9a1 | ||
|
|
3851d53f93 | ||
|
|
c896438f21 | ||
|
|
912892f94e | ||
|
|
f8bf82469c | ||
|
|
ec1664902b | ||
|
|
d09122de4a | ||
|
|
e27dc68dea | ||
|
|
1ee18d06a9 | ||
|
|
5c4edd15cc | ||
|
|
b074d3a23b | ||
|
|
205847471c | ||
|
|
6337a6673a | ||
|
|
66e1b1605b | ||
|
|
60c8f257a8 | ||
|
|
fae16e7ee1 | ||
|
|
2becd35345 | ||
|
|
33d88df6c0 | ||
|
|
ecd562e112 | ||
|
|
5cffaab8e8 | ||
|
|
a0fe1dfdad | ||
|
|
0e0d6df859 | ||
|
|
d1718aec00 |
@@ -12,7 +12,7 @@
|
||||
git push origin feature/your-feature-name
|
||||
```
|
||||
|
||||
- Creates pull request → triggers dev build: `runpod/worker-v1-vllm:dev-feature-your-feature-name`
|
||||
- Creates pull request → triggers dev build: `runpod/worker-v1-vllm:dev-refs-pull-214-merge`
|
||||
|
||||
2. **Main Branch**
|
||||
```bash
|
||||
@@ -74,6 +74,6 @@ See [README.md](../README.md) for full list of supported environment variables.
|
||||
|
||||
## 🔧 CI/CD Workflows
|
||||
|
||||
- **Dev builds**: All pull requests → `dev-<branch-name>` images
|
||||
- **Dev builds**: All pull requests → `dev-refs-pull-<PR#>-merge` images
|
||||
- **Release builds**: Git tags → versioned images + GitHub releases
|
||||
- **Manual triggers**: Available in GitHub Actions for emergency releases
|
||||
|
||||
@@ -13,7 +13,6 @@ on:
|
||||
|
||||
permissions:
|
||||
contents: write # Required for creating GitHub releases
|
||||
packages: write # Required for pushing Docker images (if using GitHub packages)
|
||||
|
||||
jobs:
|
||||
release:
|
||||
@@ -75,28 +74,13 @@ jobs:
|
||||
*.args.RELEASE_VERSION=${{ env.RELEASE_VERSION }}
|
||||
*.args.HUGGINGFACE_ACCESS_TOKEN=${{ env.HUGGINGFACE_ACCESS_TOKEN }}
|
||||
|
||||
- name: Create GitHub Release
|
||||
if: env.IS_MANUAL_RELEASE == 'false'
|
||||
uses: actions/create-release@v1
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
with:
|
||||
tag_name: ${{ github.ref_name }}
|
||||
release_name: Release ${{ github.ref_name }}
|
||||
body: |
|
||||
Release ${{ github.ref_name }}
|
||||
|
||||
Docker Image: `${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}`
|
||||
|
||||
## Changes
|
||||
See [commit history](https://github.com/${{ github.repository }}/commits/${{ github.ref_name }}) for detailed changes.
|
||||
draft: false
|
||||
prerelease: false
|
||||
|
||||
- name: Manual Release Summary
|
||||
if: env.IS_MANUAL_RELEASE == 'true'
|
||||
- name: Release Summary
|
||||
run: |
|
||||
echo "🚀 Manual release completed!"
|
||||
echo "🚀 Release completed!"
|
||||
echo "Version: ${{ env.RELEASE_VERSION }}"
|
||||
echo "Docker Image: ${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}"
|
||||
echo "Note: No GitHub release created for manual triggers"
|
||||
if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
|
||||
echo "Trigger: Manual workflow dispatch"
|
||||
else
|
||||
echo "Trigger: GitHub release (tag: ${{ github.ref_name }})"
|
||||
fi
|
||||
|
||||
@@ -4,3 +4,4 @@ runpod.toml
|
||||
.env
|
||||
test/*
|
||||
vllm-base/vllm-*
|
||||
.DS_Store
|
||||
@@ -0,0 +1,263 @@
|
||||

|
||||
|
||||
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
|
||||
|
||||
---
|
||||
|
||||
[](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
|
||||
|
||||
---
|
||||
|
||||
## Endpoint Configuration
|
||||
|
||||
All behaviour is controlled through environment variables:
|
||||
|
||||
| Environment Variable | Description | Default | Options |
|
||||
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
||||
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
||||
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
||||
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
||||
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
||||
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
||||
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
||||
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||
| `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
||||
|
||||
**Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options.
|
||||
|
||||
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
|
||||
|
||||
## API Usage
|
||||
|
||||
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
|
||||
|
||||
### RunPod Native API
|
||||
|
||||
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
|
||||
|
||||
#### Chat Completions
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"messages": [
|
||||
{ "role": "system", "content": "You are a helpful assistant." },
|
||||
{ "role": "user", "content": "What is the capital of France?" }
|
||||
],
|
||||
"sampling_params": {
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Chat Completions (Streaming)
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"messages": [
|
||||
{ "role": "user", "content": "Write a short story about a robot." }
|
||||
],
|
||||
"sampling_params": {
|
||||
"max_tokens": 500,
|
||||
"temperature": 0.8
|
||||
},
|
||||
"stream": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Text Generation
|
||||
|
||||
For direct text generation without chat format:
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"prompt": "The capital of France is",
|
||||
"sampling_params": {
|
||||
"max_tokens": 64,
|
||||
"temperature": 0.0
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### List Models
|
||||
|
||||
```json
|
||||
{
|
||||
"input": {
|
||||
"openai_route": "/v1/models"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### OpenAI-Compatible API
|
||||
|
||||
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
|
||||
|
||||
#### Chat Completions
|
||||
|
||||
**Path:** `/openai/v1/chat/completions`
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"messages": [
|
||||
{ "role": "system", "content": "You are a helpful assistant." },
|
||||
{ "role": "user", "content": "What is the capital of France?" }
|
||||
],
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}
|
||||
```
|
||||
|
||||
#### Chat Completions (Streaming)
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"messages": [
|
||||
{ "role": "user", "content": "Write a short story about a robot." }
|
||||
],
|
||||
"max_tokens": 500,
|
||||
"temperature": 0.8,
|
||||
"stream": true
|
||||
}
|
||||
```
|
||||
|
||||
#### Text Completions
|
||||
|
||||
**Path:** `/openai/v1/completions`
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"prompt": "The capital of France is",
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}
|
||||
```
|
||||
|
||||
#### List Models
|
||||
|
||||
**Path:** `/openai/v1/models`
|
||||
|
||||
```json
|
||||
{}
|
||||
```
|
||||
|
||||
#### Response Format
|
||||
|
||||
Both APIs return the same response format:
|
||||
|
||||
```json
|
||||
{
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": { "role": "assistant", "content": "Paris." },
|
||||
"finish_reason": "stop"
|
||||
}
|
||||
],
|
||||
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
Below are minimal `python` snippets so you can copy-paste to get started quickly.
|
||||
|
||||
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
|
||||
|
||||
### OpenAI compatible API
|
||||
|
||||
Minimal Python example using the official `openai` SDK:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
import os
|
||||
|
||||
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
|
||||
client = OpenAI(
|
||||
api_key=os.getenv("RUNPOD_API_KEY"),
|
||||
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
|
||||
)
|
||||
```
|
||||
|
||||
`Chat Completions (Non-Streaming)`
|
||||
|
||||
```python
|
||||
response = client.chat.completions.create(
|
||||
model="meta-llama/Llama-2-7b-chat-hf",
|
||||
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
)
|
||||
print(f"Response: {response.choices[0].message.content}")
|
||||
```
|
||||
|
||||
`Chat Completions (Streaming)`
|
||||
|
||||
```python
|
||||
response_stream = client.chat.completions.create(
|
||||
model="meta-llama/Llama-2-7b-chat-hf",
|
||||
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
stream=True
|
||||
)
|
||||
for response in response_stream:
|
||||
print(response.choices[0].delta.content or "", end="", flush=True)
|
||||
```
|
||||
|
||||
### RunPod Native API
|
||||
|
||||
```python
|
||||
import requests
|
||||
|
||||
response = requests.post(
|
||||
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
|
||||
headers={"Authorization": "Bearer <API_KEY>"},
|
||||
json={
|
||||
"input": {
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": "Explain quantum computing in simple terms"}
|
||||
],
|
||||
"sampling_params": {
|
||||
"temperature": 0.7,
|
||||
"max_tokens": 150
|
||||
}
|
||||
}
|
||||
}
|
||||
)
|
||||
|
||||
result = response.json()
|
||||
print(result["output"])
|
||||
```
|
||||
|
||||
## Compatibility
|
||||
|
||||
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
|
||||
|
||||
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
|
||||
|
||||
## Documentation
|
||||
|
||||
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
|
||||
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
|
||||
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
|
||||
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
|
||||
+67
-295
@@ -1,25 +1,15 @@
|
||||
{
|
||||
"title": "vLLM",
|
||||
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the vLLM Inference Engine on RunPod Serverless",
|
||||
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by vLLM",
|
||||
"type": "serverless",
|
||||
"category": "language",
|
||||
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
|
||||
"config": {
|
||||
"runsOn": "GPU",
|
||||
"containerDiskInGb": 200,
|
||||
"gpuIds": "ADA_80_PRO, AMPERE_80",
|
||||
"containerDiskInGb": 150,
|
||||
"gpuIds": "ADA_80_PRO,AMPERE_80",
|
||||
"gpuCount": 1,
|
||||
"allowedCudaVersions": [
|
||||
"12.9",
|
||||
"12.8",
|
||||
"12.7",
|
||||
"12.6",
|
||||
"12.5",
|
||||
"12.4",
|
||||
"12.3",
|
||||
"12.2",
|
||||
"12.1"
|
||||
],
|
||||
"allowedCudaVersions": ["12.9", "12.8"],
|
||||
"presets": [
|
||||
{
|
||||
"name": "deepseek-ai/deepseek-r1-distill-llama-8b",
|
||||
@@ -38,16 +28,6 @@
|
||||
"required": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "HF_TOKEN",
|
||||
"input": {
|
||||
"name": "Access Token",
|
||||
"type": "string",
|
||||
"description": "Hugging Face access token for gated & private models",
|
||||
"default": "",
|
||||
"required": false
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "TOKENIZER",
|
||||
"input": {
|
||||
@@ -201,41 +181,13 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "QUANTIZATION_PARAM_PATH",
|
||||
"input": {
|
||||
"name": "Quantization Param Path",
|
||||
"type": "string",
|
||||
"description": "Path to the JSON file containing the KV cache scaling factors.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "MAX_MODEL_LEN",
|
||||
"input": {
|
||||
"name": "Max Model Length",
|
||||
"type": "number",
|
||||
"description": "Model context length.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "GUIDED_DECODING_BACKEND",
|
||||
"input": {
|
||||
"name": "Guided Decoding Backend",
|
||||
"type": "string",
|
||||
"description": "Which engine will be used for guided decoding by default.",
|
||||
"options": [
|
||||
{
|
||||
"label": "outlines",
|
||||
"value": "outlines"
|
||||
},
|
||||
{
|
||||
"label": "lm-format-enforcer",
|
||||
"value": "lm-format-enforcer"
|
||||
}
|
||||
],
|
||||
"default": "outlines",
|
||||
"default": null,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
@@ -255,17 +207,8 @@
|
||||
"value": "mp"
|
||||
}
|
||||
],
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "WORKER_USE_RAY",
|
||||
"input": {
|
||||
"name": "Worker Use Ray",
|
||||
"type": "boolean",
|
||||
"description": "Deprecated, use --distributed-executor-backend=ray.",
|
||||
"default": false,
|
||||
"advanced": true
|
||||
"advanced": true,
|
||||
"default": "mp"
|
||||
}
|
||||
},
|
||||
{
|
||||
@@ -327,26 +270,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "USE_V2_BLOCK_MANAGER",
|
||||
"input": {
|
||||
"name": "Use V2 Block Manager",
|
||||
"type": "boolean",
|
||||
"description": "Use BlockSpaceMangerV2.",
|
||||
"default": false,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "NUM_LOOKAHEAD_SLOTS",
|
||||
"input": {
|
||||
"name": "Num Lookahead Slots",
|
||||
"type": "number",
|
||||
"description": "Experimental scheduling config necessary for speculative decoding.",
|
||||
"default": 0,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SEED",
|
||||
"input": {
|
||||
@@ -357,21 +280,13 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "NUM_GPU_BLOCKS_OVERRIDE",
|
||||
"input": {
|
||||
"name": "Num GPU Blocks Override",
|
||||
"type": "number",
|
||||
"description": "If specified, ignore GPU profiling result and use this number of GPU blocks.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "MAX_NUM_BATCHED_TOKENS",
|
||||
"input": {
|
||||
"name": "Max Num Batched Tokens",
|
||||
"type": "number",
|
||||
"description": "Maximum number of batched tokens per iteration.",
|
||||
"default": null,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
@@ -432,53 +347,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "ROPE_SCALING",
|
||||
"input": {
|
||||
"name": "RoPE Scaling",
|
||||
"type": "string",
|
||||
"description": "RoPE scaling configuration in JSON format.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "ROPE_THETA",
|
||||
"input": {
|
||||
"name": "RoPE Theta",
|
||||
"type": "number",
|
||||
"description": "RoPE theta. Use with rope_scaling.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "TOKENIZER_POOL_SIZE",
|
||||
"input": {
|
||||
"name": "Tokenizer Pool Size",
|
||||
"type": "number",
|
||||
"description": "Size of tokenizer pool to use for asynchronous tokenization.",
|
||||
"default": 0,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "TOKENIZER_POOL_TYPE",
|
||||
"input": {
|
||||
"name": "Tokenizer Pool Type",
|
||||
"type": "string",
|
||||
"description": "Type of tokenizer pool to use for asynchronous tokenization.",
|
||||
"default": "ray",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "TOKENIZER_POOL_EXTRA_CONFIG",
|
||||
"input": {
|
||||
"name": "Tokenizer Pool Extra Config",
|
||||
"type": "string",
|
||||
"description": "Extra config for tokenizer pool.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "ENABLE_LORA",
|
||||
"input": {
|
||||
@@ -509,16 +377,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "LORA_EXTRA_VOCAB_SIZE",
|
||||
"input": {
|
||||
"name": "LoRA Extra Vocab Size",
|
||||
"type": "number",
|
||||
"description": "Maximum size of extra vocabulary for LoRA adapters.",
|
||||
"default": 256,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "LORA_DTYPE",
|
||||
"input": {
|
||||
@@ -547,15 +405,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "LONG_LORA_SCALING_FACTORS",
|
||||
"input": {
|
||||
"name": "Long LoRA Scaling Factors",
|
||||
"type": "string",
|
||||
"description": "Specify multiple scaling factors for LoRA adapters.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "MAX_CPU_LORAS",
|
||||
"input": {
|
||||
@@ -635,6 +484,34 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SPECULATIVE_CONFIG",
|
||||
"input": {
|
||||
"name": "Speculative Config (JSON)",
|
||||
"type": "string",
|
||||
"description": "Full speculative decoding configuration as a JSON string. Overrides individual speculative env vars.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SPECULATIVE_METHOD",
|
||||
"input": {
|
||||
"name": "Speculative Method",
|
||||
"type": "string",
|
||||
"description": "Speculative decoding method to use.",
|
||||
"options": [
|
||||
{ "label": "None", "value": "" },
|
||||
{ "label": "Draft Model", "value": "draft_model" },
|
||||
{ "label": "N-gram", "value": "ngram" },
|
||||
{ "label": "EAGLE", "value": "eagle" },
|
||||
{ "label": "EAGLE3", "value": "eagle3" },
|
||||
{ "label": "Medusa", "value": "medusa" },
|
||||
{ "label": "MLP Speculator", "value": "mlp_speculator" }
|
||||
],
|
||||
"default": "",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SPECULATIVE_MODEL",
|
||||
"input": {
|
||||
@@ -653,33 +530,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE",
|
||||
"input": {
|
||||
"name": "Speculative Draft Tensor Parallel Size",
|
||||
"type": "number",
|
||||
"description": "Number of tensor parallel replicas for the draft model.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SPECULATIVE_MAX_MODEL_LEN",
|
||||
"input": {
|
||||
"name": "Speculative Max Model Length",
|
||||
"type": "number",
|
||||
"description": "The maximum sequence length supported by the draft model.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SPECULATIVE_DISABLE_BY_BATCH_SIZE",
|
||||
"input": {
|
||||
"name": "Speculative Disable by Batch Size",
|
||||
"type": "number",
|
||||
"description": "Disable speculative decoding if the number of enqueue requests is larger than this value.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "NGRAM_PROMPT_LOOKUP_MAX",
|
||||
"input": {
|
||||
@@ -689,53 +539,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "NGRAM_PROMPT_LOOKUP_MIN",
|
||||
"input": {
|
||||
"name": "Ngram Prompt Lookup Min",
|
||||
"type": "number",
|
||||
"description": "Min size of window for ngram prompt lookup in speculative decoding.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "SPEC_DECODING_ACCEPTANCE_METHOD",
|
||||
"input": {
|
||||
"name": "Speculative Decoding Acceptance Method",
|
||||
"type": "string",
|
||||
"description": "Specify the acceptance method for draft token verification in speculative decoding.",
|
||||
"options": [
|
||||
{
|
||||
"label": "rejection_sampler",
|
||||
"value": "rejection_sampler"
|
||||
},
|
||||
{
|
||||
"label": "typical_acceptance_sampler",
|
||||
"value": "typical_acceptance_sampler"
|
||||
}
|
||||
],
|
||||
"default": "rejection_sampler",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD",
|
||||
"input": {
|
||||
"name": "Typical Acceptance Sampler Posterior Threshold",
|
||||
"type": "number",
|
||||
"description": "Set the lower bound threshold for the posterior probability of a token to be accepted.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA",
|
||||
"input": {
|
||||
"name": "Typical Acceptance Sampler Posterior Alpha",
|
||||
"type": "number",
|
||||
"description": "A scaling factor for the entropy-based threshold for token acceptance.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "MODEL_LOADER_EXTRA_CONFIG",
|
||||
"input": {
|
||||
@@ -746,49 +549,11 @@
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "PREEMPTION_MODE",
|
||||
"key": "ENABLE_LOG_REQUESTS",
|
||||
"input": {
|
||||
"name": "Preemption Mode",
|
||||
"type": "string",
|
||||
"description": "If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "PREEMPTION_CHECK_PERIOD",
|
||||
"input": {
|
||||
"name": "Preemption Check Period",
|
||||
"type": "number",
|
||||
"description": "How frequently the engine checks if a preemption happens.",
|
||||
"default": 1,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "PREEMPTION_CPU_CAPACITY",
|
||||
"input": {
|
||||
"name": "Preemption CPU Capacity",
|
||||
"type": "number",
|
||||
"description": "The percentage of CPU memory used for the saved activations.",
|
||||
"default": 2,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "MAX_LOG_LEN",
|
||||
"input": {
|
||||
"name": "Max Log Length",
|
||||
"type": "number",
|
||||
"description": "Max number of characters or ID numbers being printed in log.",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "DISABLE_LOGGING_REQUEST",
|
||||
"input": {
|
||||
"name": "Disable Logging Request",
|
||||
"name": "Enable Log Requests",
|
||||
"type": "boolean",
|
||||
"description": "Disable logging requests.",
|
||||
"description": "Enable vLLM request logging.",
|
||||
"default": false,
|
||||
"advanced": true
|
||||
}
|
||||
@@ -860,16 +625,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "MAX_SEQ_LEN_TO_CAPTURE",
|
||||
"input": {
|
||||
"name": "CUDA Graph Max Content Length",
|
||||
"type": "number",
|
||||
"description": "Maximum context length covered by CUDA graphs. If a sequence has context length larger than this, we fall back to eager mode",
|
||||
"default": 8192,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "DISABLE_CUSTOM_ALL_REDUCE",
|
||||
"input": {
|
||||
@@ -945,7 +700,17 @@
|
||||
"name": "Max Concurrency",
|
||||
"type": "number",
|
||||
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
|
||||
"default": 300,
|
||||
"default": 30,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "ENABLE_EXPERT_PARALLEL",
|
||||
"input": {
|
||||
"name": "Enable Expert Parallel",
|
||||
"type": "boolean",
|
||||
"description": "Enable Expert Parallel for MoE models",
|
||||
"default": false,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
@@ -968,16 +733,6 @@
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "DISABLE_LOG_REQUESTS",
|
||||
"input": {
|
||||
"name": "Disable Log Requests",
|
||||
"type": "boolean",
|
||||
"description": "Enables or disables vLLM request logging",
|
||||
"default": true,
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "ENABLE_AUTO_TOOL_CHOICE",
|
||||
"input": {
|
||||
@@ -1023,6 +778,23 @@
|
||||
"default": "",
|
||||
"advanced": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"key": "REASONING_PARSER",
|
||||
"input": {
|
||||
"name": "Reasoning Parser",
|
||||
"type": "string",
|
||||
"description": "Parser for reasoning-capable models (enables reasoning mode)",
|
||||
"options": [
|
||||
{ "label": "None", "value": "" },
|
||||
{ "label": "DeepSeek R1", "value": "deepseek_r1" },
|
||||
{ "label": "Qwen3", "value": "qwen3" },
|
||||
{ "label": "Granite", "value": "granite" },
|
||||
{ "label": "Hunyuan A13B", "value": "hunyuan_a13b" }
|
||||
],
|
||||
"default": "",
|
||||
"advanced": true
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
@@ -12,7 +12,6 @@
|
||||
"input": {
|
||||
"openai_route": "/v1/chat/completions",
|
||||
"openai_input": {
|
||||
"model": "HuggingFaceTB/SmolLM2-135M-Instruct",
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
@@ -23,8 +22,8 @@
|
||||
"content": "Explain what a neural network is in one sentence."
|
||||
}
|
||||
],
|
||||
"max_tokens": 50,
|
||||
"temperature": 0.7
|
||||
"max_tokens": 200,
|
||||
"temperature": 0.1
|
||||
}
|
||||
},
|
||||
"timeout": 30000
|
||||
@@ -39,14 +38,6 @@
|
||||
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
|
||||
}
|
||||
],
|
||||
"allowedCudaVersions": [
|
||||
"12.7",
|
||||
"12.6",
|
||||
"12.5",
|
||||
"12.4",
|
||||
"12.3",
|
||||
"12.2",
|
||||
"12.1"
|
||||
]
|
||||
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
|
||||
}
|
||||
}
|
||||
+23
-9
@@ -1,20 +1,21 @@
|
||||
FROM nvidia/cuda:12.1.0-base-ubuntu22.04
|
||||
FROM nvidia/cuda:12.9.1-base-ubuntu22.04
|
||||
|
||||
RUN apt-get update -y \
|
||||
&& apt-get install -y python3-pip
|
||||
|
||||
RUN ldconfig /usr/local/cuda-12.1/compat/
|
||||
RUN ldconfig /usr/local/cuda-12.9/compat/
|
||||
|
||||
# Install Python dependencies
|
||||
# Install vLLM with FlashInfer from the CUDA 12.9 wheel index.
|
||||
RUN python3 -m pip install --upgrade pip && \
|
||||
python3 -m pip install "vllm[flashinfer]==0.17.0" --extra-index-url https://download.pytorch.org/whl/cu129
|
||||
|
||||
|
||||
|
||||
# Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts)
|
||||
COPY builder/requirements.txt /requirements.txt
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
python3 -m pip install --upgrade pip && \
|
||||
python3 -m pip install --upgrade -r /requirements.txt
|
||||
|
||||
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
|
||||
RUN python3 -m pip install vllm==0.10.0 && \
|
||||
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
|
||||
|
||||
# Setup for Option 2: Building the Image with the Model included
|
||||
ARG MODEL_NAME=""
|
||||
ARG TOKENIZER_NAME=""
|
||||
@@ -22,6 +23,7 @@ ARG BASE_PATH="/runpod-volume"
|
||||
ARG QUANTIZATION=""
|
||||
ARG MODEL_REVISION=""
|
||||
ARG TOKENIZER_REVISION=""
|
||||
ARG VLLM_NIGHTLY="false"
|
||||
|
||||
ENV MODEL_NAME=$MODEL_NAME \
|
||||
MODEL_REVISION=$MODEL_REVISION \
|
||||
@@ -32,10 +34,22 @@ ENV MODEL_NAME=$MODEL_NAME \
|
||||
HF_DATASETS_CACHE="${BASE_PATH}/huggingface-cache/datasets" \
|
||||
HUGGINGFACE_HUB_CACHE="${BASE_PATH}/huggingface-cache/hub" \
|
||||
HF_HOME="${BASE_PATH}/huggingface-cache/hub" \
|
||||
HF_HUB_ENABLE_HF_TRANSFER=0
|
||||
HF_HUB_ENABLE_HF_TRANSFER=0 \
|
||||
# Suppress Ray metrics agent warnings (not needed in containerized environments)
|
||||
RAY_METRICS_EXPORT_ENABLED=0 \
|
||||
RAY_DISABLE_USAGE_STATS=1 \
|
||||
# Prevent rayon thread pool panic in containers where ulimit -u < nproc
|
||||
# (tokenizers uses Rust's rayon which tries to spawn threads = CPU cores)
|
||||
TOKENIZERS_PARALLELISM=false \
|
||||
RAYON_NUM_THREADS=4
|
||||
|
||||
ENV PYTHONPATH="/:/vllm-workspace"
|
||||
|
||||
RUN if [ "${VLLM_NIGHTLY}" = "true" ]; then \
|
||||
pip install -U vllm --pre --index-url https://pypi.org/simple --extra-index-url https://wheels.vllm.ai/nightly && \
|
||||
apt-get update && apt-get install -y git && rm -rf /var/lib/apt/lists/* && \
|
||||
pip install git+https://github.com/huggingface/transformers.git; \
|
||||
fi
|
||||
|
||||
COPY src /src
|
||||
RUN --mount=type=secret,id=HF_TOKEN,required=false \
|
||||
|
||||
@@ -9,16 +9,10 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
||||
## Table of Contents
|
||||
|
||||
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
|
||||
- [Option 1: Deploy Any Model Using Pre-Built Docker Image **[RECOMMENDED]**](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
||||
- [Environment Variables](#environment-variables)
|
||||
- [LLM Settings](#llm-settings)
|
||||
- [Tokenizer Settings](#tokenizer-settings)
|
||||
- [System and Parallelism Settings](#system-and-parallelism-settings)
|
||||
- [Streaming Batch Size Settings](#streaming-batch-size-settings)
|
||||
- [OpenAI Settings](#openai-settings)
|
||||
- [Serverless Settings](#serverless-settings)
|
||||
- [Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended]](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
||||
- [Configuration](#configuration)
|
||||
- [Option 2: Build Docker Image with Model Inside](#option-2-build-docker-image-with-model-inside)
|
||||
- [Prerequisites](#prerequisites-1)
|
||||
- [Prerequisites](#prerequisites)
|
||||
- [Arguments](#arguments)
|
||||
- [Example: Building an image with OpenChat-3.5](#example-building-an-image-with-openchat-35)
|
||||
- [(Optional) Including Huggingface Token](#optional-including-huggingface-token)
|
||||
@@ -26,12 +20,14 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
||||
- [Usage: OpenAI Compatibility](#usage-openai-compatibility)
|
||||
- [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker)
|
||||
- [OpenAI Request Input Parameters](#openai-request-input-parameters)
|
||||
- [Chat Completions](#chat-completions)
|
||||
- [Chat Completions [RECOMMENDED]](#chat-completions-recommended)
|
||||
- [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai)
|
||||
- [Usage: standard](#non-openai-usage)
|
||||
- [Input Request Parameters](#input-request-parameters)
|
||||
- [Text Input Formats](#text-input-formats)
|
||||
- [Chat Completions](#chat-completions)
|
||||
- [Getting a list of names for available models](#getting-a-list-of-names-for-available-models)
|
||||
- [Usage: Standard (Non-OpenAI)](#usage-standard-non-openai)
|
||||
- [Request Input Parameters](#request-input-parameters)
|
||||
- [Sampling Parameters](#sampling-parameters)
|
||||
- [Text Input Formats](#text-input-formats)
|
||||
|
||||
# Setting up the Serverless Worker
|
||||
|
||||
@@ -39,129 +35,41 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
||||
|
||||
**🚀 Deploy Guide**: Follow our [step-by-step deployment guide](https://docs.runpod.io/serverless/vllm/get-started) to deploy using the RunPod Console.
|
||||
|
||||
**📦 Docker Image**: `runpod/worker-v1-vllm:<version>stable-cuda12.1.0`
|
||||
**📦 Docker Image**: `runpod/worker-v1-vllm:<version>`
|
||||
|
||||
- **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases)
|
||||
- **CUDA Compatibility**: Requires CUDA >= 12.1
|
||||
|
||||
### Environment Variables
|
||||
### Configuration
|
||||
|
||||
Use these to configure worker-vllm so it works for your use case / model.
|
||||
Configure worker-vllm using environment variables:
|
||||
|
||||
#### LLM Settings
|
||||
| Environment Variable | Description | Default | Options |
|
||||
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
||||
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
||||
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
||||
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
||||
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
||||
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
||||
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
||||
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ------------------------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
|
||||
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
|
||||
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
|
||||
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
|
||||
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
|
||||
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
|
||||
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
|
||||
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
|
||||
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
|
||||
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
|
||||
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
|
||||
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
|
||||
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
|
||||
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
|
||||
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
|
||||
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
|
||||
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
|
||||
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
|
||||
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
|
||||
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
|
||||
| `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. |
|
||||
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
|
||||
| `SEED` | 0 | `int` | Random seed for operations. |
|
||||
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
|
||||
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
|
||||
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
|
||||
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
|
||||
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
|
||||
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
|
||||
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
|
||||
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
|
||||
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
|
||||
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
|
||||
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
|
||||
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
|
||||
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
|
||||
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
|
||||
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
|
||||
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
|
||||
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}` |
|
||||
| `SCHEDULER_DELAY_FACTOR` | 0.0 | `float` | Apply a delay before scheduling next prompt. |
|
||||
| `ENABLE_CHUNKED_PREFILL` | False | `bool` | Enable chunked prefill requests. |
|
||||
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
|
||||
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
|
||||
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
|
||||
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
|
||||
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
|
||||
| `SPEC_DECODING_ACCEPTANCE_METHOD` | 'rejection_sampler' | ['rejection_sampler', 'typical_acceptance_sampler'] | Specify the acceptance method for draft token verification in speculative decoding. |
|
||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD` | None | `float` | Set the lower bound threshold for the posterior probability of a token to be accepted. |
|
||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA` | None | `float` | A scaling factor for the entropy-based threshold for token acceptance. |
|
||||
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
|
||||
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
|
||||
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
|
||||
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
|
||||
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
|
||||
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
|
||||
**Pass any vLLM engine arg** not listed above by setting an environment variable with the **UPPERCASED** field name (same names vLLM uses). The worker auto-discovers all `AsyncEngineArgs` fields from env. For example:
|
||||
|
||||
#### Tokenizer Settings
|
||||
| Environment Variable | vLLM Engine Arg | Example Value |
|
||||
| ------------------------- | ------------------------ | ------------- |
|
||||
| `MAX_MODEL_LEN` | `max_model_len` | `4096` |
|
||||
| `ENFORCE_EAGER` | `enforce_eager` | `true` |
|
||||
| `ENABLE_CHUNKED_PREFILL` | `enable_chunked_prefill` | `true` |
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ---------------------- | --------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
|
||||
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
|
||||
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
|
||||
Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is applied automatically. Backward-compat aliases: `MODEL_NAME`, `TOKENIZER_NAME`, `MAX_CONTEXT_LEN_TO_CAPTURE`. This lets you configure any vLLM option without waiting for explicit worker support.
|
||||
|
||||
#### System and Parallelism Settings
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ------------------------------ | --------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
|
||||
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
|
||||
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
|
||||
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
||||
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
|
||||
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
|
||||
|
||||
#### Streaming Batch Size Settings
|
||||
|
||||
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ---------------------------------- | --------- | -------------- | --------------------------------------------------------------------------------------------------------- |
|
||||
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
|
||||
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
|
||||
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
|
||||
|
||||
#### OpenAI Settings
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
|
||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||
|
||||
#### Serverless Settings
|
||||
|
||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
||||
| ---------------------- | --------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
||||
|
||||
## Option 2: Build Docker Image with Model Inside
|
||||
|
||||
@@ -182,6 +90,7 @@ To build an image with the model baked in, you must specify the following docker
|
||||
- `WORKER_CUDA_VERSION`: `12.1.0` (`12.1.0` is recommended for optimal performance).
|
||||
- `TOKENIZER_NAME`: Tokenizer repository if you would like to use a different tokenizer than the one that comes with the model. (default: `None`, which uses the model's tokenizer)
|
||||
- `TOKENIZER_REVISION`: Tokenizer revision to load (default: `main`).
|
||||
- `VLLM_NIGHTLY`: Set to `true` to replace the pinned vLLM release with the latest nightly build and the latest `transformers` from source. Useful for testing unreleased vLLM features. (default: `false`)
|
||||
|
||||
For the remaining settings, you may apply them as environment variables when running the container. Supported environment variables are listed in the [Environment Variables](#environment-variables) section.
|
||||
|
||||
@@ -191,6 +100,20 @@ For the remaining settings, you may apply them as environment variables when run
|
||||
docker build -t username/image:tag --build-arg MODEL_NAME="openchat/openchat_3.5" --build-arg BASE_PATH="/models" .
|
||||
```
|
||||
|
||||
### Example: Building with vLLM Nightly
|
||||
|
||||
To use the latest unreleased vLLM build (installs from the nightly wheel index and `transformers` from source):
|
||||
|
||||
```bash
|
||||
docker build -t username/image:tag --build-arg VLLM_NIGHTLY=true .
|
||||
```
|
||||
|
||||
You can combine it with other arguments:
|
||||
|
||||
```bash
|
||||
docker build -t username/image:tag --build-arg VLLM_NIGHTLY=true --build-arg MODEL_NAME="meta-llama/Llama-3.1-8B-Instruct" --build-arg BASE_PATH="/models" .
|
||||
```
|
||||
|
||||
### (Optional) Including Huggingface Token
|
||||
|
||||
If the model you would like to deploy is private or gated, you will need to include it during build time as a Docker secret, which will protect it from being exposed in the image and on DockerHub.
|
||||
|
||||
@@ -1,14 +1,14 @@
|
||||
ray
|
||||
pandas
|
||||
pyarrow
|
||||
runpod~=1.7.7
|
||||
runpod
|
||||
huggingface-hub
|
||||
packaging
|
||||
typing-extensions>=4.8.0
|
||||
pydantic
|
||||
pydantic-settings
|
||||
hf-transfer
|
||||
transformers>=4.55.0
|
||||
transformers>=4.57.0
|
||||
bitsandbytes>=0.45.0
|
||||
kernels
|
||||
torch==2.6.0
|
||||
torch-c-dlpack-ext
|
||||
|
||||
@@ -0,0 +1,202 @@
|
||||
# Configuration Reference
|
||||
|
||||
Complete guide to all environment variables and configuration options for worker-vllm.
|
||||
|
||||
## LLM Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------- |
|
||||
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
|
||||
| `MODEL_REVISION` | 'main' | `str` | Model revision to load (default: main). |
|
||||
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
|
||||
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
|
||||
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
|
||||
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
|
||||
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
|
||||
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
|
||||
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
|
||||
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
|
||||
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
|
||||
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
|
||||
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
|
||||
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
|
||||
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
|
||||
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
|
||||
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
|
||||
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
|
||||
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
|
||||
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
|
||||
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
|
||||
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
|
||||
| `SEED` | 0 | `int` | Random seed for operations. |
|
||||
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
|
||||
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
|
||||
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
|
||||
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
|
||||
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
|
||||
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
|
||||
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
|
||||
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
|
||||
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
|
||||
|
||||
## LoRA (Low-Rank Adaptation) Settings
|
||||
|
||||
| Variable | Default | Type | Description |
|
||||
| --------------------------- | ------- | ------------------------------------------ | --------------------------------------------------------------------------------------------------------- |
|
||||
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
|
||||
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
|
||||
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
|
||||
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
|
||||
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
|
||||
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
|
||||
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
|
||||
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
|
||||
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}]` |
|
||||
|
||||
> **Note (Serverless)**: When LoRA adapters are configured via `LORA_MODULES`, initialization is deferred to the first request to ensure compatibility with RunPod Serverless. This means the first request will include LoRA loading time. Subsequent requests are unaffected. Check logs for "LoRA mode: X adapter(s) will load on first request" at startup.
|
||||
|
||||
## Speculative Decoding Settings
|
||||
|
||||
Speculative decoding can be configured in two ways:
|
||||
|
||||
### Option 1: JSON Configuration
|
||||
|
||||
Set `SPECULATIVE_CONFIG` to a JSON string with your full speculative decoding configuration:
|
||||
|
||||
```bash
|
||||
SPECULATIVE_CONFIG='{"method": "ngram", "num_speculative_tokens": 5, "prompt_lookup_max": 4}'
|
||||
```
|
||||
|
||||
### Option 2: Individual Environment Variables
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------------------------- | ------- | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------- |
|
||||
| `SPECULATIVE_METHOD` | None | ['draft_model', 'ngram', 'eagle', 'eagle3', 'medusa', 'mlp_speculator'] | Speculative decoding method to use. |
|
||||
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
|
||||
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
|
||||
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
|
||||
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
|
||||
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
|
||||
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
|
||||
|
||||
If `SPECULATIVE_CONFIG` is set, it takes priority over individual env vars. When using individual env vars without `SPECULATIVE_METHOD`, the method is auto-detected from the model name or configuration.
|
||||
|
||||
## Scheduling & Performance Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ------------------------------ | ------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
|
||||
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
|
||||
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
|
||||
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
||||
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
|
||||
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
|
||||
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models. |
|
||||
| `ATTENTION_BACKEND` | `None` | `str` | Attention backend to use (e.g., `FLASH_ATTN`, `FLASHINFER`, `TRITON_FLASH_ATTN`). Replaces deprecated `VLLM_ATTENTION_BACKEND`. |
|
||||
| `ASYNC_SCHEDULING` | `None` | `bool` | Enable async scheduling (overlaps engine scheduling with GPU execution). Default: enabled in vLLM 0.14.0+. Set to `false` to disable. |
|
||||
| `STREAM_INTERVAL` | `1` | `int` | Controls how often to yield streaming results. Lower = more frequent updates. |
|
||||
|
||||
## Tokenizer Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------- | ------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
|
||||
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
|
||||
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
|
||||
|
||||
## Streaming & Batch Settings
|
||||
|
||||
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker.
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------------------- | ------- | ------------ | --------------------------------------------------------------------------------------------------------- |
|
||||
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
|
||||
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
|
||||
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
|
||||
|
||||
## OpenAI Compatibility Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
|
||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
|
||||
| `TRUST_REQUEST_CHAT_TEMPLATE` | `false` | `bool` | Allow clients to send custom chat templates in API requests. **Security consideration:** Only enable if you trust your API clients. |
|
||||
| `RETURN_TOKENS_AS_TOKEN_IDS` | `false` | `bool` | Return token IDs instead of decoded text strings in responses. |
|
||||
| `EXCLUDE_TOOLS_WHEN_TOOL_CHOICE_NONE` | `false` | `bool` | Exclude tool definitions from the prompt when `tool_choice` is set to `none`. |
|
||||
| `ENABLE_PROMPT_TOKENS_DETAILS` | `false` | `bool` | Include detailed prompt token information in API responses. |
|
||||
| `ENABLE_FORCE_INCLUDE_USAGE` | `false` | `bool` | Always include usage statistics in API responses, even when not requested. |
|
||||
| `ENABLE_LOG_OUTPUTS` | `false` | `bool` | Log model outputs for debugging purposes. |
|
||||
| `LOG_ERROR_STACK` | `false` | `bool` | Include full stack traces in error responses for debugging. |
|
||||
|
||||
## Serverless & Concurrency Settings
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||
| `ENABLE_LOG_REQUESTS` | False | `bool` | Enables vLLM request logging. (Replaces deprecated `DISABLE_LOG_REQUESTS` in vLLM 0.15.0) |
|
||||
|
||||
## Advanced Settings
|
||||
|
||||
| Variable | Default | Type | Description |
|
||||
| --------------------------- | ------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
|
||||
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
|
||||
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
|
||||
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
|
||||
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
|
||||
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
|
||||
|
||||
## UPPERCASED env vars: Pass any engine arg
|
||||
|
||||
Any vLLM `AsyncEngineArgs` field can be set via an environment variable using the **UPPERCASED** field name (the same names vLLM uses). The worker auto-discovers all fields from env — no prefix.
|
||||
|
||||
**Format:** `<FIELD_NAME_UPPERCASED>=<value>` (e.g. `MAX_MODEL_LEN=4096`)
|
||||
|
||||
**Examples:**
|
||||
|
||||
| Environment Variable | vLLM Engine Arg | Value Example |
|
||||
| ------------------------ | ------------------------ | ------------- |
|
||||
| `MAX_MODEL_LEN` | `max_model_len` | `4096` |
|
||||
| `ENFORCE_EAGER` | `enforce_eager` | `true` |
|
||||
| `ENABLE_CHUNKED_PREFILL` | `enable_chunked_prefill` | `true` |
|
||||
| `NUM_SCHEDULER_STEPS` | `num_scheduler_steps` | `8` |
|
||||
| `TOKENIZER_POOL_SIZE` | `tokenizer_pool_size` | `4` |
|
||||
|
||||
**Backward-compat aliases:** `MODEL_NAME` → `model`, `TOKENIZER_NAME` → `tokenizer`, `MAX_CONTEXT_LEN_TO_CAPTURE` → `max_seq_len_to_capture`, `MODEL_REVISION` → `revision`.
|
||||
|
||||
**Notes:**
|
||||
- Only valid `AsyncEngineArgs` fields are applied. Unknown keys are silently ignored.
|
||||
- Values are automatically cast to the correct type (`int`, `float`, `bool`, `str`, or JSON for `dict`/`list`/`tuple`).
|
||||
- For a full list of available engine args, see the [vLLM AsyncEngineArgs documentation](https://docs.vllm.ai/en/latest/configuration/engine_args/).
|
||||
|
||||
## Docker Build Arguments
|
||||
|
||||
These variables are used when building custom Docker images with models baked in:
|
||||
|
||||
| Variable | Default | Type | Description |
|
||||
| --------------------- | ---------------- | ----- | ------------------------------------------------- |
|
||||
| `BASE_PATH` | `/runpod-volume` | `str` | Storage directory for huggingface cache and model |
|
||||
| `WORKER_CUDA_VERSION` | `12.1.0` | `str` | CUDA version for the worker image |
|
||||
|
||||
## Deprecated Variables
|
||||
|
||||
⚠️ **The following variables are deprecated and will be removed in future versions:**
|
||||
|
||||
| Old Variable | New Variable | Note |
|
||||
| ---------------------------- | ------------------------ | -------------------------------------------------------------------- |
|
||||
| `MAX_CONTEXT_LEN_TO_CAPTURE` | `MAX_SEQ_LEN_TO_CAPTURE` | Use new variable name |
|
||||
| `kv_cache_dtype=fp8_e5m2` | `kv_cache_dtype=fp8` | Simplified fp8 format |
|
||||
| `USE_V2_BLOCK_MANAGER` | *(removed)* | V2 block manager is now the default in vLLM 0.13.0, setting ignored |
|
||||
| `VLLM_ATTENTION_BACKEND` | `ATTENTION_BACKEND` | Use new env var name (old still works with deprecation warning) |
|
||||
| `DISABLE_LOG_REQUESTS` | `ENABLE_LOG_REQUESTS` | Inverted logic in vLLM 0.15.0 (old still works with deprecation warning) |
|
||||
|
||||
+3
-2
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
|
||||
|
||||
- `src/engine_args.py`: Centralized configuration management
|
||||
- `src/constants.py`: Default values for core settings
|
||||
- `worker-config.json`: UI form generation for RunPod console
|
||||
- `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
|
||||
- `worker-config.json`: UI form generation for RunPod console (if exists)
|
||||
|
||||
## Core Development Concepts
|
||||
|
||||
@@ -222,7 +223,7 @@ src/
|
||||
|
||||
### 2. **Concurrency Patterns**
|
||||
|
||||
- **Max Concurrency**: 300 concurrent requests by default
|
||||
- **Max Concurrency**: 30 concurrent requests by default
|
||||
- **vLLM Queuing**: Internal request batching and scheduling
|
||||
- **RunPod Integration**: Concurrency modifier for auto-scaling
|
||||
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 86 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 1.1 MiB |
+1
-1
@@ -1,4 +1,4 @@
|
||||
DEFAULT_BATCH_SIZE = 50
|
||||
DEFAULT_MAX_CONCURRENCY = 300
|
||||
DEFAULT_MAX_CONCURRENCY = 30
|
||||
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
|
||||
DEFAULT_MIN_BATCH_SIZE = 1
|
||||
+65
-26
@@ -1,24 +1,25 @@
|
||||
import os
|
||||
import logging
|
||||
import json
|
||||
import asyncio
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
from typing import AsyncGenerator, Optional
|
||||
|
||||
from dotenv import load_dotenv
|
||||
from typing import AsyncGenerator, Optional
|
||||
import time
|
||||
|
||||
from vllm import AsyncLLMEngine
|
||||
from vllm.entrypoints.logger import RequestLogger
|
||||
from vllm.entrypoints.openai.serving_chat import OpenAIServingChat
|
||||
from vllm.entrypoints.openai.serving_completion import OpenAIServingCompletion
|
||||
from vllm.entrypoints.openai.protocol import ChatCompletionRequest, CompletionRequest, ErrorResponse
|
||||
from vllm.entrypoints.openai.serving_models import BaseModelPath, LoRAModulePath, OpenAIServingModels
|
||||
from vllm.entrypoints.openai.chat_completion.protocol import ChatCompletionRequest
|
||||
from vllm.entrypoints.openai.chat_completion.serving import OpenAIServingChat
|
||||
from vllm.entrypoints.openai.completion.protocol import CompletionRequest
|
||||
from vllm.entrypoints.openai.completion.serving import OpenAIServingCompletion
|
||||
from vllm.entrypoints.openai.engine.protocol import ErrorResponse
|
||||
from vllm.entrypoints.openai.models.protocol import BaseModelPath, LoRAModulePath
|
||||
from vllm.entrypoints.openai.models.serving import OpenAIServingModels
|
||||
|
||||
|
||||
from utils import DummyRequest, JobInput, BatchSize, create_error_response
|
||||
from constants import DEFAULT_MAX_CONCURRENCY, DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MIN_BATCH_SIZE
|
||||
from tokenizer import TokenizerWrapper
|
||||
from constants import DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MAX_CONCURRENCY, DEFAULT_MIN_BATCH_SIZE
|
||||
from engine_args import get_engine_args
|
||||
from tokenizer import TokenizerWrapper
|
||||
from utils import BatchSize, DummyRequest, JobInput, create_error_response
|
||||
|
||||
class vLLMEngine:
|
||||
def __init__(self, engine = None):
|
||||
@@ -174,10 +175,24 @@ class vLLMEngine:
|
||||
class OpenAIvLLMEngine(vLLMEngine):
|
||||
def __init__(self, vllm_engine):
|
||||
super().__init__(vllm_engine)
|
||||
self.served_model_name = os.getenv("OPENAI_SERVED_MODEL_NAME_OVERRIDE") or self.engine_args.model
|
||||
self.served_model_name = os.getenv("OPENAI_SERVED_MODEL_NAME_OVERRIDE") or self.engine_args.served_model_name or self.engine_args.model
|
||||
self.response_role = os.getenv("OPENAI_RESPONSE_ROLE") or "assistant"
|
||||
self.lora_adapters = self._load_lora_adapters()
|
||||
asyncio.run(self._initialize_engines())
|
||||
|
||||
# Always defer OpenAI engine initialization to the first request.
|
||||
# asyncio.run() creates a temporary event loop that gets closed, but async
|
||||
# components (tokenizer pool, serving engines) bind futures to that loop.
|
||||
# When Runpod's serverless handler runs in its own event loop, those futures
|
||||
# are "attached to a different loop" causing RuntimeError.
|
||||
# This affects all configurations, not just LoRA.
|
||||
self._engines_initialized = False
|
||||
if self.lora_adapters:
|
||||
logging.info(f"LoRA mode: {len(self.lora_adapters)} adapter(s) will load on first request")
|
||||
for adapter in self.lora_adapters:
|
||||
logging.info(f" - {adapter.name}: {adapter.path}")
|
||||
else:
|
||||
logging.info("OpenAI engines will initialize on first request")
|
||||
|
||||
# Handle both integer and boolean string values for RAW_OPENAI_OUTPUT
|
||||
raw_output_env = os.getenv("RAW_OPENAI_OUTPUT", "1")
|
||||
if raw_output_env.lower() in ('true', 'false'):
|
||||
@@ -201,15 +216,28 @@ class OpenAIvLLMEngine(vLLMEngine):
|
||||
continue
|
||||
return adapters
|
||||
|
||||
async def _ensure_engines_initialized(self):
|
||||
"""Initialize engines on first request to avoid event loop mismatch.
|
||||
|
||||
In Runpod Serverless, the startup code runs outside the handler's event
|
||||
loop. Deferring initialization to the first request ensures all async
|
||||
components (tokenizer pool, serving engines, LoRA state) are created in
|
||||
the correct event loop context.
|
||||
"""
|
||||
if not self._engines_initialized:
|
||||
logging.info("Initializing OpenAI serving engines...")
|
||||
await self._initialize_engines()
|
||||
self._engines_initialized = True
|
||||
logging.info("OpenAI serving engines initialized successfully")
|
||||
|
||||
async def _initialize_engines(self):
|
||||
self.model_config = await self.llm.get_model_config()
|
||||
self.model_config = self.llm.model_config
|
||||
self.base_model_paths = [
|
||||
BaseModelPath(name=self.engine_args.model, model_path=self.engine_args.model)
|
||||
BaseModelPath(name=self.served_model_name, model_path=self.engine_args.model)
|
||||
]
|
||||
|
||||
self.serving_models = OpenAIServingModels(
|
||||
engine_client=self.llm,
|
||||
model_config=self.model_config,
|
||||
base_model_paths=self.base_model_paths,
|
||||
lora_modules=self.lora_adapters,
|
||||
)
|
||||
@@ -222,28 +250,39 @@ class OpenAIvLLMEngine(vLLMEngine):
|
||||
|
||||
self.chat_engine = OpenAIServingChat(
|
||||
engine_client=self.llm,
|
||||
model_config=self.model_config,
|
||||
models=self.serving_models,
|
||||
response_role=self.response_role,
|
||||
request_logger=None,
|
||||
chat_template=chat_template,
|
||||
chat_template_content_format="auto",
|
||||
# enable_reasoning=os.getenv('ENABLE_REASONING', 'false').lower() == 'true',
|
||||
reasoning_parser= os.getenv('REASONING_PARSER', "") or None,
|
||||
# return_token_as_token_ids=False,
|
||||
trust_request_chat_template=os.getenv('TRUST_REQUEST_CHAT_TEMPLATE', 'false').lower() == 'true',
|
||||
return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
|
||||
reasoning_parser=os.getenv('REASONING_PARSER', "") or "",
|
||||
enable_auto_tools=os.getenv('ENABLE_AUTO_TOOL_CHOICE', 'false').lower() == 'true',
|
||||
exclude_tools_when_tool_choice_none=os.getenv('EXCLUDE_TOOLS_WHEN_TOOL_CHOICE_NONE', 'false').lower() == 'true',
|
||||
tool_parser=os.getenv('TOOL_CALL_PARSER', "") or None,
|
||||
enable_prompt_tokens_details=False
|
||||
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
|
||||
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
|
||||
enable_log_outputs=os.getenv('ENABLE_LOG_OUTPUTS', 'false').lower() == 'true',
|
||||
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
|
||||
)
|
||||
self.completion_engine = OpenAIServingCompletion(
|
||||
engine_client=self.llm,
|
||||
model_config=self.model_config,
|
||||
models=self.serving_models,
|
||||
request_logger=None,
|
||||
# return_token_as_token_ids=False,
|
||||
return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
|
||||
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
|
||||
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
|
||||
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
|
||||
)
|
||||
|
||||
if hasattr(self.chat_engine, 'warmup'):
|
||||
await self.chat_engine.warmup()
|
||||
|
||||
async def generate(self, openai_request: JobInput):
|
||||
# Ensure engines are ready (no-op if already initialized at startup)
|
||||
await self._ensure_engines_initialized()
|
||||
|
||||
if openai_request.openai_route == "/v1/models":
|
||||
yield await self._handle_model_request()
|
||||
elif openai_request.openai_route in ["/v1/chat/completions", "/v1/completions"]:
|
||||
|
||||
+387
-104
@@ -1,117 +1,335 @@
|
||||
import os
|
||||
import json
|
||||
import logging
|
||||
from typing import get_origin, get_args
|
||||
from torch.cuda import device_count
|
||||
from vllm import AsyncEngineArgs
|
||||
from vllm.model_executor.model_loader.tensorizer import TensorizerConfig
|
||||
from src.utils import convert_limit_mm_per_prompt
|
||||
|
||||
RENAME_ARGS_MAP = {
|
||||
# Backward-compat: env var names users already know → engine arg name
|
||||
ENV_ALIASES = {
|
||||
"MODEL_NAME": "model",
|
||||
"MODEL_REVISION": "revision",
|
||||
"TOKENIZER_NAME": "tokenizer",
|
||||
"MAX_CONTEXT_LEN_TO_CAPTURE": "max_seq_len_to_capture"
|
||||
}
|
||||
|
||||
# Literal defaults from original worker (used when env/local do not set a value)
|
||||
DEFAULT_ARGS = {
|
||||
"disable_log_stats": os.getenv('DISABLE_LOG_STATS', 'False').lower() == 'true',
|
||||
"disable_log_requests": os.getenv('DISABLE_LOG_REQUESTS', 'False').lower() == 'true',
|
||||
"gpu_memory_utilization": float(os.getenv('GPU_MEMORY_UTILIZATION', 0.95)),
|
||||
"pipeline_parallel_size": int(os.getenv('PIPELINE_PARALLEL_SIZE', 1)),
|
||||
"tensor_parallel_size": int(os.getenv('TENSOR_PARALLEL_SIZE', 1)),
|
||||
"served_model_name": os.getenv('SERVED_MODEL_NAME', None),
|
||||
"tokenizer": os.getenv('TOKENIZER', None),
|
||||
"skip_tokenizer_init": os.getenv('SKIP_TOKENIZER_INIT', 'False').lower() == 'true',
|
||||
"tokenizer_mode": os.getenv('TOKENIZER_MODE', 'auto'),
|
||||
"trust_remote_code": os.getenv('TRUST_REMOTE_CODE', 'False').lower() == 'true',
|
||||
"download_dir": os.getenv('DOWNLOAD_DIR', None),
|
||||
"load_format": os.getenv('LOAD_FORMAT', 'auto'),
|
||||
"config_format": os.getenv('CONFIG_FORMAT', 'auto'),
|
||||
"dtype": os.getenv('DTYPE', 'auto'),
|
||||
"kv_cache_dtype": os.getenv('KV_CACHE_DTYPE', 'auto'),
|
||||
"quantization_param_path": os.getenv('QUANTIZATION_PARAM_PATH', None),
|
||||
"seed": int(os.getenv('SEED', 0)),
|
||||
"max_model_len": int(os.getenv('MAX_MODEL_LEN', 0)) or None,
|
||||
"worker_use_ray": os.getenv('WORKER_USE_RAY', 'False').lower() == 'true',
|
||||
"distributed_executor_backend": os.getenv('DISTRIBUTED_EXECUTOR_BACKEND', None),
|
||||
"max_parallel_loading_workers": int(os.getenv('MAX_PARALLEL_LOADING_WORKERS', 0)) or None,
|
||||
"block_size": int(os.getenv('BLOCK_SIZE', 16)),
|
||||
"enable_prefix_caching": os.getenv('ENABLE_PREFIX_CACHING', 'False').lower() == 'true',
|
||||
"disable_sliding_window": os.getenv('DISABLE_SLIDING_WINDOW', 'False').lower() == 'true',
|
||||
"use_v2_block_manager": os.getenv('USE_V2_BLOCK_MANAGER', 'False').lower() == 'true',
|
||||
"swap_space": int(os.getenv('SWAP_SPACE', 4)), # GiB
|
||||
"cpu_offload_gb": int(os.getenv('CPU_OFFLOAD_GB', 0)), # GiB
|
||||
"max_num_batched_tokens": int(os.getenv('MAX_NUM_BATCHED_TOKENS', 0)) or None,
|
||||
"max_num_seqs": int(os.getenv('MAX_NUM_SEQS', 256)),
|
||||
"max_logprobs": int(os.getenv('MAX_LOGPROBS', 20)), # Default value for OpenAI Chat Completions API
|
||||
"revision": os.getenv('REVISION', None),
|
||||
"code_revision": os.getenv('CODE_REVISION', None),
|
||||
"rope_scaling": os.getenv('ROPE_SCALING', None),
|
||||
"rope_theta": float(os.getenv('ROPE_THETA', 0)) or None,
|
||||
"tokenizer_revision": os.getenv('TOKENIZER_REVISION', None),
|
||||
"quantization": os.getenv('QUANTIZATION', None),
|
||||
"enforce_eager": os.getenv('ENFORCE_EAGER', 'False').lower() == 'true',
|
||||
"max_context_len_to_capture": int(os.getenv('MAX_CONTEXT_LEN_TO_CAPTURE', 0)) or None,
|
||||
"max_seq_len_to_capture": int(os.getenv('MAX_SEQ_LEN_TO_CAPTURE', 8192)),
|
||||
"disable_custom_all_reduce": os.getenv('DISABLE_CUSTOM_ALL_REDUCE', 'False').lower() == 'true',
|
||||
"tokenizer_pool_size": int(os.getenv('TOKENIZER_POOL_SIZE', 0)),
|
||||
"tokenizer_pool_type": os.getenv('TOKENIZER_POOL_TYPE', 'ray'),
|
||||
"tokenizer_pool_extra_config": os.getenv('TOKENIZER_POOL_EXTRA_CONFIG', None),
|
||||
"enable_lora": os.getenv('ENABLE_LORA', 'False').lower() == 'true',
|
||||
"max_loras": int(os.getenv('MAX_LORAS', 1)),
|
||||
"max_lora_rank": int(os.getenv('MAX_LORA_RANK', 16)),
|
||||
"enable_prompt_adapter": os.getenv('ENABLE_PROMPT_ADAPTER', 'False').lower() == 'true',
|
||||
"max_prompt_adapters": int(os.getenv('MAX_PROMPT_ADAPTERS', 1)),
|
||||
"max_prompt_adapter_token": int(os.getenv('MAX_PROMPT_ADAPTER_TOKEN', 0)),
|
||||
"fully_sharded_loras": os.getenv('FULLY_SHARDED_LORAS', 'False').lower() == 'true',
|
||||
"lora_extra_vocab_size": int(os.getenv('LORA_EXTRA_VOCAB_SIZE', 256)),
|
||||
"long_lora_scaling_factors": tuple(map(float, os.getenv('LONG_LORA_SCALING_FACTORS', '').split(','))) if os.getenv('LONG_LORA_SCALING_FACTORS') else None,
|
||||
"lora_dtype": os.getenv('LORA_DTYPE', 'auto'),
|
||||
"max_cpu_loras": int(os.getenv('MAX_CPU_LORAS', 0)) or None,
|
||||
"device": os.getenv('DEVICE', 'auto'),
|
||||
"ray_workers_use_nsight": os.getenv('RAY_WORKERS_USE_NSIGHT', 'False').lower() == 'true',
|
||||
"num_gpu_blocks_override": int(os.getenv('NUM_GPU_BLOCKS_OVERRIDE', 0)) or None,
|
||||
"num_lookahead_slots": int(os.getenv('NUM_LOOKAHEAD_SLOTS', 0)),
|
||||
"model_loader_extra_config": os.getenv('MODEL_LOADER_EXTRA_CONFIG', None),
|
||||
"ignore_patterns": os.getenv('IGNORE_PATTERNS', None),
|
||||
"preemption_mode": os.getenv('PREEMPTION_MODE', None),
|
||||
"scheduler_delay_factor": float(os.getenv('SCHEDULER_DELAY_FACTOR', 0.0)),
|
||||
"enable_chunked_prefill": os.getenv('ENABLE_CHUNKED_PREFILL', None),
|
||||
"guided_decoding_backend": os.getenv('GUIDED_DECODING_BACKEND', 'outlines'),
|
||||
"speculative_model": os.getenv('SPECULATIVE_MODEL', None),
|
||||
"speculative_draft_tensor_parallel_size": int(os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE', 0)) or None,
|
||||
"num_speculative_tokens": int(os.getenv('NUM_SPECULATIVE_TOKENS', 0)) or None,
|
||||
"speculative_max_model_len": int(os.getenv('SPECULATIVE_MAX_MODEL_LEN', 0)) or None,
|
||||
"speculative_disable_by_batch_size": int(os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE', 0)) or None,
|
||||
"ngram_prompt_lookup_max": int(os.getenv('NGRAM_PROMPT_LOOKUP_MAX', 0)) or None,
|
||||
"ngram_prompt_lookup_min": int(os.getenv('NGRAM_PROMPT_LOOKUP_MIN', 0)) or None,
|
||||
"spec_decoding_acceptance_method": os.getenv('SPEC_DECODING_ACCEPTANCE_METHOD', 'rejection_sampler'),
|
||||
"typical_acceptance_sampler_posterior_threshold": float(os.getenv('TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD', 0)) or None,
|
||||
"typical_acceptance_sampler_posterior_alpha": float(os.getenv('TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA', 0)) or None,
|
||||
"qlora_adapter_name_or_path": os.getenv('QLORA_ADAPTER_NAME_OR_PATH', None),
|
||||
"disable_logprobs_during_spec_decoding": os.getenv('DISABLE_LOGPROBS_DURING_SPEC_DECODING', None),
|
||||
"otlp_traces_endpoint": os.getenv('OTLP_TRACES_ENDPOINT', None),
|
||||
"use_v2_block_manager": os.getenv('USE_V2_BLOCK_MANAGER', 'true'),
|
||||
"disable_log_stats": False,
|
||||
"enable_log_requests": False,
|
||||
"gpu_memory_utilization": 0.95,
|
||||
"pipeline_parallel_size": 1,
|
||||
"tensor_parallel_size": 1,
|
||||
"skip_tokenizer_init": False,
|
||||
"tokenizer_mode": "auto",
|
||||
"trust_remote_code": False,
|
||||
"load_format": "auto",
|
||||
"dtype": "auto",
|
||||
"kv_cache_dtype": "auto",
|
||||
"seed": 0,
|
||||
"worker_use_ray": False,
|
||||
"block_size": 16,
|
||||
"enable_prefix_caching": False,
|
||||
"disable_sliding_window": False,
|
||||
"swap_space": 4,
|
||||
"cpu_offload_gb": 0,
|
||||
"max_num_seqs": 256,
|
||||
"max_logprobs": 20,
|
||||
"enforce_eager": False,
|
||||
"max_seq_len_to_capture": 8192,
|
||||
"disable_custom_all_reduce": False,
|
||||
"tokenizer_pool_size": 0,
|
||||
"tokenizer_pool_type": "ray",
|
||||
"enable_lora": False,
|
||||
"max_loras": 1,
|
||||
"max_lora_rank": 16,
|
||||
"enable_prompt_adapter": False,
|
||||
"max_prompt_adapters": 1,
|
||||
"max_prompt_adapter_token": 0,
|
||||
"fully_sharded_loras": False,
|
||||
"lora_extra_vocab_size": 256,
|
||||
"lora_dtype": "auto",
|
||||
"device": "auto",
|
||||
"ray_workers_use_nsight": False,
|
||||
"num_lookahead_slots": 0,
|
||||
"scheduler_delay_factor": 0.0,
|
||||
"guided_decoding_backend": "outlines",
|
||||
"spec_decoding_acceptance_method": "rejection_sampler",
|
||||
"stream_interval": 1,
|
||||
|
||||
}
|
||||
limit_mm_env = os.getenv('LIMIT_MM_PER_PROMPT')
|
||||
if limit_mm_env is not None:
|
||||
DEFAULT_ARGS["limit_mm_per_prompt"] = convert_limit_mm_per_prompt(limit_mm_env)
|
||||
|
||||
def match_vllm_args(args):
|
||||
"""Rename args to match vllm by:
|
||||
1. Renaming keys to lower case
|
||||
2. Renaming keys to match vllm
|
||||
3. Filtering args to match vllm's AsyncEngineArgs
|
||||
|
||||
Args:
|
||||
args (dict): Dictionary of args
|
||||
def _resolve_field_type(field_type: type) -> type:
|
||||
"""Resolve Optional/Union to the concrete type for conversion."""
|
||||
origin = get_origin(field_type)
|
||||
args = get_args(field_type) if hasattr(field_type, "__args__") else ()
|
||||
if origin is not None:
|
||||
# Optional[X] is Union[X, None]; X | None is UnionType
|
||||
non_none = [a for a in args if a is not type(None)]
|
||||
if non_none:
|
||||
return non_none[0]
|
||||
return field_type
|
||||
|
||||
Returns:
|
||||
dict: Dictionary of args with renamed keys
|
||||
|
||||
def _convert_env_value_to_field_type(value: str, field_name: str, field_type: type):
|
||||
"""Convert env var string to the type expected by AsyncEngineArgs for this field."""
|
||||
val = value.strip() if isinstance(value, str) else value
|
||||
if val in ("", "None", "none"):
|
||||
args = get_args(field_type) if hasattr(field_type, "__args__") else ()
|
||||
if type(None) in (args or ()):
|
||||
return None
|
||||
raise ValueError("empty value not allowed for non-optional field")
|
||||
effective_type = _resolve_field_type(field_type)
|
||||
# bool
|
||||
if effective_type is bool:
|
||||
return str(val).lower() in ("true", "1", "yes", "on")
|
||||
# int
|
||||
if effective_type is int:
|
||||
return int(val)
|
||||
# float
|
||||
if effective_type is float:
|
||||
return float(val)
|
||||
# str
|
||||
if effective_type is str:
|
||||
return str(val)
|
||||
# dict, list, or complex (try JSON)
|
||||
origin = get_origin(effective_type)
|
||||
if effective_type in (dict, list) or origin in (dict, list):
|
||||
try:
|
||||
return json.loads(val)
|
||||
except json.JSONDecodeError:
|
||||
return val
|
||||
# tuple (e.g. long_lora_scaling_factors) — comma-separated or JSON array
|
||||
if effective_type is tuple or origin is tuple:
|
||||
args = get_args(field_type) if hasattr(field_type, "__args__") else ()
|
||||
elem_types = [a for a in args if a is not Ellipsis]
|
||||
elem_type = elem_types[0] if elem_types else str
|
||||
try:
|
||||
parsed = json.loads(val)
|
||||
if isinstance(parsed, list):
|
||||
return tuple(elem_type(x) for x in parsed)
|
||||
except (json.JSONDecodeError, TypeError):
|
||||
pass
|
||||
return tuple(elem_type(x.strip()) for x in str(val).split(",") if x.strip())
|
||||
# Fallback: try int, float, then str
|
||||
try:
|
||||
return int(val)
|
||||
except ValueError:
|
||||
pass
|
||||
try:
|
||||
return float(val)
|
||||
except ValueError:
|
||||
pass
|
||||
return str(val)
|
||||
|
||||
|
||||
def _get_args_from_env_auto_discover() -> dict:
|
||||
"""Auto-discover engine args from env vars using UPPERCASED field names.
|
||||
|
||||
For every field in AsyncEngineArgs, check os.getenv(FIELD_NAME).
|
||||
E.g. MAX_MODEL_LEN=4096 -> max_model_len=4096.
|
||||
Uses same type conversion as before; supports all vLLM engine args without manual listing.
|
||||
"""
|
||||
renamed_args = {RENAME_ARGS_MAP.get(k, k): v for k, v in args.items()}
|
||||
matched_args = {k: v for k, v in renamed_args.items() if k in AsyncEngineArgs.__dataclass_fields__}
|
||||
return {k: v for k, v in matched_args.items() if v not in [None, "", "None"]}
|
||||
args = {}
|
||||
valid_fields = AsyncEngineArgs.__dataclass_fields__
|
||||
for field_name, field in valid_fields.items():
|
||||
env_key = field_name.upper()
|
||||
value = os.environ.get(env_key)
|
||||
if value is None:
|
||||
continue
|
||||
try:
|
||||
args[field_name] = _convert_env_value_to_field_type(
|
||||
value, field_name, field.type
|
||||
)
|
||||
except (ValueError, TypeError, json.JSONDecodeError) as e:
|
||||
logging.warning(
|
||||
"Skip env %s=%r: %s", env_key, value, e
|
||||
)
|
||||
return args
|
||||
|
||||
|
||||
def _apply_env_aliases(args: dict) -> None:
|
||||
"""Apply ENV_ALIASES: if MODEL_NAME etc. are set, set the target engine arg."""
|
||||
valid_fields = AsyncEngineArgs.__dataclass_fields__
|
||||
for alias, target in ENV_ALIASES.items():
|
||||
value = os.environ.get(alias)
|
||||
if value is None or target not in valid_fields:
|
||||
continue
|
||||
try:
|
||||
args[target] = _convert_env_value_to_field_type(
|
||||
value, target, valid_fields[target].type
|
||||
)
|
||||
except (ValueError, TypeError, json.JSONDecodeError) as e:
|
||||
logging.warning("Skip env alias %s=%r: %s", alias, value, e)
|
||||
|
||||
def get_speculative_config():
|
||||
"""Build speculative decoding configuration from environment variables.
|
||||
|
||||
Supports two modes:
|
||||
1. Full JSON config via SPECULATIVE_CONFIG env var
|
||||
2. Individual env vars for common settings
|
||||
"""
|
||||
# Option 1: Full JSON configuration
|
||||
spec_config_json = os.getenv('SPECULATIVE_CONFIG')
|
||||
if spec_config_json:
|
||||
try:
|
||||
config = json.loads(spec_config_json)
|
||||
logging.info(f"Using speculative config from SPECULATIVE_CONFIG: {config}")
|
||||
return config
|
||||
except json.JSONDecodeError as e:
|
||||
logging.error(f"Failed to parse SPECULATIVE_CONFIG JSON: {e}")
|
||||
return None
|
||||
|
||||
# Option 2: Build config from individual environment variables
|
||||
spec_method = os.getenv('SPECULATIVE_METHOD')
|
||||
spec_model = os.getenv('SPECULATIVE_MODEL')
|
||||
_num_spec_tokens = os.getenv('NUM_SPECULATIVE_TOKENS')
|
||||
_ngram_max = os.getenv('NGRAM_PROMPT_LOOKUP_MAX')
|
||||
_ngram_min = os.getenv('NGRAM_PROMPT_LOOKUP_MIN')
|
||||
|
||||
# Convert numeric vars to int so '0' (hub.json default) is treated as unset
|
||||
num_spec_tokens = (int(_num_spec_tokens) or None) if _num_spec_tokens else None
|
||||
ngram_max = (int(_ngram_max) or None) if _ngram_max else None
|
||||
ngram_min = (int(_ngram_min) or None) if _ngram_min else None
|
||||
|
||||
if not any([spec_method, spec_model, ngram_max]):
|
||||
return None
|
||||
|
||||
config = {}
|
||||
|
||||
# Determine method
|
||||
if spec_method:
|
||||
config['method'] = spec_method
|
||||
elif ngram_max and not spec_model:
|
||||
config['method'] = 'ngram'
|
||||
elif spec_model:
|
||||
model_lower = spec_model.lower()
|
||||
if 'eagle3' in model_lower:
|
||||
config['method'] = 'eagle3'
|
||||
elif 'eagle' in model_lower:
|
||||
config['method'] = 'eagle'
|
||||
elif 'medusa' in model_lower:
|
||||
config['method'] = 'medusa'
|
||||
else:
|
||||
config['method'] = 'draft_model'
|
||||
|
||||
if spec_model:
|
||||
config['model'] = spec_model
|
||||
if num_spec_tokens:
|
||||
config['num_speculative_tokens'] = num_spec_tokens
|
||||
if ngram_max:
|
||||
config['prompt_lookup_max'] = ngram_max
|
||||
if ngram_min:
|
||||
config['prompt_lookup_min'] = ngram_min
|
||||
|
||||
draft_tp = os.getenv('SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE')
|
||||
if draft_tp:
|
||||
config['draft_tensor_parallel_size'] = int(draft_tp)
|
||||
|
||||
spec_max_len = os.getenv('SPECULATIVE_MAX_MODEL_LEN')
|
||||
if spec_max_len:
|
||||
config['max_model_len'] = int(spec_max_len)
|
||||
|
||||
disable_batch = os.getenv('SPECULATIVE_DISABLE_BY_BATCH_SIZE')
|
||||
if disable_batch:
|
||||
config['disable_by_batch_size'] = int(disable_batch)
|
||||
|
||||
spec_quant = os.getenv('SPECULATIVE_QUANTIZATION')
|
||||
if spec_quant:
|
||||
config['quantization'] = spec_quant
|
||||
|
||||
spec_revision = os.getenv('SPECULATIVE_MODEL_REVISION')
|
||||
if spec_revision:
|
||||
config['revision'] = spec_revision
|
||||
|
||||
spec_eager = os.getenv('SPECULATIVE_ENFORCE_EAGER')
|
||||
if spec_eager:
|
||||
config['enforce_eager'] = spec_eager.lower() == 'true'
|
||||
|
||||
if config:
|
||||
logging.info(f"Built speculative config from env vars: {config}")
|
||||
return config
|
||||
|
||||
return None
|
||||
|
||||
|
||||
def _resolve_max_model_len(model, trust_remote_code=False, revision=None):
|
||||
"""Resolve max_model_len from the model's HuggingFace config."""
|
||||
try:
|
||||
from transformers import AutoConfig
|
||||
config = AutoConfig.from_pretrained(
|
||||
model,
|
||||
trust_remote_code=trust_remote_code,
|
||||
revision=revision,
|
||||
)
|
||||
for attr in ('max_position_embeddings', 'n_positions', 'max_seq_len', 'seq_length'):
|
||||
val = getattr(config, attr, None)
|
||||
if val is not None:
|
||||
logging.info(f"Resolved max_model_len={val} from model config ({attr})")
|
||||
return val
|
||||
except Exception as e:
|
||||
logging.warning(f"Could not resolve max_model_len from model config: {e}")
|
||||
return None
|
||||
|
||||
|
||||
def _local_args_to_engine_args(local: dict) -> dict:
|
||||
"""Map local args (e.g. from /local_model_args.json) to engine arg names and filter."""
|
||||
valid = AsyncEngineArgs.__dataclass_fields__
|
||||
out = {}
|
||||
for k, v in local.items():
|
||||
target = ENV_ALIASES.get(k, k.lower().replace("-", "_"))
|
||||
if target not in valid or v in (None, "", "None"):
|
||||
continue
|
||||
out[target] = v
|
||||
return out
|
||||
|
||||
|
||||
def _sanitize_hf_overrides(hf_overrides: dict) -> dict | None:
|
||||
"""Strip rope_scaling from hf_overrides sub-configs if vLLM rejects them.
|
||||
|
||||
Older vLLM (<0.7) required explicit mrope rope_scaling in hf_overrides for
|
||||
models like Qwen2-VL. Newer vLLM auto-detects mrope and raises a ValueError
|
||||
in patch_rope_scaling_dict when it finds conflicting rope_type values. Strip
|
||||
the offending rope_scaling so the model loads with its native config.
|
||||
"""
|
||||
if not isinstance(hf_overrides, dict):
|
||||
return hf_overrides
|
||||
|
||||
try:
|
||||
from vllm.transformers_utils.config import patch_rope_scaling_dict
|
||||
except ImportError:
|
||||
return hf_overrides
|
||||
|
||||
import copy
|
||||
cleaned = {}
|
||||
changed = False
|
||||
for key, value in hf_overrides.items():
|
||||
if isinstance(value, dict) and "rope_scaling" in value:
|
||||
rope_scaling = value.get("rope_scaling")
|
||||
if isinstance(rope_scaling, dict):
|
||||
try:
|
||||
patch_rope_scaling_dict(copy.deepcopy(rope_scaling))
|
||||
except (ValueError, Exception) as e:
|
||||
logging.warning(
|
||||
"Stripping hf_overrides['%s']['rope_scaling'] because vLLM "
|
||||
"rejected it (%s). Newer vLLM auto-detects rope scaling from "
|
||||
"the model config.", key, e
|
||||
)
|
||||
stripped = {k: v for k, v in value.items() if k != "rope_scaling"}
|
||||
cleaned[key] = stripped if stripped else None
|
||||
changed = True
|
||||
continue
|
||||
cleaned[key] = value
|
||||
|
||||
if not changed:
|
||||
return hf_overrides
|
||||
|
||||
result = {k: v for k, v in cleaned.items() if v is not None}
|
||||
return result or None
|
||||
|
||||
|
||||
def get_local_args():
|
||||
"""
|
||||
Retrieve local arguments from a JSON file.
|
||||
@@ -134,23 +352,43 @@ def get_local_args():
|
||||
|
||||
return local_args
|
||||
def get_engine_args():
|
||||
# Start with default args
|
||||
args = DEFAULT_ARGS
|
||||
# Start with worker custom defaults (only where we differ from vLLM)
|
||||
args = dict(DEFAULT_ARGS)
|
||||
|
||||
# Get env args that match keys in AsyncEngineArgs
|
||||
args.update(os.environ)
|
||||
# Auto-discover: every AsyncEngineArgs field from env UPPERCASED (e.g. MAX_MODEL_LEN)
|
||||
args.update(_get_args_from_env_auto_discover())
|
||||
|
||||
# Get local args if model is baked in and overwrite env args
|
||||
args.update(get_local_args())
|
||||
# Backward-compat aliases (MODEL_NAME → model, etc.)
|
||||
_apply_env_aliases(args)
|
||||
|
||||
# Local baked-in model overrides
|
||||
local = get_local_args()
|
||||
if local:
|
||||
args.update(_local_args_to_engine_args(local))
|
||||
|
||||
# Filter to valid engine args and drop sentinel empty values
|
||||
valid_fields = AsyncEngineArgs.__dataclass_fields__
|
||||
args = {
|
||||
k: v for k, v in args.items()
|
||||
if k in valid_fields and v not in (None, "", "None")
|
||||
}
|
||||
|
||||
# Special conversion for limit_mm_per_prompt (e.g. "image=1,video=0")
|
||||
limit_mm_env = os.getenv("LIMIT_MM_PER_PROMPT")
|
||||
if limit_mm_env is not None:
|
||||
args["limit_mm_per_prompt"] = convert_limit_mm_per_prompt(limit_mm_env)
|
||||
|
||||
# if args.get("TENSORIZER_URI"): TODO: add back once tensorizer is ready
|
||||
# args["load_format"] = "tensorizer"
|
||||
# args["model_loader_extra_config"] = TensorizerConfig(tensorizer_uri=args["TENSORIZER_URI"], num_readers=None)
|
||||
# logging.info(f"Using tensorized model from {args['TENSORIZER_URI']}")
|
||||
|
||||
|
||||
# Rename and match to vllm args
|
||||
args = match_vllm_args(args)
|
||||
if "hf_overrides" in args:
|
||||
sanitized = _sanitize_hf_overrides(args["hf_overrides"])
|
||||
if sanitized:
|
||||
args["hf_overrides"] = sanitized
|
||||
else:
|
||||
del args["hf_overrides"]
|
||||
|
||||
if args.get("load_format") == "bitsandbytes":
|
||||
args["quantization"] = args["load_format"]
|
||||
@@ -175,4 +413,49 @@ def get_engine_args():
|
||||
# os.environ["VLLM_ATTENTION_BACKEND"] = "FLASHINFER"
|
||||
# logging.info("Using FLASHINFER for gemma-2 model.")
|
||||
|
||||
# Set max_num_batched_tokens to max_model_len for unlimited batching.
|
||||
# vLLM defaults max_num_batched_tokens to 2048 when None, which is too low.
|
||||
|
||||
if args.get("max_model_len") == 0:
|
||||
args["max_model_len"] = None
|
||||
|
||||
if args.get("max_num_batched_tokens") == 0:
|
||||
args["max_num_batched_tokens"] = None
|
||||
|
||||
if args.get("max_num_batched_tokens") is None:
|
||||
max_model_len = args.get("max_model_len")
|
||||
if max_model_len is None:
|
||||
max_model_len = _resolve_max_model_len(
|
||||
args.get("model"),
|
||||
trust_remote_code=args.get("trust_remote_code", False),
|
||||
revision=args.get("revision"),
|
||||
)
|
||||
if max_model_len is not None:
|
||||
args["max_num_batched_tokens"] = max_model_len
|
||||
logging.info(f"Setting max_num_batched_tokens to {max_model_len}")
|
||||
|
||||
# VLLM_ATTENTION_BACKEND is deprecated, migrate to attention_backend
|
||||
if os.getenv('VLLM_ATTENTION_BACKEND'):
|
||||
logging.warning(
|
||||
"VLLM_ATTENTION_BACKEND env var is deprecated. "
|
||||
"Use ATTENTION_BACKEND instead (maps to --attention-backend CLI arg)."
|
||||
)
|
||||
if not args.get('attention_backend'):
|
||||
args['attention_backend'] = os.getenv('VLLM_ATTENTION_BACKEND')
|
||||
|
||||
# DISABLE_LOG_REQUESTS is deprecated, use ENABLE_LOG_REQUESTS instead
|
||||
if os.getenv('DISABLE_LOG_REQUESTS'):
|
||||
logging.warning(
|
||||
"DISABLE_LOG_REQUESTS env var is deprecated. "
|
||||
"Use ENABLE_LOG_REQUESTS instead (default: False)."
|
||||
)
|
||||
# Honor old behavior: if DISABLE_LOG_REQUESTS=true, don't enable logging
|
||||
if os.getenv('DISABLE_LOG_REQUESTS', 'False').lower() == 'true':
|
||||
args['enable_log_requests'] = False
|
||||
|
||||
# Add speculative decoding configuration if present
|
||||
speculative_config = get_speculative_config()
|
||||
if speculative_config:
|
||||
args["speculative_config"] = speculative_config
|
||||
|
||||
return AsyncEngineArgs(**args)
|
||||
|
||||
+42
-9
@@ -1,22 +1,55 @@
|
||||
import os
|
||||
import sys
|
||||
import multiprocessing
|
||||
import traceback
|
||||
import runpod
|
||||
from utils import JobInput
|
||||
from engine import vLLMEngine, OpenAIvLLMEngine
|
||||
from runpod import RunPodLogger
|
||||
|
||||
log = RunPodLogger()
|
||||
|
||||
vllm_engine = None
|
||||
openai_engine = None
|
||||
|
||||
vllm_engine = vLLMEngine()
|
||||
OpenAIvLLMEngine = OpenAIvLLMEngine(vllm_engine)
|
||||
|
||||
async def handler(job):
|
||||
try:
|
||||
from utils import JobInput
|
||||
job_input = JobInput(job["input"])
|
||||
engine = OpenAIvLLMEngine if job_input.openai_route else vllm_engine
|
||||
engine = openai_engine if job_input.openai_route else vllm_engine
|
||||
results_generator = engine.generate(job_input)
|
||||
async for batch in results_generator:
|
||||
yield batch
|
||||
except Exception as e:
|
||||
error_str = str(e)
|
||||
full_traceback = traceback.format_exc()
|
||||
|
||||
runpod.serverless.start(
|
||||
log.error(f"Error during inference: {error_str}")
|
||||
log.error(f"Full traceback:\n{full_traceback}")
|
||||
|
||||
# CUDA errors = worker is broken, exit to let RunPod spin up a healthy one
|
||||
if "CUDA" in error_str or "cuda" in error_str:
|
||||
log.error("Terminating worker due to CUDA/GPU error")
|
||||
sys.exit(1)
|
||||
|
||||
yield {"error": error_str}
|
||||
|
||||
|
||||
# Only run in main process to prevent re-initialization when vLLM spawns worker subprocesses
|
||||
if __name__ == "__main__" or multiprocessing.current_process().name == "MainProcess":
|
||||
|
||||
try:
|
||||
from engine import vLLMEngine, OpenAIvLLMEngine
|
||||
|
||||
vllm_engine = vLLMEngine()
|
||||
openai_engine = OpenAIvLLMEngine(vllm_engine)
|
||||
log.info("vLLM engines initialized successfully")
|
||||
except Exception as e:
|
||||
log.error(f"Worker startup failed: {e}\n{traceback.format_exc()}")
|
||||
sys.exit(1)
|
||||
|
||||
runpod.serverless.start(
|
||||
{
|
||||
"handler": handler,
|
||||
"concurrency_modifier": lambda x: vllm_engine.max_concurrency,
|
||||
"concurrency_modifier": lambda x: vllm_engine.max_concurrency if vllm_engine else 1,
|
||||
"return_aggregate_stream": True,
|
||||
}
|
||||
)
|
||||
)
|
||||
|
||||
+3
-4
@@ -3,11 +3,10 @@ import logging
|
||||
from http import HTTPStatus
|
||||
from functools import wraps
|
||||
from time import time
|
||||
from vllm.entrypoints.openai.protocol import RequestResponseMetadata
|
||||
|
||||
try:
|
||||
from vllm.utils import random_uuid
|
||||
from vllm.entrypoints.openai.protocol import ErrorResponse
|
||||
from vllm.entrypoints.openai.engine.protocol import ErrorResponse, ErrorInfo, RequestResponseMetadata
|
||||
from vllm import SamplingParams
|
||||
except ImportError:
|
||||
logging.warning("Error importing vllm, skipping related imports. This is ONLY expected when baking model into docker image from a machine without GPUs")
|
||||
@@ -88,9 +87,9 @@ class BatchSize:
|
||||
self.current_batch_size = min(self.current_batch_size*self.batch_size_growth_factor, self.max_batch_size)
|
||||
|
||||
def create_error_response(message: str, err_type: str = "BadRequestError", status_code: HTTPStatus = HTTPStatus.BAD_REQUEST) -> ErrorResponse:
|
||||
return ErrorResponse(message=message,
|
||||
return ErrorResponse(error=ErrorInfo(message=message,
|
||||
type=err_type,
|
||||
code=status_code.value)
|
||||
code=status_code.value))
|
||||
|
||||
def get_int_bool_env(env_var: str, default: bool) -> bool:
|
||||
return int(os.getenv(env_var, int(default))) == 1
|
||||
|
||||
-1514
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user