Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c896438f21 | ||
|
|
912892f94e | ||
|
|
f8bf82469c | ||
|
|
ec1664902b | ||
|
|
d09122de4a | ||
|
|
e27dc68dea | ||
|
|
1ee18d06a9 | ||
|
|
5c4edd15cc | ||
|
|
b074d3a23b | ||
|
|
205847471c | ||
|
|
6337a6673a | ||
|
|
66e1b1605b | ||
|
|
60c8f257a8 | ||
|
|
fae16e7ee1 | ||
|
|
2becd35345 | ||
|
|
33d88df6c0 | ||
|
|
ecd562e112 | ||
|
|
5cffaab8e8 |
+261
-260
@@ -1,260 +1,261 @@
|
|||||||

|

|
||||||
|
|
||||||
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
|
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
[](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
|
[](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Endpoint Configuration
|
## Endpoint Configuration
|
||||||
|
|
||||||
All behaviour is controlled through environment variables:
|
All behaviour is controlled through environment variables:
|
||||||
|
|
||||||
| Environment Variable | Description | Default | Options |
|
| Environment Variable | Description | Default | Options |
|
||||||
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
||||||
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
||||||
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
||||||
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
||||||
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
||||||
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
||||||
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
||||||
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
||||||
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
| `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
|
||||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||||
|
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
||||||
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
|
|
||||||
|
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
|
||||||
## API Usage
|
|
||||||
|
## API Usage
|
||||||
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
|
|
||||||
|
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
|
||||||
### RunPod Native API
|
|
||||||
|
### RunPod Native API
|
||||||
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
|
|
||||||
|
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
|
||||||
#### Chat Completions
|
|
||||||
|
#### Chat Completions
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"input": {
|
{
|
||||||
"messages": [
|
"input": {
|
||||||
{ "role": "system", "content": "You are a helpful assistant." },
|
"messages": [
|
||||||
{ "role": "user", "content": "What is the capital of France?" }
|
{ "role": "system", "content": "You are a helpful assistant." },
|
||||||
],
|
{ "role": "user", "content": "What is the capital of France?" }
|
||||||
"sampling_params": {
|
],
|
||||||
"max_tokens": 100,
|
"sampling_params": {
|
||||||
"temperature": 0.7
|
"max_tokens": 100,
|
||||||
}
|
"temperature": 0.7
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
```
|
}
|
||||||
|
```
|
||||||
#### Chat Completions (Streaming)
|
|
||||||
|
#### Chat Completions (Streaming)
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"input": {
|
{
|
||||||
"messages": [
|
"input": {
|
||||||
{ "role": "user", "content": "Write a short story about a robot." }
|
"messages": [
|
||||||
],
|
{ "role": "user", "content": "Write a short story about a robot." }
|
||||||
"sampling_params": {
|
],
|
||||||
"max_tokens": 500,
|
"sampling_params": {
|
||||||
"temperature": 0.8
|
"max_tokens": 500,
|
||||||
},
|
"temperature": 0.8
|
||||||
"stream": true
|
},
|
||||||
}
|
"stream": true
|
||||||
}
|
}
|
||||||
```
|
}
|
||||||
|
```
|
||||||
#### Text Generation
|
|
||||||
|
#### Text Generation
|
||||||
For direct text generation without chat format:
|
|
||||||
|
For direct text generation without chat format:
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"input": {
|
{
|
||||||
"prompt": "The capital of France is",
|
"input": {
|
||||||
"sampling_params": {
|
"prompt": "The capital of France is",
|
||||||
"max_tokens": 64,
|
"sampling_params": {
|
||||||
"temperature": 0.0
|
"max_tokens": 64,
|
||||||
}
|
"temperature": 0.0
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
```
|
}
|
||||||
|
```
|
||||||
#### List Models
|
|
||||||
|
#### List Models
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"input": {
|
{
|
||||||
"openai_route": "/v1/models"
|
"input": {
|
||||||
}
|
"openai_route": "/v1/models"
|
||||||
}
|
}
|
||||||
```
|
}
|
||||||
|
```
|
||||||
---
|
|
||||||
|
---
|
||||||
### OpenAI-Compatible API
|
|
||||||
|
### OpenAI-Compatible API
|
||||||
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
|
|
||||||
|
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
|
||||||
#### Chat Completions
|
|
||||||
|
#### Chat Completions
|
||||||
**Path:** `/openai/v1/chat/completions`
|
|
||||||
|
**Path:** `/openai/v1/chat/completions`
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
{
|
||||||
"messages": [
|
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||||
{ "role": "system", "content": "You are a helpful assistant." },
|
"messages": [
|
||||||
{ "role": "user", "content": "What is the capital of France?" }
|
{ "role": "system", "content": "You are a helpful assistant." },
|
||||||
],
|
{ "role": "user", "content": "What is the capital of France?" }
|
||||||
"max_tokens": 100,
|
],
|
||||||
"temperature": 0.7
|
"max_tokens": 100,
|
||||||
}
|
"temperature": 0.7
|
||||||
```
|
}
|
||||||
|
```
|
||||||
#### Chat Completions (Streaming)
|
|
||||||
|
#### Chat Completions (Streaming)
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
{
|
||||||
"messages": [
|
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||||
{ "role": "user", "content": "Write a short story about a robot." }
|
"messages": [
|
||||||
],
|
{ "role": "user", "content": "Write a short story about a robot." }
|
||||||
"max_tokens": 500,
|
],
|
||||||
"temperature": 0.8,
|
"max_tokens": 500,
|
||||||
"stream": true
|
"temperature": 0.8,
|
||||||
}
|
"stream": true
|
||||||
```
|
}
|
||||||
|
```
|
||||||
#### Text Completions
|
|
||||||
|
#### Text Completions
|
||||||
**Path:** `/openai/v1/completions`
|
|
||||||
|
**Path:** `/openai/v1/completions`
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
{
|
||||||
"prompt": "The capital of France is",
|
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||||
"max_tokens": 100,
|
"prompt": "The capital of France is",
|
||||||
"temperature": 0.7
|
"max_tokens": 100,
|
||||||
}
|
"temperature": 0.7
|
||||||
```
|
}
|
||||||
|
```
|
||||||
#### List Models
|
|
||||||
|
#### List Models
|
||||||
**Path:** `/openai/v1/models`
|
|
||||||
|
**Path:** `/openai/v1/models`
|
||||||
```json
|
|
||||||
{}
|
```json
|
||||||
```
|
{}
|
||||||
|
```
|
||||||
#### Response Format
|
|
||||||
|
#### Response Format
|
||||||
Both APIs return the same response format:
|
|
||||||
|
Both APIs return the same response format:
|
||||||
```json
|
|
||||||
{
|
```json
|
||||||
"choices": [
|
{
|
||||||
{
|
"choices": [
|
||||||
"index": 0,
|
{
|
||||||
"message": { "role": "assistant", "content": "Paris." },
|
"index": 0,
|
||||||
"finish_reason": "stop"
|
"message": { "role": "assistant", "content": "Paris." },
|
||||||
}
|
"finish_reason": "stop"
|
||||||
],
|
}
|
||||||
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
|
],
|
||||||
}
|
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
|
||||||
```
|
}
|
||||||
|
```
|
||||||
---
|
|
||||||
|
---
|
||||||
## Usage
|
|
||||||
|
## Usage
|
||||||
Below are minimal `python` snippets so you can copy-paste to get started quickly.
|
|
||||||
|
Below are minimal `python` snippets so you can copy-paste to get started quickly.
|
||||||
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
|
|
||||||
|
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
|
||||||
### OpenAI compatible API
|
|
||||||
|
### OpenAI compatible API
|
||||||
Minimal Python example using the official `openai` SDK:
|
|
||||||
|
Minimal Python example using the official `openai` SDK:
|
||||||
```python
|
|
||||||
from openai import OpenAI
|
```python
|
||||||
import os
|
from openai import OpenAI
|
||||||
|
import os
|
||||||
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
|
|
||||||
client = OpenAI(
|
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
|
||||||
api_key=os.getenv("RUNPOD_API_KEY"),
|
client = OpenAI(
|
||||||
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
|
api_key=os.getenv("RUNPOD_API_KEY"),
|
||||||
)
|
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
|
||||||
```
|
)
|
||||||
|
```
|
||||||
`Chat Completions (Non-Streaming)`
|
|
||||||
|
`Chat Completions (Non-Streaming)`
|
||||||
```python
|
|
||||||
response = client.chat.completions.create(
|
```python
|
||||||
model="meta-llama/Llama-2-7b-chat-hf",
|
response = client.chat.completions.create(
|
||||||
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
model="meta-llama/Llama-2-7b-chat-hf",
|
||||||
temperature=0,
|
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||||
max_tokens=100,
|
temperature=0,
|
||||||
)
|
max_tokens=100,
|
||||||
print(f"Response: {response.choices[0].message.content}")
|
)
|
||||||
```
|
print(f"Response: {response.choices[0].message.content}")
|
||||||
|
```
|
||||||
`Chat Completions (Streaming)`
|
|
||||||
|
`Chat Completions (Streaming)`
|
||||||
```python
|
|
||||||
response_stream = client.chat.completions.create(
|
```python
|
||||||
model="meta-llama/Llama-2-7b-chat-hf",
|
response_stream = client.chat.completions.create(
|
||||||
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
model="meta-llama/Llama-2-7b-chat-hf",
|
||||||
temperature=0,
|
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||||
max_tokens=100,
|
temperature=0,
|
||||||
stream=True
|
max_tokens=100,
|
||||||
)
|
stream=True
|
||||||
for response in response_stream:
|
)
|
||||||
print(response.choices[0].delta.content or "", end="", flush=True)
|
for response in response_stream:
|
||||||
```
|
print(response.choices[0].delta.content or "", end="", flush=True)
|
||||||
|
```
|
||||||
### RunPod Native API
|
|
||||||
|
### RunPod Native API
|
||||||
```python
|
|
||||||
import requests
|
```python
|
||||||
|
import requests
|
||||||
response = requests.post(
|
|
||||||
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
|
response = requests.post(
|
||||||
headers={"Authorization": "Bearer <API_KEY>"},
|
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
|
||||||
json={
|
headers={"Authorization": "Bearer <API_KEY>"},
|
||||||
"input": {
|
json={
|
||||||
"messages": [
|
"input": {
|
||||||
{"role": "system", "content": "You are a helpful assistant."},
|
"messages": [
|
||||||
{"role": "user", "content": "Explain quantum computing in simple terms"}
|
{"role": "system", "content": "You are a helpful assistant."},
|
||||||
],
|
{"role": "user", "content": "Explain quantum computing in simple terms"}
|
||||||
"sampling_params": {
|
],
|
||||||
"temperature": 0.7,
|
"sampling_params": {
|
||||||
"max_tokens": 150
|
"temperature": 0.7,
|
||||||
}
|
"max_tokens": 150
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
)
|
}
|
||||||
|
)
|
||||||
result = response.json()
|
|
||||||
print(result["output"])
|
result = response.json()
|
||||||
```
|
print(result["output"])
|
||||||
|
```
|
||||||
## Compatibility
|
|
||||||
|
## Compatibility
|
||||||
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
|
|
||||||
|
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
|
||||||
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
|
|
||||||
|
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
|
||||||
## Documentation
|
|
||||||
|
## Documentation
|
||||||
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
|
|
||||||
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
|
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
|
||||||
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
|
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
|
||||||
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
|
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
|
||||||
|
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
|
||||||
|
|||||||
+22
-25
@@ -1,25 +1,15 @@
|
|||||||
{
|
{
|
||||||
"title": "vLLM",
|
"title": "vLLM",
|
||||||
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the vLLM Inference Engine on RunPod Serverless",
|
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by vLLM",
|
||||||
"type": "serverless",
|
"type": "serverless",
|
||||||
"category": "language",
|
"category": "language",
|
||||||
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
|
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
|
||||||
"config": {
|
"config": {
|
||||||
"runsOn": "GPU",
|
"runsOn": "GPU",
|
||||||
"containerDiskInGb": 200,
|
"containerDiskInGb": 150,
|
||||||
"gpuIds": "ADA_80_PRO, AMPERE_80",
|
"gpuIds": "ADA_80_PRO,AMPERE_80",
|
||||||
"gpuCount": 1,
|
"gpuCount": 1,
|
||||||
"allowedCudaVersions": [
|
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5", "12.4"],
|
||||||
"12.9",
|
|
||||||
"12.8",
|
|
||||||
"12.7",
|
|
||||||
"12.6",
|
|
||||||
"12.5",
|
|
||||||
"12.4",
|
|
||||||
"12.3",
|
|
||||||
"12.2",
|
|
||||||
"12.1"
|
|
||||||
],
|
|
||||||
"presets": [
|
"presets": [
|
||||||
{
|
{
|
||||||
"name": "deepseek-ai/deepseek-r1-distill-llama-8b",
|
"name": "deepseek-ai/deepseek-r1-distill-llama-8b",
|
||||||
@@ -38,16 +28,6 @@
|
|||||||
"required": true
|
"required": true
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"key": "HF_TOKEN",
|
|
||||||
"input": {
|
|
||||||
"name": "Access Token",
|
|
||||||
"type": "string",
|
|
||||||
"description": "Hugging Face access token for gated & private models",
|
|
||||||
"default": "",
|
|
||||||
"required": false
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"key": "TOKENIZER",
|
"key": "TOKENIZER",
|
||||||
"input": {
|
"input": {
|
||||||
@@ -945,7 +925,7 @@
|
|||||||
"name": "Max Concurrency",
|
"name": "Max Concurrency",
|
||||||
"type": "number",
|
"type": "number",
|
||||||
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
|
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
|
||||||
"default": 300,
|
"default": 30,
|
||||||
"advanced": true
|
"advanced": true
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
@@ -1023,6 +1003,23 @@
|
|||||||
"default": "",
|
"default": "",
|
||||||
"advanced": true
|
"advanced": true
|
||||||
}
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "REASONING_PARSER",
|
||||||
|
"input": {
|
||||||
|
"name": "Reasoning Parser",
|
||||||
|
"type": "string",
|
||||||
|
"description": "Parser for reasoning-capable models (enables reasoning mode)",
|
||||||
|
"options": [
|
||||||
|
{ "label": "None", "value": "" },
|
||||||
|
{ "label": "DeepSeek R1", "value": "deepseek_r1" },
|
||||||
|
{ "label": "Qwen3", "value": "qwen3" },
|
||||||
|
{ "label": "Granite", "value": "granite" },
|
||||||
|
{ "label": "Hunyuan A13B", "value": "hunyuan_a13b" }
|
||||||
|
],
|
||||||
|
"default": "",
|
||||||
|
"advanced": true
|
||||||
|
}
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
|
|||||||
+1
-9
@@ -38,14 +38,6 @@
|
|||||||
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
|
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"allowedCudaVersions": [
|
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
|
||||||
"12.7",
|
|
||||||
"12.6",
|
|
||||||
"12.5",
|
|
||||||
"12.4",
|
|
||||||
"12.3",
|
|
||||||
"12.2",
|
|
||||||
"12.1"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
+1
-1
@@ -12,7 +12,7 @@ RUN --mount=type=cache,target=/root/.cache/pip \
|
|||||||
python3 -m pip install --upgrade -r /requirements.txt
|
python3 -m pip install --upgrade -r /requirements.txt
|
||||||
|
|
||||||
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
|
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
|
||||||
RUN python3 -m pip install vllm==0.10.0 && \
|
RUN python3 -m pip install vllm==0.11.0 && \
|
||||||
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
|
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
|
||||||
|
|
||||||
# Setup for Option 2: Building the Image with the Model included
|
# Setup for Option 2: Building the Image with the Model included
|
||||||
|
|||||||
@@ -57,7 +57,7 @@ Configure worker-vllm using environment variables:
|
|||||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
| `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
|
||||||
|
|
||||||
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
||||||
|
|
||||||
|
|||||||
@@ -8,7 +8,7 @@ typing-extensions>=4.8.0
|
|||||||
pydantic
|
pydantic
|
||||||
pydantic-settings
|
pydantic-settings
|
||||||
hf-transfer
|
hf-transfer
|
||||||
transformers>=4.55.0
|
transformers>=4.57.0
|
||||||
bitsandbytes>=0.45.0
|
bitsandbytes>=0.45.0
|
||||||
kernels
|
kernels
|
||||||
torch==2.6.0
|
torch==2.6.0
|
||||||
|
|||||||
@@ -113,12 +113,13 @@ The way this works is that the first request will have a batch size of `DEFAULT_
|
|||||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
||||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||||
|
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
|
||||||
|
|
||||||
## Serverless & Concurrency Settings
|
## Serverless & Concurrency Settings
|
||||||
|
|
||||||
| Variable | Default | Type/Choices | Description |
|
| Variable | Default | Type/Choices | Description |
|
||||||
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||||
|
|
||||||
|
|||||||
+3
-2
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
|
|||||||
|
|
||||||
- `src/engine_args.py`: Centralized configuration management
|
- `src/engine_args.py`: Centralized configuration management
|
||||||
- `src/constants.py`: Default values for core settings
|
- `src/constants.py`: Default values for core settings
|
||||||
- `worker-config.json`: UI form generation for RunPod console
|
- `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
|
||||||
|
- `worker-config.json`: UI form generation for RunPod console (if exists)
|
||||||
|
|
||||||
## Core Development Concepts
|
## Core Development Concepts
|
||||||
|
|
||||||
@@ -222,7 +223,7 @@ src/
|
|||||||
|
|
||||||
### 2. **Concurrency Patterns**
|
### 2. **Concurrency Patterns**
|
||||||
|
|
||||||
- **Max Concurrency**: 300 concurrent requests by default
|
- **Max Concurrency**: 30 concurrent requests by default
|
||||||
- **vLLM Queuing**: Internal request batching and scheduling
|
- **vLLM Queuing**: Internal request batching and scheduling
|
||||||
- **RunPod Integration**: Concurrency modifier for auto-scaling
|
- **RunPod Integration**: Concurrency modifier for auto-scaling
|
||||||
|
|
||||||
|
|||||||
+1
-1
@@ -1,4 +1,4 @@
|
|||||||
DEFAULT_BATCH_SIZE = 50
|
DEFAULT_BATCH_SIZE = 50
|
||||||
DEFAULT_MAX_CONCURRENCY = 300
|
DEFAULT_MAX_CONCURRENCY = 30
|
||||||
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
|
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
|
||||||
DEFAULT_MIN_BATCH_SIZE = 1
|
DEFAULT_MIN_BATCH_SIZE = 1
|
||||||
-1514
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user