Compare commits

...
2 Commits
Author SHA1 Message Date
Tim Pietrusky ecd562e112 fix: remove "access token" as this is handled by the platform
Release / release (push) Waiting to run
2025-09-19 21:00:47 +02:00
5cffaab8e8 docs: how to use the reasoning parser (#218)
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-17 07:49:37 +02:00
3 changed files with 280 additions and 271 deletions
+261 -260
View File
@@ -1,260 +1,261 @@
![vLLM worker banner](https://cpjrphpz3t5wbwfe.public.blob.vercel-storage.com/worker-vllm_banner.jpeg) ![vLLM worker banner](https://cpjrphpz3t5wbwfe.public.blob.vercel-storage.com/worker-vllm_banner.jpeg)
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
--- ---
[![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm) [![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
--- ---
## Endpoint Configuration ## Endpoint Configuration
All behaviour is controlled through environment variables: All behaviour is controlled through environment variables:
| Environment Variable | Description | Default | Options | | Environment Variable | Description | Default | Options |
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ | | ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID | | `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token | | `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) | | `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" | | `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer | | `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 | | `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer | | `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string | | `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) | | `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. | | `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String | | `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer | | `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
## API Usage
## API Usage
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
### RunPod Native API
### RunPod Native API
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
#### Chat Completions
#### Chat Completions
```json
{ ```json
"input": { {
"messages": [ "input": {
{ "role": "system", "content": "You are a helpful assistant." }, "messages": [
{ "role": "user", "content": "What is the capital of France?" } { "role": "system", "content": "You are a helpful assistant." },
], { "role": "user", "content": "What is the capital of France?" }
"sampling_params": { ],
"max_tokens": 100, "sampling_params": {
"temperature": 0.7 "max_tokens": 100,
} "temperature": 0.7
} }
} }
``` }
```
#### Chat Completions (Streaming)
#### Chat Completions (Streaming)
```json
{ ```json
"input": { {
"messages": [ "input": {
{ "role": "user", "content": "Write a short story about a robot." } "messages": [
], { "role": "user", "content": "Write a short story about a robot." }
"sampling_params": { ],
"max_tokens": 500, "sampling_params": {
"temperature": 0.8 "max_tokens": 500,
}, "temperature": 0.8
"stream": true },
} "stream": true
} }
``` }
```
#### Text Generation
#### Text Generation
For direct text generation without chat format:
For direct text generation without chat format:
```json
{ ```json
"input": { {
"prompt": "The capital of France is", "input": {
"sampling_params": { "prompt": "The capital of France is",
"max_tokens": 64, "sampling_params": {
"temperature": 0.0 "max_tokens": 64,
} "temperature": 0.0
} }
} }
``` }
```
#### List Models
#### List Models
```json
{ ```json
"input": { {
"openai_route": "/v1/models" "input": {
} "openai_route": "/v1/models"
} }
``` }
```
---
---
### OpenAI-Compatible API
### OpenAI-Compatible API
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
#### Chat Completions
#### Chat Completions
**Path:** `/openai/v1/chat/completions`
**Path:** `/openai/v1/chat/completions`
```json
{ ```json
"model": "meta-llama/Llama-2-7b-chat-hf", {
"messages": [ "model": "meta-llama/Llama-2-7b-chat-hf",
{ "role": "system", "content": "You are a helpful assistant." }, "messages": [
{ "role": "user", "content": "What is the capital of France?" } { "role": "system", "content": "You are a helpful assistant." },
], { "role": "user", "content": "What is the capital of France?" }
"max_tokens": 100, ],
"temperature": 0.7 "max_tokens": 100,
} "temperature": 0.7
``` }
```
#### Chat Completions (Streaming)
#### Chat Completions (Streaming)
```json
{ ```json
"model": "meta-llama/Llama-2-7b-chat-hf", {
"messages": [ "model": "meta-llama/Llama-2-7b-chat-hf",
{ "role": "user", "content": "Write a short story about a robot." } "messages": [
], { "role": "user", "content": "Write a short story about a robot." }
"max_tokens": 500, ],
"temperature": 0.8, "max_tokens": 500,
"stream": true "temperature": 0.8,
} "stream": true
``` }
```
#### Text Completions
#### Text Completions
**Path:** `/openai/v1/completions`
**Path:** `/openai/v1/completions`
```json
{ ```json
"model": "meta-llama/Llama-2-7b-chat-hf", {
"prompt": "The capital of France is", "model": "meta-llama/Llama-2-7b-chat-hf",
"max_tokens": 100, "prompt": "The capital of France is",
"temperature": 0.7 "max_tokens": 100,
} "temperature": 0.7
``` }
```
#### List Models
#### List Models
**Path:** `/openai/v1/models`
**Path:** `/openai/v1/models`
```json
{} ```json
``` {}
```
#### Response Format
#### Response Format
Both APIs return the same response format:
Both APIs return the same response format:
```json
{ ```json
"choices": [ {
{ "choices": [
"index": 0, {
"message": { "role": "assistant", "content": "Paris." }, "index": 0,
"finish_reason": "stop" "message": { "role": "assistant", "content": "Paris." },
} "finish_reason": "stop"
], }
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 } ],
} "usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
``` }
```
---
---
## Usage
## Usage
Below are minimal `python` snippets so you can copy-paste to get started quickly.
Below are minimal `python` snippets so you can copy-paste to get started quickly.
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
### OpenAI compatible API
### OpenAI compatible API
Minimal Python example using the official `openai` SDK:
Minimal Python example using the official `openai` SDK:
```python
from openai import OpenAI ```python
import os from openai import OpenAI
import os
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
client = OpenAI( # Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
api_key=os.getenv("RUNPOD_API_KEY"), client = OpenAI(
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1", api_key=os.getenv("RUNPOD_API_KEY"),
) base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
``` )
```
`Chat Completions (Non-Streaming)`
`Chat Completions (Non-Streaming)`
```python
response = client.chat.completions.create( ```python
model="meta-llama/Llama-2-7b-chat-hf", response = client.chat.completions.create(
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}], model="meta-llama/Llama-2-7b-chat-hf",
temperature=0, messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
max_tokens=100, temperature=0,
) max_tokens=100,
print(f"Response: {response.choices[0].message.content}") )
``` print(f"Response: {response.choices[0].message.content}")
```
`Chat Completions (Streaming)`
`Chat Completions (Streaming)`
```python
response_stream = client.chat.completions.create( ```python
model="meta-llama/Llama-2-7b-chat-hf", response_stream = client.chat.completions.create(
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}], model="meta-llama/Llama-2-7b-chat-hf",
temperature=0, messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
max_tokens=100, temperature=0,
stream=True max_tokens=100,
) stream=True
for response in response_stream: )
print(response.choices[0].delta.content or "", end="", flush=True) for response in response_stream:
``` print(response.choices[0].delta.content or "", end="", flush=True)
```
### RunPod Native API
### RunPod Native API
```python
import requests ```python
import requests
response = requests.post(
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run", response = requests.post(
headers={"Authorization": "Bearer <API_KEY>"}, "https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
json={ headers={"Authorization": "Bearer <API_KEY>"},
"input": { json={
"messages": [ "input": {
{"role": "system", "content": "You are a helpful assistant."}, "messages": [
{"role": "user", "content": "Explain quantum computing in simple terms"} {"role": "system", "content": "You are a helpful assistant."},
], {"role": "user", "content": "Explain quantum computing in simple terms"}
"sampling_params": { ],
"temperature": 0.7, "sampling_params": {
"max_tokens": 150 "temperature": 0.7,
} "max_tokens": 150
} }
} }
) }
)
result = response.json()
print(result["output"]) result = response.json()
``` print(result["output"])
```
## Compatibility
## Compatibility
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
## Documentation
## Documentation
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables - **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies - **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns - **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
+18 -11
View File
@@ -1,6 +1,6 @@
{ {
"title": "vLLM", "title": "vLLM",
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the vLLM Inference Engine on RunPod Serverless", "description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by vLLM",
"type": "serverless", "type": "serverless",
"category": "language", "category": "language",
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png", "iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
@@ -38,16 +38,6 @@
"required": true "required": true
} }
}, },
{
"key": "HF_TOKEN",
"input": {
"name": "Access Token",
"type": "string",
"description": "Hugging Face access token for gated & private models",
"default": "",
"required": false
}
},
{ {
"key": "TOKENIZER", "key": "TOKENIZER",
"input": { "input": {
@@ -1023,6 +1013,23 @@
"default": "", "default": "",
"advanced": true "advanced": true
} }
},
{
"key": "REASONING_PARSER",
"input": {
"name": "Reasoning Parser",
"type": "string",
"description": "Parser for reasoning-capable models (enables reasoning mode)",
"options": [
{ "label": "None", "value": "" },
{ "label": "DeepSeek R1", "value": "deepseek_r1" },
{ "label": "Qwen3", "value": "qwen3" },
{ "label": "Granite", "value": "granite" },
{ "label": "Hunyuan A13B", "value": "hunyuan_a13b" }
],
"default": "",
"advanced": true
}
} }
] ]
} }
+1
View File
@@ -113,6 +113,7 @@ The way this works is that the first request will have a batch size of `DEFAULT_
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. | | `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. | | `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` | | `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
## Serverless & Concurrency Settings ## Serverless & Concurrency Settings