Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ecd562e112 | ||
|
|
5cffaab8e8 |
@@ -24,6 +24,7 @@ All behaviour is controlled through environment variables:
|
|||||||
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||||
|
| `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
|
||||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
||||||
|
|
||||||
|
|||||||
+18
-11
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"title": "vLLM",
|
"title": "vLLM",
|
||||||
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the vLLM Inference Engine on RunPod Serverless",
|
"description": "Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by vLLM",
|
||||||
"type": "serverless",
|
"type": "serverless",
|
||||||
"category": "language",
|
"category": "language",
|
||||||
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
|
"iconUrl": "https://registry.npmmirror.com/@lobehub/icons-static-png/latest/files/dark/vllm-color.png",
|
||||||
@@ -38,16 +38,6 @@
|
|||||||
"required": true
|
"required": true
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"key": "HF_TOKEN",
|
|
||||||
"input": {
|
|
||||||
"name": "Access Token",
|
|
||||||
"type": "string",
|
|
||||||
"description": "Hugging Face access token for gated & private models",
|
|
||||||
"default": "",
|
|
||||||
"required": false
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"key": "TOKENIZER",
|
"key": "TOKENIZER",
|
||||||
"input": {
|
"input": {
|
||||||
@@ -1023,6 +1013,23 @@
|
|||||||
"default": "",
|
"default": "",
|
||||||
"advanced": true
|
"advanced": true
|
||||||
}
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "REASONING_PARSER",
|
||||||
|
"input": {
|
||||||
|
"name": "Reasoning Parser",
|
||||||
|
"type": "string",
|
||||||
|
"description": "Parser for reasoning-capable models (enables reasoning mode)",
|
||||||
|
"options": [
|
||||||
|
{ "label": "None", "value": "" },
|
||||||
|
{ "label": "DeepSeek R1", "value": "deepseek_r1" },
|
||||||
|
{ "label": "Qwen3", "value": "qwen3" },
|
||||||
|
{ "label": "Granite", "value": "granite" },
|
||||||
|
{ "label": "Hunyuan A13B", "value": "hunyuan_a13b" }
|
||||||
|
],
|
||||||
|
"default": "",
|
||||||
|
"advanced": true
|
||||||
|
}
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -113,6 +113,7 @@ The way this works is that the first request will have a batch size of `DEFAULT_
|
|||||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
||||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||||
|
| `REASONING_PARSER` | `None` | `str` | Parser for reasoning-capable models (enables reasoning mode). Examples: `deepseek_r1`, `qwen3`, `granite`, `hunyuan_a13b`. Leave unset to disable. |
|
||||||
|
|
||||||
## Serverless & Concurrency Settings
|
## Serverless & Concurrency Settings
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user