feat: better hub support & concise README for the main repo (#215)
Release / release (push) Waiting to run
Release / release (push) Waiting to run
* feat: moved config into docs; added banner; auto detect "messages" in input * docs: moved config into docs * chore: added .DS_Store * chore: get the original stuff working again * chore: remove all changes * docs: reduced toc and added small config table --------- Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
This commit is contained in:
co-authored by
Tim Pietrusky
parent
0e0d6df859
commit
a0fe1dfdad
@@ -4,3 +4,4 @@ runpod.toml
|
|||||||
.env
|
.env
|
||||||
test/*
|
test/*
|
||||||
vllm-base/vllm-*
|
vllm-base/vllm-*
|
||||||
|
.DS_Store
|
||||||
@@ -0,0 +1,260 @@
|
|||||||
|

|
||||||
|
|
||||||
|
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
[](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Endpoint Configuration
|
||||||
|
|
||||||
|
All behaviour is controlled through environment variables:
|
||||||
|
|
||||||
|
| Environment Variable | Description | Default | Options |
|
||||||
|
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
||||||
|
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
||||||
|
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
||||||
|
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
||||||
|
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
||||||
|
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
||||||
|
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
||||||
|
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
||||||
|
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||||
|
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||||
|
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||||
|
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||||
|
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
||||||
|
|
||||||
|
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
|
||||||
|
|
||||||
|
## API Usage
|
||||||
|
|
||||||
|
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
|
||||||
|
|
||||||
|
### RunPod Native API
|
||||||
|
|
||||||
|
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
|
||||||
|
|
||||||
|
#### Chat Completions
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"input": {
|
||||||
|
"messages": [
|
||||||
|
{ "role": "system", "content": "You are a helpful assistant." },
|
||||||
|
{ "role": "user", "content": "What is the capital of France?" }
|
||||||
|
],
|
||||||
|
"sampling_params": {
|
||||||
|
"max_tokens": 100,
|
||||||
|
"temperature": 0.7
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Chat Completions (Streaming)
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"input": {
|
||||||
|
"messages": [
|
||||||
|
{ "role": "user", "content": "Write a short story about a robot." }
|
||||||
|
],
|
||||||
|
"sampling_params": {
|
||||||
|
"max_tokens": 500,
|
||||||
|
"temperature": 0.8
|
||||||
|
},
|
||||||
|
"stream": true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Text Generation
|
||||||
|
|
||||||
|
For direct text generation without chat format:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"input": {
|
||||||
|
"prompt": "The capital of France is",
|
||||||
|
"sampling_params": {
|
||||||
|
"max_tokens": 64,
|
||||||
|
"temperature": 0.0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### List Models
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"input": {
|
||||||
|
"openai_route": "/v1/models"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### OpenAI-Compatible API
|
||||||
|
|
||||||
|
For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod API key.
|
||||||
|
|
||||||
|
#### Chat Completions
|
||||||
|
|
||||||
|
**Path:** `/openai/v1/chat/completions`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||||
|
"messages": [
|
||||||
|
{ "role": "system", "content": "You are a helpful assistant." },
|
||||||
|
{ "role": "user", "content": "What is the capital of France?" }
|
||||||
|
],
|
||||||
|
"max_tokens": 100,
|
||||||
|
"temperature": 0.7
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Chat Completions (Streaming)
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||||
|
"messages": [
|
||||||
|
{ "role": "user", "content": "Write a short story about a robot." }
|
||||||
|
],
|
||||||
|
"max_tokens": 500,
|
||||||
|
"temperature": 0.8,
|
||||||
|
"stream": true
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Text Completions
|
||||||
|
|
||||||
|
**Path:** `/openai/v1/completions`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||||
|
"prompt": "The capital of France is",
|
||||||
|
"max_tokens": 100,
|
||||||
|
"temperature": 0.7
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### List Models
|
||||||
|
|
||||||
|
**Path:** `/openai/v1/models`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{}
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Response Format
|
||||||
|
|
||||||
|
Both APIs return the same response format:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"choices": [
|
||||||
|
{
|
||||||
|
"index": 0,
|
||||||
|
"message": { "role": "assistant", "content": "Paris." },
|
||||||
|
"finish_reason": "stop"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
Below are minimal `python` snippets so you can copy-paste to get started quickly.
|
||||||
|
|
||||||
|
> Replace `<ENDPOINT_ID>` with your endpoint ID and `<API_KEY>` with a [RunPod API key](https://docs.runpod.io/get-started/api-keys).
|
||||||
|
|
||||||
|
### OpenAI compatible API
|
||||||
|
|
||||||
|
Minimal Python example using the official `openai` SDK:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from openai import OpenAI
|
||||||
|
import os
|
||||||
|
|
||||||
|
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
|
||||||
|
client = OpenAI(
|
||||||
|
api_key=os.getenv("RUNPOD_API_KEY"),
|
||||||
|
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
`Chat Completions (Non-Streaming)`
|
||||||
|
|
||||||
|
```python
|
||||||
|
response = client.chat.completions.create(
|
||||||
|
model="meta-llama/Llama-2-7b-chat-hf",
|
||||||
|
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||||
|
temperature=0,
|
||||||
|
max_tokens=100,
|
||||||
|
)
|
||||||
|
print(f"Response: {response.choices[0].message.content}")
|
||||||
|
```
|
||||||
|
|
||||||
|
`Chat Completions (Streaming)`
|
||||||
|
|
||||||
|
```python
|
||||||
|
response_stream = client.chat.completions.create(
|
||||||
|
model="meta-llama/Llama-2-7b-chat-hf",
|
||||||
|
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
|
||||||
|
temperature=0,
|
||||||
|
max_tokens=100,
|
||||||
|
stream=True
|
||||||
|
)
|
||||||
|
for response in response_stream:
|
||||||
|
print(response.choices[0].delta.content or "", end="", flush=True)
|
||||||
|
```
|
||||||
|
|
||||||
|
### RunPod Native API
|
||||||
|
|
||||||
|
```python
|
||||||
|
import requests
|
||||||
|
|
||||||
|
response = requests.post(
|
||||||
|
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
|
||||||
|
headers={"Authorization": "Bearer <API_KEY>"},
|
||||||
|
json={
|
||||||
|
"input": {
|
||||||
|
"messages": [
|
||||||
|
{"role": "system", "content": "You are a helpful assistant."},
|
||||||
|
{"role": "user", "content": "Explain quantum computing in simple terms"}
|
||||||
|
],
|
||||||
|
"sampling_params": {
|
||||||
|
"temperature": 0.7,
|
||||||
|
"max_tokens": 150
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
result = response.json()
|
||||||
|
print(result["output"])
|
||||||
|
```
|
||||||
|
|
||||||
|
## Compatibility
|
||||||
|
|
||||||
|
For supported models, see the [vLLM supported models documentation](https://docs.vllm.ai/en/latest/models/supported_models.html).
|
||||||
|
|
||||||
|
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
|
||||||
|
|
||||||
|
## Documentation
|
||||||
|
|
||||||
|
- **[🚀 Deployment Guide](https://docs.runpod.io/serverless/vllm/get-started)** - Step-by-step setup
|
||||||
|
- **[📖 Configuration Reference](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md)** - All environment variables
|
||||||
|
- **[🏗️ Advanced Deployment](https://github.com/runpod-workers/worker-vllm/blob/main/docs/deployment.md)** - Custom builds and strategies
|
||||||
|
- **[🔧 Development Guide](https://github.com/runpod-workers/worker-vllm/blob/main/docs/conventions.md)** - Architecture and patterns
|
||||||
+2
-3
@@ -12,7 +12,6 @@
|
|||||||
"input": {
|
"input": {
|
||||||
"openai_route": "/v1/chat/completions",
|
"openai_route": "/v1/chat/completions",
|
||||||
"openai_input": {
|
"openai_input": {
|
||||||
"model": "HuggingFaceTB/SmolLM2-135M-Instruct",
|
|
||||||
"messages": [
|
"messages": [
|
||||||
{
|
{
|
||||||
"role": "system",
|
"role": "system",
|
||||||
@@ -23,8 +22,8 @@
|
|||||||
"content": "Explain what a neural network is in one sentence."
|
"content": "Explain what a neural network is in one sentence."
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"max_tokens": 50,
|
"max_tokens": 200,
|
||||||
"temperature": 0.7
|
"temperature": 0.1
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
"timeout": 30000
|
"timeout": 30000
|
||||||
|
|||||||
@@ -9,16 +9,10 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
|||||||
## Table of Contents
|
## Table of Contents
|
||||||
|
|
||||||
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
|
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
|
||||||
- [Option 1: Deploy Any Model Using Pre-Built Docker Image **[RECOMMENDED]**](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
- [Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended]](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
||||||
- [Environment Variables](#environment-variables)
|
- [Configuration](#configuration)
|
||||||
- [LLM Settings](#llm-settings)
|
|
||||||
- [Tokenizer Settings](#tokenizer-settings)
|
|
||||||
- [System and Parallelism Settings](#system-and-parallelism-settings)
|
|
||||||
- [Streaming Batch Size Settings](#streaming-batch-size-settings)
|
|
||||||
- [OpenAI Settings](#openai-settings)
|
|
||||||
- [Serverless Settings](#serverless-settings)
|
|
||||||
- [Option 2: Build Docker Image with Model Inside](#option-2-build-docker-image-with-model-inside)
|
- [Option 2: Build Docker Image with Model Inside](#option-2-build-docker-image-with-model-inside)
|
||||||
- [Prerequisites](#prerequisites-1)
|
- [Prerequisites](#prerequisites)
|
||||||
- [Arguments](#arguments)
|
- [Arguments](#arguments)
|
||||||
- [Example: Building an image with OpenChat-3.5](#example-building-an-image-with-openchat-35)
|
- [Example: Building an image with OpenChat-3.5](#example-building-an-image-with-openchat-35)
|
||||||
- [(Optional) Including Huggingface Token](#optional-including-huggingface-token)
|
- [(Optional) Including Huggingface Token](#optional-including-huggingface-token)
|
||||||
@@ -26,12 +20,14 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
|||||||
- [Usage: OpenAI Compatibility](#usage-openai-compatibility)
|
- [Usage: OpenAI Compatibility](#usage-openai-compatibility)
|
||||||
- [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker)
|
- [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker)
|
||||||
- [OpenAI Request Input Parameters](#openai-request-input-parameters)
|
- [OpenAI Request Input Parameters](#openai-request-input-parameters)
|
||||||
- [Chat Completions](#chat-completions)
|
- [Chat Completions [RECOMMENDED]](#chat-completions-recommended)
|
||||||
- [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai)
|
- [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai)
|
||||||
- [Usage: standard](#non-openai-usage)
|
- [Chat Completions](#chat-completions)
|
||||||
- [Input Request Parameters](#input-request-parameters)
|
- [Getting a list of names for available models](#getting-a-list-of-names-for-available-models)
|
||||||
- [Text Input Formats](#text-input-formats)
|
- [Usage: Standard (Non-OpenAI)](#usage-standard-non-openai)
|
||||||
|
- [Request Input Parameters](#request-input-parameters)
|
||||||
- [Sampling Parameters](#sampling-parameters)
|
- [Sampling Parameters](#sampling-parameters)
|
||||||
|
- [Text Input Formats](#text-input-formats)
|
||||||
|
|
||||||
# Setting up the Serverless Worker
|
# Setting up the Serverless Worker
|
||||||
|
|
||||||
@@ -44,124 +40,26 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
|||||||
- **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases)
|
- **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases)
|
||||||
- **CUDA Compatibility**: Requires CUDA >= 12.1
|
- **CUDA Compatibility**: Requires CUDA >= 12.1
|
||||||
|
|
||||||
### Environment Variables
|
### Configuration
|
||||||
|
|
||||||
Use these to configure worker-vllm so it works for your use case / model.
|
Configure worker-vllm using environment variables:
|
||||||
|
|
||||||
#### LLM Settings
|
| Environment Variable | Description | Default | Options |
|
||||||
|
| ----------------------------------- | ------------------------------------------------- | ------------------- | ------------------------------------------------------------------ |
|
||||||
|
| `MODEL_NAME` | Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
|
||||||
|
| `HF_TOKEN` | HuggingFace access token for gated/private models | | Your HuggingFace access token |
|
||||||
|
| `MAX_MODEL_LEN` | Model's maximum context length | | Integer (e.g., 4096) |
|
||||||
|
| `QUANTIZATION` | Quantization method | | "awq", "gptq", "squeezellm", "bitsandbytes" |
|
||||||
|
| `TENSOR_PARALLEL_SIZE` | Number of GPUs | 1 | Integer |
|
||||||
|
| `GPU_MEMORY_UTILIZATION` | Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
|
||||||
|
| `MAX_NUM_SEQS` | Maximum number of sequences per iteration | 256 | Integer |
|
||||||
|
| `CUSTOM_CHAT_TEMPLATE` | Custom chat template override | | Jinja2 template string |
|
||||||
|
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||||
|
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||||
|
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||||
|
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
||||||
|
|
||||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
||||||
| ------------------------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
||||||
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
|
|
||||||
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
|
|
||||||
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
|
|
||||||
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
|
|
||||||
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
|
|
||||||
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
|
|
||||||
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
|
|
||||||
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
|
|
||||||
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
|
|
||||||
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
|
|
||||||
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
|
|
||||||
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
|
|
||||||
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
|
|
||||||
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
|
|
||||||
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
|
|
||||||
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
|
|
||||||
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
|
|
||||||
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
|
|
||||||
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
|
|
||||||
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
|
|
||||||
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
|
|
||||||
| `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. |
|
|
||||||
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
|
|
||||||
| `SEED` | 0 | `int` | Random seed for operations. |
|
|
||||||
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
|
|
||||||
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
|
|
||||||
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
|
|
||||||
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
|
|
||||||
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
|
|
||||||
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
|
|
||||||
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
|
|
||||||
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
|
|
||||||
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
|
|
||||||
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
|
|
||||||
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
|
|
||||||
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
|
|
||||||
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
|
|
||||||
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
|
|
||||||
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
|
|
||||||
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
|
|
||||||
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
|
|
||||||
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
|
|
||||||
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
|
|
||||||
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}` |
|
|
||||||
| `SCHEDULER_DELAY_FACTOR` | 0.0 | `float` | Apply a delay before scheduling next prompt. |
|
|
||||||
| `ENABLE_CHUNKED_PREFILL` | False | `bool` | Enable chunked prefill requests. |
|
|
||||||
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
|
|
||||||
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
|
|
||||||
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
|
|
||||||
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
|
|
||||||
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
|
|
||||||
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
|
|
||||||
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
|
|
||||||
| `SPEC_DECODING_ACCEPTANCE_METHOD` | 'rejection_sampler' | ['rejection_sampler', 'typical_acceptance_sampler'] | Specify the acceptance method for draft token verification in speculative decoding. |
|
|
||||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD` | None | `float` | Set the lower bound threshold for the posterior probability of a token to be accepted. |
|
|
||||||
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA` | None | `float` | A scaling factor for the entropy-based threshold for token acceptance. |
|
|
||||||
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
|
|
||||||
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
|
|
||||||
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
|
|
||||||
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
|
|
||||||
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
|
|
||||||
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
|
|
||||||
|
|
||||||
#### Tokenizer Settings
|
|
||||||
|
|
||||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
|
||||||
| ---------------------- | --------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
|
|
||||||
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
|
|
||||||
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
|
|
||||||
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
|
|
||||||
|
|
||||||
#### System and Parallelism Settings
|
|
||||||
|
|
||||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
|
||||||
| ------------------------------ | --------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
|
|
||||||
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
|
|
||||||
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
|
|
||||||
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
|
|
||||||
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
|
||||||
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
|
|
||||||
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
|
|
||||||
|
|
||||||
#### Streaming Batch Size Settings
|
|
||||||
|
|
||||||
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker
|
|
||||||
|
|
||||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
|
||||||
| ---------------------------------- | --------- | -------------- | --------------------------------------------------------------------------------------------------------- |
|
|
||||||
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
|
|
||||||
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
|
|
||||||
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
|
|
||||||
|
|
||||||
#### OpenAI Settings
|
|
||||||
|
|
||||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
|
||||||
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
|
|
||||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
|
|
||||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
|
||||||
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
|
||||||
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
|
||||||
|
|
||||||
#### Serverless Settings
|
|
||||||
|
|
||||||
| `Name` | `Default` | `Type/Choices` | `Description` |
|
|
||||||
| ---------------------- | --------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
|
||||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
|
||||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
|
||||||
|
|
||||||
## Option 2: Build Docker Image with Model Inside
|
## Option 2: Build Docker Image with Model Inside
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,152 @@
|
|||||||
|
# Configuration Reference
|
||||||
|
|
||||||
|
Complete guide to all environment variables and configuration options for worker-vllm.
|
||||||
|
|
||||||
|
## LLM Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type/Choices | Description |
|
||||||
|
| ------------------------------ | ------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------- |
|
||||||
|
| `MODEL_NAME` | 'facebook/opt-125m' | `str` | Name or path of the Hugging Face model to use. |
|
||||||
|
| `MODEL_REVISION` | 'main' | `str` | Model revision to load (default: main). |
|
||||||
|
| `TOKENIZER` | None | `str` | Name or path of the Hugging Face tokenizer to use. |
|
||||||
|
| `SKIP_TOKENIZER_INIT` | False | `bool` | Skip initialization of tokenizer and detokenizer. |
|
||||||
|
| `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. |
|
||||||
|
| `TRUST_REMOTE_CODE` | `False` | `bool` | Trust remote code from Hugging Face. |
|
||||||
|
| `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. |
|
||||||
|
| `LOAD_FORMAT` | 'auto' | `str` | The format of the model weights to load. |
|
||||||
|
| `HF_TOKEN` | - | `str` | Hugging Face token for private and gated models. |
|
||||||
|
| `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. |
|
||||||
|
| `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8'] | Data type for KV cache storage. |
|
||||||
|
| `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. |
|
||||||
|
| `MAX_MODEL_LEN` | None | `int` | Model context length. |
|
||||||
|
| `GUIDED_DECODING_BACKEND` | 'outlines' | ['outlines', 'lm-format-enforcer'] | Which engine will be used for guided decoding by default. |
|
||||||
|
| `DISTRIBUTED_EXECUTOR_BACKEND` | None | ['ray', 'mp'] | Backend to use for distributed serving. |
|
||||||
|
| `WORKER_USE_RAY` | False | `bool` | Deprecated, use --distributed-executor-backend=ray. |
|
||||||
|
| `PIPELINE_PARALLEL_SIZE` | 1 | `int` | Number of pipeline stages. |
|
||||||
|
| `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. |
|
||||||
|
| `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. |
|
||||||
|
| `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. |
|
||||||
|
| `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. |
|
||||||
|
| `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. |
|
||||||
|
| `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. |
|
||||||
|
| `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. |
|
||||||
|
| `SEED` | 0 | `int` | Random seed for operations. |
|
||||||
|
| `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. |
|
||||||
|
| `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. |
|
||||||
|
| `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. |
|
||||||
|
| `MAX_LOGPROBS` | 20 | `int` | Max number of log probs to return when logprobs is specified in SamplingParams. |
|
||||||
|
| `DISABLE_LOG_STATS` | False | `bool` | Disable logging statistics. |
|
||||||
|
| `QUANTIZATION` | None | ['awq', 'squeezellm', 'gptq', 'bitsandbytes'] | Method used to quantize the weights. |
|
||||||
|
| `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. |
|
||||||
|
| `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. |
|
||||||
|
| `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. |
|
||||||
|
| `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. |
|
||||||
|
| `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |
|
||||||
|
|
||||||
|
## LoRA (Low-Rank Adaptation) Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type | Description |
|
||||||
|
| --------------------------- | ------- | ------------------------------------------ | --------------------------------------------------------------------------------------------------------- |
|
||||||
|
| `ENABLE_LORA` | False | `bool` | If True, enable handling of LoRA adapters. |
|
||||||
|
| `MAX_LORAS` | 1 | `int` | Max number of LoRAs in a single batch. |
|
||||||
|
| `MAX_LORA_RANK` | 16 | `int` | Max LoRA rank. |
|
||||||
|
| `LORA_EXTRA_VOCAB_SIZE` | 256 | `int` | Maximum size of extra vocabulary for LoRA adapters. |
|
||||||
|
| `LORA_DTYPE` | 'auto' | ['auto', 'float16', 'bfloat16', 'float32'] | Data type for LoRA. |
|
||||||
|
| `LONG_LORA_SCALING_FACTORS` | None | `tuple` | Specify multiple scaling factors for LoRA adapters. |
|
||||||
|
| `MAX_CPU_LORAS` | None | `int` | Maximum number of LoRAs to store in CPU memory. |
|
||||||
|
| `FULLY_SHARDED_LORAS` | False | `bool` | Enable fully sharded LoRA layers. |
|
||||||
|
| `LORA_MODULES` | `[]` | `list[dict]` | Add lora adapters from Hugging Face `[{"name": "xx", "path": "xxx/xxxx", "base_model_name": "xxx/xxxx"}]` |
|
||||||
|
|
||||||
|
## Speculative Decoding Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type/Choices | Description |
|
||||||
|
| ------------------------------------------------ | ------------------- | --------------------------------------------------- | ----------------------------------------------------------------------------------------- |
|
||||||
|
| `SCHEDULER_DELAY_FACTOR` | 0.0 | `float` | Apply a delay before scheduling next prompt. |
|
||||||
|
| `ENABLE_CHUNKED_PREFILL` | False | `bool` | Enable chunked prefill requests. |
|
||||||
|
| `SPECULATIVE_MODEL` | None | `str` | The name of the draft model to be used in speculative decoding. |
|
||||||
|
| `NUM_SPECULATIVE_TOKENS` | None | `int` | The number of speculative tokens to sample from the draft model. |
|
||||||
|
| `SPECULATIVE_DRAFT_TENSOR_PARALLEL_SIZE` | None | `int` | Number of tensor parallel replicas for the draft model. |
|
||||||
|
| `SPECULATIVE_MAX_MODEL_LEN` | None | `int` | The maximum sequence length supported by the draft model. |
|
||||||
|
| `SPECULATIVE_DISABLE_BY_BATCH_SIZE` | None | `int` | Disable speculative decoding if the number of enqueue requests is larger than this value. |
|
||||||
|
| `NGRAM_PROMPT_LOOKUP_MAX` | None | `int` | Max size of window for ngram prompt lookup in speculative decoding. |
|
||||||
|
| `NGRAM_PROMPT_LOOKUP_MIN` | None | `int` | Min size of window for ngram prompt lookup in speculative decoding. |
|
||||||
|
| `SPEC_DECODING_ACCEPTANCE_METHOD` | 'rejection_sampler' | ['rejection_sampler', 'typical_acceptance_sampler'] | Specify the acceptance method for draft token verification in speculative decoding. |
|
||||||
|
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_THRESHOLD` | None | `float` | Set the lower bound threshold for the posterior probability of a token to be accepted. |
|
||||||
|
| `TYPICAL_ACCEPTANCE_SAMPLER_POSTERIOR_ALPHA` | None | `float` | A scaling factor for the entropy-based threshold for token acceptance. |
|
||||||
|
|
||||||
|
## System Performance Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type/Choices | Description |
|
||||||
|
| ------------------------------ | ------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
|
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` | Sets GPU VRAM utilization. |
|
||||||
|
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` | Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
|
||||||
|
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` | Token block size for contiguous chunks of tokens. |
|
||||||
|
| `SWAP_SPACE` | `4` | `int` | CPU swap space size (GiB) per GPU. |
|
||||||
|
| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
||||||
|
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
|
||||||
|
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
|
||||||
|
|
||||||
|
## Tokenizer Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type/Choices | Description |
|
||||||
|
| ---------------------- | ------- | ----------------------------------- | ------------------------------------------------------------------------------------------------- |
|
||||||
|
| `TOKENIZER_NAME` | `None` | `str` | Tokenizer repository to use a different tokenizer than the model's default. |
|
||||||
|
| `TOKENIZER_REVISION` | `None` | `str` | Tokenizer revision to load. |
|
||||||
|
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template | Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
|
||||||
|
|
||||||
|
## Streaming & Batch Settings
|
||||||
|
|
||||||
|
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker.
|
||||||
|
|
||||||
|
| Variable | Default | Type/Choices | Description |
|
||||||
|
| ---------------------------------- | ------- | ------------ | --------------------------------------------------------------------------------------------------------- |
|
||||||
|
| `DEFAULT_BATCH_SIZE` | `50` | `int` | Default and Maximum batch size for token streaming to reduce HTTP calls. |
|
||||||
|
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` | Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
|
||||||
|
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` | Growth factor for dynamic batch size. |
|
||||||
|
|
||||||
|
## OpenAI Compatibility Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type/Choices | Description |
|
||||||
|
| ----------------------------------- | ----------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
|
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` | Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
|
||||||
|
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` | Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
|
||||||
|
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` | Role of the LLM's Response in OpenAI Chat Completions. |
|
||||||
|
| `ENABLE_AUTO_TOOL_CHOICE` | `false` | `bool` | Enables automatic tool selection for supported models. Set to `true` to activate. |
|
||||||
|
| `TOOL_CALL_PARSER` | `None` | `str` | Specifies the parser for tool calls. Options: `mistral`, `hermes`, `llama3_json`, `llama4_json`, `llama4_pythonic`, `granite`, `granite-20b-fc`, `deepseek_v3`, `internlm`, `jamba`, `phi4_mini_json`, `pythonic` |
|
||||||
|
|
||||||
|
## Serverless & Concurrency Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type/Choices | Description |
|
||||||
|
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
|
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||||
|
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||||
|
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||||
|
|
||||||
|
## Advanced Settings
|
||||||
|
|
||||||
|
| Variable | Default | Type | Description |
|
||||||
|
| --------------------------- | ------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||||
|
| `MODEL_LOADER_EXTRA_CONFIG` | None | `dict` | Extra config for model loader. |
|
||||||
|
| `PREEMPTION_MODE` | None | `str` | If 'recompute', the engine performs preemption-aware recomputation. If 'save', the engine saves activations into the CPU memory as preemption happens. |
|
||||||
|
| `PREEMPTION_CHECK_PERIOD` | 1.0 | `float` | How frequently the engine checks if a preemption happens. |
|
||||||
|
| `PREEMPTION_CPU_CAPACITY` | 2 | `float` | The percentage of CPU memory used for the saved activations. |
|
||||||
|
| `DISABLE_LOGGING_REQUEST` | False | `bool` | Disable logging requests. |
|
||||||
|
| `MAX_LOG_LEN` | None | `int` | Max number of prompt characters or prompt ID numbers being printed in log. |
|
||||||
|
|
||||||
|
## Docker Build Arguments
|
||||||
|
|
||||||
|
These variables are used when building custom Docker images with models baked in:
|
||||||
|
|
||||||
|
| Variable | Default | Type | Description |
|
||||||
|
| --------------------- | ---------------- | ----- | ------------------------------------------------- |
|
||||||
|
| `BASE_PATH` | `/runpod-volume` | `str` | Storage directory for huggingface cache and model |
|
||||||
|
| `WORKER_CUDA_VERSION` | `12.1.0` | `str` | CUDA version for the worker image |
|
||||||
|
|
||||||
|
## Deprecated Variables
|
||||||
|
|
||||||
|
⚠️ **The following variables are deprecated and will be removed in future versions:**
|
||||||
|
|
||||||
|
| Old Variable | New Variable | Note |
|
||||||
|
| ---------------------------- | ------------------------ | --------------------- |
|
||||||
|
| `MAX_CONTEXT_LEN_TO_CAPTURE` | `MAX_SEQ_LEN_TO_CAPTURE` | Use new variable name |
|
||||||
|
| `kv_cache_dtype=fp8_e5m2` | `kv_cache_dtype=fp8` | Simplified fp8 format |
|
||||||
Binary file not shown.
|
After Width: | Height: | Size: 86 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 1.1 MiB |
Reference in New Issue
Block a user