Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
66e1b1605b | ||
|
|
60c8f257a8 | ||
|
|
fae16e7ee1 | ||
|
|
2becd35345 |
+1
-11
@@ -38,16 +38,6 @@
|
|||||||
"required": true
|
"required": true
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"key": "HF_TOKEN",
|
|
||||||
"input": {
|
|
||||||
"name": "Access Token",
|
|
||||||
"type": "string",
|
|
||||||
"description": "Hugging Face access token for gated & private models",
|
|
||||||
"default": "",
|
|
||||||
"required": false
|
|
||||||
}
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"key": "TOKENIZER",
|
"key": "TOKENIZER",
|
||||||
"input": {
|
"input": {
|
||||||
@@ -945,7 +935,7 @@
|
|||||||
"name": "Max Concurrency",
|
"name": "Max Concurrency",
|
||||||
"type": "number",
|
"type": "number",
|
||||||
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
|
"description": "Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency",
|
||||||
"default": 300,
|
"default": 30,
|
||||||
"advanced": true
|
"advanced": true
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
|||||||
+1
-1
@@ -12,7 +12,7 @@ RUN --mount=type=cache,target=/root/.cache/pip \
|
|||||||
python3 -m pip install --upgrade -r /requirements.txt
|
python3 -m pip install --upgrade -r /requirements.txt
|
||||||
|
|
||||||
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
|
# Install vLLM (switching back to pip installs since issues that required building fork are fixed and space optimization is not as important since caching) and FlashInfer
|
||||||
RUN python3 -m pip install vllm==0.10.0 && \
|
RUN python3 -m pip install vllm==0.11.0 && \
|
||||||
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
|
python3 -m pip install flashinfer -i https://flashinfer.ai/whl/cu121/torch2.3
|
||||||
|
|
||||||
# Setup for Option 2: Building the Image with the Model included
|
# Setup for Option 2: Building the Image with the Model included
|
||||||
|
|||||||
@@ -57,7 +57,7 @@ Configure worker-vllm using environment variables:
|
|||||||
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
| `ENABLE_AUTO_TOOL_CHOICE` | Enable automatic tool selection | false | boolean (true or false) |
|
||||||
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
| `TOOL_CALL_PARSER` | Parser for tool calls | | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. |
|
||||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
|
||||||
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
|
| `MAX_CONCURRENCY` | Maximum concurrent requests | 30 | Integer |
|
||||||
|
|
||||||
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
|
||||||
|
|
||||||
|
|||||||
@@ -119,7 +119,7 @@ The way this works is that the first request will have a batch size of `DEFAULT_
|
|||||||
|
|
||||||
| Variable | Default | Type/Choices | Description |
|
| Variable | Default | Type/Choices | Description |
|
||||||
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||||
|
|
||||||
|
|||||||
+3
-2
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
|
|||||||
|
|
||||||
- `src/engine_args.py`: Centralized configuration management
|
- `src/engine_args.py`: Centralized configuration management
|
||||||
- `src/constants.py`: Default values for core settings
|
- `src/constants.py`: Default values for core settings
|
||||||
- `worker-config.json`: UI form generation for RunPod console
|
- `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
|
||||||
|
- `worker-config.json`: UI form generation for RunPod console (if exists)
|
||||||
|
|
||||||
## Core Development Concepts
|
## Core Development Concepts
|
||||||
|
|
||||||
@@ -222,7 +223,7 @@ src/
|
|||||||
|
|
||||||
### 2. **Concurrency Patterns**
|
### 2. **Concurrency Patterns**
|
||||||
|
|
||||||
- **Max Concurrency**: 300 concurrent requests by default
|
- **Max Concurrency**: 30 concurrent requests by default
|
||||||
- **vLLM Queuing**: Internal request batching and scheduling
|
- **vLLM Queuing**: Internal request batching and scheduling
|
||||||
- **RunPod Integration**: Concurrency modifier for auto-scaling
|
- **RunPod Integration**: Concurrency modifier for auto-scaling
|
||||||
|
|
||||||
|
|||||||
+1
-1
@@ -1,4 +1,4 @@
|
|||||||
DEFAULT_BATCH_SIZE = 50
|
DEFAULT_BATCH_SIZE = 50
|
||||||
DEFAULT_MAX_CONCURRENCY = 300
|
DEFAULT_MAX_CONCURRENCY = 30
|
||||||
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
|
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
|
||||||
DEFAULT_MIN_BATCH_SIZE = 1
|
DEFAULT_MIN_BATCH_SIZE = 1
|
||||||
-1514
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user