Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
Release / release (push) Waiting to run
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
This commit is contained in:
@@ -119,7 +119,7 @@ The way this works is that the first request will have a batch size of `DEFAULT_
|
||||
|
||||
| Variable | Default | Type/Choices | Description |
|
||||
| ---------------------- | ------- | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `MAX_CONCURRENCY` | `300` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `MAX_CONCURRENCY` | `30` | `int` | Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `DISABLE_LOG_STATS` | False | `bool` | Enables or disables vLLM stats logging. |
|
||||
| `DISABLE_LOG_REQUESTS` | False | `bool` | Enables or disables vLLM request logging. |
|
||||
|
||||
|
||||
+3
-2
@@ -51,7 +51,8 @@ RunPod Request → handler.py → JobInput → Engine Selection → vLLM Generat
|
||||
|
||||
- `src/engine_args.py`: Centralized configuration management
|
||||
- `src/constants.py`: Default values for core settings
|
||||
- `worker-config.json`: UI form generation for RunPod console
|
||||
- `.runpod/hub.json`: Hub UI configuration (CRITICAL: always update when changing defaults)
|
||||
- `worker-config.json`: UI form generation for RunPod console (if exists)
|
||||
|
||||
## Core Development Concepts
|
||||
|
||||
@@ -222,7 +223,7 @@ src/
|
||||
|
||||
### 2. **Concurrency Patterns**
|
||||
|
||||
- **Max Concurrency**: 300 concurrent requests by default
|
||||
- **Max Concurrency**: 30 concurrent requests by default
|
||||
- **vLLM Queuing**: Internal request batching and scheduling
|
||||
- **RunPod Integration**: Concurrency modifier for auto-scaling
|
||||
|
||||
|
||||
Reference in New Issue
Block a user