Three coupled fixes verified end-to-end on a private fork (TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing): 1. tests.json allowedCudaVersions: 12.x → 13.0 The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0 (#288, #289), but tests.json was still pinned to 12.5–12.9, so the test pod was scheduled on a GPU with driver < 13.0 and container init failed at the nvidia-container-cli hook with "unsatisfied condition: cuda>=13.0". 2. requirements.txt kernels<0.15 huggingface/kernels v0.15.1 tightened LayerRepository to require a revision or version argument (https://github.com/huggingface/kernels/pull/544). transformers >=5 still constructs LayerRepository(repo_id=..., layer_name=...) without either, so worker import raised ValueError during `from transformers import ...`. 0.14.1 is the last safe release. 3. tests.json timeout 30000 → 300000 vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090 for SmolLM2-135M takes ~60–70s before the first request can be served. The previous 30s per-test timeout fired before the worker came up, producing "context cancelled or timed out: context deadline exceeded" for every test even when the worker was healthy. 300s gives enough headroom for cold start + the actual inference call. Refs: DR-1161
Run LLMs using vLLM with an OpenAI-compatible API
Current vLLM version: 0.20.2
Endpoint Configuration
All behaviour is controlled through environment variables:
| Environment Variable | Description | Default | Options |
|---|---|---|---|
MODEL_NAME |
Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
HF_TOKEN |
HuggingFace access token for gated/private models | Your HuggingFace access token | |
MAX_MODEL_LEN |
Model's maximum context length | Integer (e.g., 4096) | |
QUANTIZATION |
Quantization method | "awq", "gptq", "squeezellm", "bitsandbytes" | |
TENSOR_PARALLEL_SIZE |
Number of GPUs | 1 | Integer |
GPU_MEMORY_UTILIZATION |
Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
MAX_NUM_SEQS |
Maximum number of sequences per iteration | 256 | Integer |
CUSTOM_CHAT_TEMPLATE |
Custom chat template override | Jinja2 template string | |
ENABLE_AUTO_TOOL_CHOICE |
Enable automatic tool selection | false | boolean (true or false) |
TOOL_CALL_PARSER |
Parser for tool calls | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. | |
REASONING_PARSER |
Parser for reasoning-capable models | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" | |
OPENAI_SERVED_MODEL_NAME_OVERRIDE |
Override served model name in API | String | |
MAX_CONCURRENCY |
Maximum concurrent requests | 300 | Integer |
ENFORCE_EAGER |
If True, we will disable CUDA graph and always execute the model in eager mode. If False, we will use CUDA graph and eager execution in hybrid for maximal performance and flexibility. | true | boolean (true or false) |
Pass any vLLM engine arg not listed above by setting an env var with the UPPERCASED field name (e.g. MAX_MODEL_LEN=4096, ENABLE_CHUNKED_PREFILL=true). The worker auto-discovers all AsyncEngineArgs fields from env. See the vLLM engine args docs for all available options.
For complete configuration options, see the full configuration documentation.
Specify Transformers Version
To change the version of the Transformers library use the TRANSFORMERS_VERSION environment variable to specify the version you want to use. Note this might break the handler, so use for development purposes.
API Usage
This worker supports two API formats: RunPod native and OpenAI-compatible.
RunPod Native API
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
Chat Completions
{
"input": {
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "What is the capital of France?" }
],
"sampling_params": {
"max_tokens": 100,
"temperature": 0.7
}
}
}
Chat Completions (Streaming)
{
"input": {
"messages": [
{ "role": "user", "content": "Write a short story about a robot." }
],
"sampling_params": {
"max_tokens": 500,
"temperature": 0.8
},
"stream": true
}
}
Text Generation
For direct text generation without chat format:
{
"input": {
"prompt": "The capital of France is",
"sampling_params": {
"max_tokens": 64,
"temperature": 0.0
}
}
}
List Models
{
"input": {
"openai_route": "/v1/models"
}
}
OpenAI-Compatible API
For external clients and SDKs, use the /openai/v1 path prefix with your RunPod API key.
Chat Completions
Path: /openai/v1/chat/completions
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "What is the capital of France?" }
],
"max_tokens": 100,
"temperature": 0.7
}
Chat Completions (Streaming)
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"messages": [
{ "role": "user", "content": "Write a short story about a robot." }
],
"max_tokens": 500,
"temperature": 0.8,
"stream": true
}
Text Completions
Path: /openai/v1/completions
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"prompt": "The capital of France is",
"max_tokens": 100,
"temperature": 0.7
}
List Models
Path: /openai/v1/models
{}
OpenAI Responses API
Path: /openai/v1/responses
Supports the OpenAI Responses API format. Note: this route bypasses the RunPod queue and is served directly — use /openai/ prefixed paths rather than the RunPod job queue for these endpoints.
{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"input": "Tell me a joke."
}
Anthropic Messages API
Path: /openai/v1/messages
Supports the Anthropic Messages API format. Served directly, bypassing the RunPod queue.
{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"max_tokens": 256,
"messages": [
{"role": "user", "content": "Hello!"}
]
}
Response Format
Both APIs return the same response format:
{
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Paris." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
}
Usage
Below are minimal python snippets so you can copy-paste to get started quickly.
Replace
<ENDPOINT_ID>with your endpoint ID and<API_KEY>with a RunPod API key.
OpenAI compatible API
Minimal Python example using the official openai SDK:
from openai import OpenAI
import os
# Initialize the OpenAI Client with your Runpod API Key and Endpoint URL
client = OpenAI(
api_key=os.getenv("RUNPOD_API_KEY"),
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
)
Chat Completions (Non-Streaming)
response = client.chat.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
temperature=0,
max_tokens=100,
)
print(f"Response: {response.choices[0].message.content}")
Chat Completions (Streaming)
response_stream = client.chat.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
temperature=0,
max_tokens=100,
stream=True
)
for response in response_stream:
print(response.choices[0].delta.content or "", end="", flush=True)
RunPod Native API
import requests
response = requests.post(
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
headers={"Authorization": "Bearer <API_KEY>"},
json={
"input": {
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in simple terms"}
],
"sampling_params": {
"temperature": 0.7,
"max_tokens": 150
}
}
}
)
result = response.json()
print(result["output"])
Compatibility
For supported models, see the vLLM supported models documentation.
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
Documentation
- 🚀 Deployment Guide - Step-by-step setup
- 📖 Configuration Reference - All environment variables
- 🏗️ Advanced Deployment - Custom builds and strategies
- 🔧 Development Guide - Architecture and patterns
