8.6 KiB
Run LLMs using vLLM with an OpenAI-compatible API
Endpoint Configuration
All behaviour is controlled through environment variables:
| Environment Variable | Description | Default | Options |
|---|---|---|---|
MODEL_NAME |
Path of the model weights | "facebook/opt-125m" | Local folder or Hugging Face repo ID |
HF_TOKEN |
HuggingFace access token for gated/private models | Your HuggingFace access token | |
MAX_MODEL_LEN |
Model's maximum context length | Integer (e.g., 4096) | |
QUANTIZATION |
Quantization method | "awq", "gptq", "squeezellm", "bitsandbytes" | |
TENSOR_PARALLEL_SIZE |
Number of GPUs | 1 | Integer |
GPU_MEMORY_UTILIZATION |
Fraction of GPU memory to use | 0.95 | Float between 0.0 and 1.0 |
MAX_NUM_SEQS |
Maximum number of sequences per iteration | 256 | Integer |
CUSTOM_CHAT_TEMPLATE |
Custom chat template override | Jinja2 template string | |
ENABLE_AUTO_TOOL_CHOICE |
Enable automatic tool selection | false | boolean (true or false) |
TOOL_CALL_PARSER |
Parser for tool calls | "mistral", "hermes", "llama3_json", "granite", "deepseek_v3", etc. | |
REASONING_PARSER |
Parser for reasoning-capable models | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" | |
OPENAI_SERVED_MODEL_NAME_OVERRIDE |
Override served model name in API | String | |
MAX_CONCURRENCY |
Maximum concurrent requests | 300 | Integer |
Pass any vLLM engine arg not listed above by prefixing it with VLLM_RUNPOD_. The suffix maps to the vLLM AsyncEngineArgs field name (case-insensitive). For example, VLLM_RUNPOD_ENABLE_CHUNKED_PREFILL=true sets enable_chunked_prefill. See the vLLM engine args docs for all available options.
For complete configuration options, see the full configuration documentation.
API Usage
This worker supports two API formats: RunPod native and OpenAI-compatible.
RunPod Native API
For testing directly in the RunPod UI, use these examples in your endpoint's request tab.
Chat Completions
{
"input": {
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "What is the capital of France?" }
],
"sampling_params": {
"max_tokens": 100,
"temperature": 0.7
}
}
}
Chat Completions (Streaming)
{
"input": {
"messages": [
{ "role": "user", "content": "Write a short story about a robot." }
],
"sampling_params": {
"max_tokens": 500,
"temperature": 0.8
},
"stream": true
}
}
Text Generation
For direct text generation without chat format:
{
"input": {
"prompt": "The capital of France is",
"sampling_params": {
"max_tokens": 64,
"temperature": 0.0
}
}
}
List Models
{
"input": {
"openai_route": "/v1/models"
}
}
OpenAI-Compatible API
For external clients and SDKs, use the /openai/v1 path prefix with your RunPod API key.
Chat Completions
Path: /openai/v1/chat/completions
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "What is the capital of France?" }
],
"max_tokens": 100,
"temperature": 0.7
}
Chat Completions (Streaming)
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"messages": [
{ "role": "user", "content": "Write a short story about a robot." }
],
"max_tokens": 500,
"temperature": 0.8,
"stream": true
}
Text Completions
Path: /openai/v1/completions
{
"model": "meta-llama/Llama-2-7b-chat-hf",
"prompt": "The capital of France is",
"max_tokens": 100,
"temperature": 0.7
}
List Models
Path: /openai/v1/models
{}
Response Format
Both APIs return the same response format:
{
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Paris." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 9, "completion_tokens": 1, "total_tokens": 10 }
}
Usage
Below are minimal python snippets so you can copy-paste to get started quickly.
Replace
<ENDPOINT_ID>with your endpoint ID and<API_KEY>with a RunPod API key.
OpenAI compatible API
Minimal Python example using the official openai SDK:
from openai import OpenAI
import os
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
client = OpenAI(
api_key=os.getenv("RUNPOD_API_KEY"),
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
)
Chat Completions (Non-Streaming)
response = client.chat.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
temperature=0,
max_tokens=100,
)
print(f"Response: {response.choices[0].message.content}")
Chat Completions (Streaming)
response_stream = client.chat.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
messages=[{"role": "user", "content": "Explain quantum computing in simple terms"}],
temperature=0,
max_tokens=100,
stream=True
)
for response in response_stream:
print(response.choices[0].delta.content or "", end="", flush=True)
RunPod Native API
import requests
response = requests.post(
"https://api.runpod.ai/v2/<ENDPOINT_ID>/run",
headers={"Authorization": "Bearer <API_KEY>"},
json={
"input": {
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in simple terms"}
],
"sampling_params": {
"temperature": 0.7,
"max_tokens": 150
}
}
}
)
result = response.json()
print(result["output"])
Compatibility
For supported models, see the vLLM supported models documentation.
Anything not recognized by worker-vllm is forwarded to vLLM's engine, so advanced options in the vLLM docs (guided generation, LoRA, speculative decoding, etc.) also work.
Documentation
- 🚀 Deployment Guide - Step-by-step setup
- 📖 Configuration Reference - All environment variables
- 🏗️ Advanced Deployment - Custom builds and strategies
- 🔧 Development Guide - Architecture and patterns
