Compare commits

..
57 Commits
Author SHA1 Message Date
chrisvelaandGitHub 50aba8fb57 Merge pull request #293 from runpod-workers/fix/update-deep-gemm-hub-value
Release / release (push) Waiting to run
fix: update VLLM_USE_DEEP_GEMM hub to default to 0
2026-05-27 11:51:48 -05:00
velaraptor-runpod 9edc5715ce fix: update configuration.md 2026-05-27 11:33:37 -05:00
velaraptor-runpod 4c91f2c5b5 fix: update VLLM_USE_DEEP_GEMM hub to default to 0 2026-05-27 11:26:52 -05:00
chrisvelaandGitHub 6265b99348 Merge pull request #288 from runpod-workers/feat/0.20.0
feat: update to 0.20.2
2026-05-26 17:17:45 -05:00
velaraptor-runpod 026f8d700b fix: specify deepgemm commit version 2026-05-21 18:46:54 -05:00
chrisvelaandGitHub 146bdb0252 Merge branch 'main' into feat/0.20.0 2026-05-21 16:04:39 -05:00
velaraptor-runpod da01193a3d chore: fix readme 2026-05-21 15:57:40 -05:00
velaraptor-runpod c2e6cc9f61 chore: update readme with correct vllm version 2026-05-21 15:39:19 -05:00
velaraptor-runpod 69968a6b39 chore: fix logging, warning for text prompt 2026-05-21 15:34:09 -05:00
velaraptor-runpod 32b29d4c6c fix: add deepgemm, update base image and hub for cuda 13.0 2026-05-20 17:43:42 -05:00
velaraptor-runpod dcea4fc4f9 fix dockerfile 2026-05-15 12:08:31 -04:00
velaraptor-runpod 9c139e8ceb update: update to 0.20.1, update dockerfile to cuda 13 2026-05-15 11:35:53 -04:00
velaraptor-runpod 678bb4be8f feat: update to 0.20.1 for patch fixes 2026-05-07 11:58:22 -05:00
chrisvelaandGitHub 87d7365126 Merge pull request #292 from runpod-workers/fix/open-ai
Release / release (push) Waiting to run
fix: fix warmup
2026-05-01 18:06:11 -05:00
velaraptor-runpod 0e83616f93 fix: fix warmup 2026-05-01 17:39:51 -05:00
chrisvelaandGitHub ab6d39dcf8 Merge pull request #291 from runpod-workers/bug/281-hf-token
Release / release (push) Waiting to run
bug: fix hf-token being passed in engineargs
2026-05-01 15:32:04 -05:00
velaraptor-runpod ed315a175e merge main 2026-05-01 15:28:05 -05:00
velaraptor-runpod 73f030ae5e Merge branch 'main' into bug/281-hf-token 2026-05-01 14:33:51 -05:00
chrisvelaandGitHub 8a099c1723 Merge pull request #287 from runpod-workers/feat/0.19.1
feat: update to 0.19.1
2026-05-01 14:29:39 -05:00
velaraptor-runpod ff87840a58 fix: add enforce_eager as true, add pytorch_alloc_conf to expandle_segments to True for OOM, for hub defaults 2026-05-01 14:13:03 -05:00
velaraptor-runpod 7dc853b1fe chore: remove release to trigger on release publish, just use tags. duplicate 2026-05-01 10:13:32 -05:00
velaraptor-runpod a8b754b92a merge main 2026-05-01 10:12:43 -05:00
chrisvelaandGitHub 6357aeda51 Merge pull request #290 from runpod-workers/fix/fix-old-actions
Release / release (push) Waiting to run
fix: fix old github actions, trigger release on publish release
2026-05-01 10:01:22 -05:00
velaraptor-runpod 0140b29c44 chore: add release notes to slack notification 2026-05-01 09:48:43 -05:00
velaraptor-runpod 0cb8aeae77 chore: add specific runpod version 2026-05-01 09:47:23 -05:00
chrisvelaandGitHub cd8f9e9560 Merge branch 'main' into fix/fix-old-actions 2026-04-30 21:57:45 -05:00
chrisvelaandGitHub 895fd25fac Merge pull request #286 from runpod-workers/feat/0.18.1
feat: update vllm to 0.18.1
2026-04-30 21:56:07 -05:00
velaraptor-runpod 7bb8df73af bug: fix hf-token being passed in engineargs 2026-04-30 20:37:30 -05:00
velaraptor-runpod 747cdf5891 chore: update readme vllm version 2026-04-30 20:26:23 -05:00
velaraptor-runpod 3d4af5df9b chore: update readme vllm version 2026-04-30 20:25:35 -05:00
velaraptor-runpod 72547aa3bb chore: update readme vllm version 2026-04-30 20:25:02 -05:00
velaraptor-runpod 577fd8c3c3 fix: fix old github actions, trigger release on publish release 2026-04-30 20:22:24 -05:00
chrisvelaandGitHub f49f35456e Merge pull request #285 from runpod-workers/feat/0.17.1
feat: update vllm to 0.17.1
2026-04-30 20:05:45 -05:00
chrisvelaandGitHub cff7b09ef2 Merge pull request #289 from runpod-workers/feat/add-notifications
feat: add notifications for new prs, issues, and new releases of vllm
2026-04-30 20:03:48 -05:00
velaraptor-runpod 04b342c675 fix permissions 2026-04-30 20:00:48 -05:00
velaraptor-runpod 5b29643799 feat: add notifications for new prs, issues, and new releases of vllm 2026-04-30 19:55:29 -05:00
velaraptor-runpod 4f8a16df5d fix: update transformers to >=5 2026-04-30 19:41:09 -05:00
velaraptor-runpodandClaude Sonnet 4.6 22356ee2b3 feat: upgrade vLLM to 0.20.0
- Bump vllm[flashinfer] to 0.20.0 in Dockerfile
- Remove io_processor param from OpenAIServingRender (dropped in 0.20.0)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 18:39:07 -05:00
velaraptor-runpodandClaude Sonnet 4.6 fa42ecd79a fix: resolve lowercase HF cache paths when MODEL_NAME uses original casing
Fixes FDE-174. Some model stores (e.g. RunPod pre-cached network volumes)
normalize repo IDs to lowercase. HuggingFace Hub caches using the original
casing, so MODEL_NAME=Qwen/Qwen2.5-Coder-32B-Instruct-AWQ would miss a
cache stored as models--qwen--qwen2.5-coder-32b-instruct-awq/ and attempt
a redundant download that fails on limited container storage.

If the exact-case HF cache directory is absent but a lowercase variant
exists, the latest snapshot path is returned directly so vLLM loads from
disk. Absolute paths and models with no lowercase cache are unchanged.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 17:54:14 -05:00
velaraptor-runpodandClaude Sonnet 4.6 178c72238e fix: surface LORA_MODULES parse failures instead of silently loading zero adapters
Fixes FDE-194. Previously a malformed LORA_MODULES value was swallowed at
info level and the engine would start with no LoRA adapters, causing 500s
on any request using an adapter model name (e.g. npc-sim-*).

Changes:
- Log at error level when LORA_MODULES cannot be parsed as JSON
- Log at error level when individual adapter dicts fail LoRAModulePath validation
- Log a final error when all adapters fail to load so the cause is obvious
- Accept a single adapter dict (not just an array) for convenience
- Return early when LORA_MODULES is unset to skip unnecessary parsing

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 17:49:09 -05:00
velaraptor-runpodandClaude Sonnet 4.6 e6950bdebd feat: upgrade vLLM to 0.19.1
- Bump vllm[flashinfer] to 0.19.1 in Dockerfile
- Add OpenAIServingRender (new required dependency in 0.19.x serving layer)
- Pass openai_serving_render to all four serving class constructors
- Remove log_error_stack param (removed upstream in 0.19.x)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 17:41:11 -05:00
velaraptor-runpod a774cefe85 fix: clean workspace file 2026-04-30 17:29:00 -05:00
velaraptor-runpod 9de17d49b7 feat: update vllm to 0.18.1 2026-04-30 17:25:23 -05:00
velaraptor-runpod 296556a6f7 feat: update vllm to 0.17.1 2026-04-30 17:21:03 -05:00
chrisvelaandGitHub a1544ea70d Merge pull request #277 from runpod-workers/feat/lmcache
feat: uv installer, LMCache support, add /v1/responses and /v1/messages endpoints
2026-04-28 19:44:00 -03:00
Tim Pietrusky 3403889528 fix: address review comments on responses/messages handlers and lmcache guard
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
  except blocks of _handle_responses_request and _handle_messages_request;
  emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
  blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
  actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
  and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
2026-04-23 11:28:50 +02:00
Tim Pietrusky f299204770 docs: fix anthropic messages path and missing comma in routes list 2026-04-23 10:54:24 +02:00
velaraptor-runpod 1ed25eea20 update readme with responses and messages routes 2026-04-16 15:30:30 -05:00
velaraptor-runpod dc4ad7ddeb fix: kv_transfer_config is dataclass not json, fix for env var 2026-04-09 14:33:11 -05:00
velaraptor-runpod 9035b0e07f fix: lmcache version 2026-04-09 14:32:19 -05:00
velaraptor-runpod c979f0020f update readmes on TRANSFORMERS_VERSION 2026-04-06 16:54:08 -05:00
velaraptor-runpod 6fbd480a26 fix: allow for transformers_version 2026-04-06 16:49:32 -05:00
velaraptor-runpod 3ef1fb8e7b requested changes 2026-04-03 15:53:40 -05:00
velaraptor-runpod 30e8514d63 Runpod not RunPod 2026-03-17 20:59:52 -05:00
velaraptor-runpod d9808815ee feat: update requirements.txt 2026-03-17 19:56:39 -05:00
velaraptor-runpod d8ed3b5353 feat: update 0.16.0, add lmcache 2026-03-17 19:54:11 -05:00
velaraptor-runpod 4c4e039565 feat: add messages route for anthropic/claude 2026-03-17 19:51:50 -05:00
14 changed files with 756 additions and 114 deletions
+32 -31
View File
@@ -9,59 +9,60 @@ on:
workflow_dispatch: workflow_dispatch:
permissions:
contents: write
pull-requests: write
jobs: jobs:
check_dep: check_dep:
runs-on: ubuntu-latest runs-on: ubuntu-latest
name: Check python requirements file and update name: Check python requirements file and update
steps: steps:
- name: Checkout - name: Checkout
uses: actions/checkout@v2 uses: actions/checkout@v4
- name: Check for new package version and update - name: Check for new package version and update
run: | run: |
echo "Fetching the current runpod version from requirements.txt..." echo "Fetching current runpod version from requirements.txt..."
# Get current version, allowing both == and ~= in the search pattern # Match runpod with any version specifier or no specifier at all
current_version=$(grep -oP 'runpod[~=]{1,2}\K[^"]+' ./builder/requirements.txt) current_version=$(grep -oP '^runpod([~>=!<]{1,2}\K[\d.]+)?' ./builder/requirements.txt | grep -oP '[\d.]+' || echo "")
echo "Current version: $current_version" echo "Current version: ${current_version:-unset}"
# Extract major and minor from current version echo "Fetching latest runpod version from PyPI..."
current_major_minor=$(echo $current_version | cut -d. -f1,2) new_version=$(curl -sf https://pypi.org/pypi/runpod/json | jq -r .info.version)
echo "Current major.minor: $current_major_minor"
echo "Fetching the latest runpod version from PyPI..."
# Get new version from PyPI
new_version=$(curl -s https://pypi.org/pypi/runpod/json | jq -r .info.version)
echo "NEW_VERSION_ENV=$new_version" >> $GITHUB_ENV echo "NEW_VERSION_ENV=$new_version" >> $GITHUB_ENV
echo "New version: $new_version" echo "New version: $new_version"
# Extract major and minor from new version
new_major_minor=$(echo $new_version | cut -d. -f1,2)
echo "New major.minor: $new_major_minor"
if [ -z "$new_version" ]; then if [ -z "$new_version" ]; then
echo "ERROR: Failed to fetch the new version from PyPI." echo "ERROR: Failed to fetch new version from PyPI."
exit 1 exit 1
fi fi
# Check if the major or minor version is different if [ -z "$current_version" ]; then
if [ "$current_major_minor" = "$new_major_minor" ]; then echo "No version pin found — pinning to $new_version."
echo "No update needed. The new version ($new_major_minor) is within the allowed range (~= $current_major_minor)." else
current_major_minor=$(echo "$current_version" | cut -d. -f1,2)
new_major_minor=$(echo "$new_version" | cut -d. -f1,2)
echo "Current major.minor: $current_major_minor New major.minor: $new_major_minor"
if [ "$current_major_minor" = "$new_major_minor" ]; then
echo "No update needed. New version ($new_version) is within ~= $current_major_minor range."
exit 0 exit 0
fi
echo "New major/minor detected ($new_major_minor). Updating requirements.txt..."
fi fi
echo "New major/minor detected ($new_major_minor). Updating requirements.txt..." # Replace any `runpod`, `runpod==x`, `runpod~=x`, etc. with pinned version
sed -i "s|^runpod.*|runpod~=$new_version|" ./builder/requirements.txt
# Update requirements.txt, preserving the existing constraint type (~= or ==) echo "requirements.txt updated."
sed -i "s/runpod[~=][^ ]*/runpod~=$new_version/" ./builder/requirements.txt
echo "requirements.txt has been updated."
- name: Create Pull Request - name: Create Pull Request
uses: peter-evans/create-pull-request@v3 uses: peter-evans/create-pull-request@v7
with: with:
token: ${{ secrets.GITHUB_TOKEN }} token: ${{ secrets.GITHUB_TOKEN }}
commit-message: Update runpod package version commit-message: "chore: update runpod to ${{ env.NEW_VERSION_ENV }}"
title: Update runpod package version title: "chore: update runpod to ${{ env.NEW_VERSION_ENV }}"
body: The package version has been updated to ${{ env.NEW_VERSION_ENV }} body: The `runpod` package has been updated to `${{ env.NEW_VERSION_ENV }}`.
branch: runpod-package-update branch: runpod-package-update
+43 -12
View File
@@ -3,7 +3,7 @@ name: Release
on: on:
push: push:
tags: tags:
- "v[0-9]+.[0-9]+.[0-9]+*" # Trigger on version tags like v1.0.0, v2.1.0, etc. - "v[0-9]+.[0-9]+.[0-9]+*"
workflow_dispatch: workflow_dispatch:
inputs: inputs:
version: version:
@@ -53,16 +53,13 @@ jobs:
# Determine version based on trigger type # Determine version based on trigger type
if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
# Manual trigger: use input version
VERSION="${{ github.event.inputs.version }}" VERSION="${{ github.event.inputs.version }}"
echo "RELEASE_VERSION=${VERSION}" >> $GITHUB_ENV elif [[ "${{ github.event_name }}" == "release" ]]; then
echo "IS_MANUAL_RELEASE=true" >> $GITHUB_ENV VERSION="${{ github.event.release.tag_name }}"
else else
# Tag trigger: use tag name (remove refs/tags/ prefix)
VERSION=${GITHUB_REF#refs/tags/} VERSION=${GITHUB_REF#refs/tags/}
echo "RELEASE_VERSION=${VERSION}" >> $GITHUB_ENV
echo "IS_MANUAL_RELEASE=false" >> $GITHUB_ENV
fi fi
echo "RELEASE_VERSION=${VERSION}" >> $GITHUB_ENV
- name: Build and push the images to Docker Hub - name: Build and push the images to Docker Hub
uses: docker/bake-action@v2 uses: docker/bake-action@v2
@@ -76,11 +73,45 @@ jobs:
- name: Release Summary - name: Release Summary
run: | run: |
echo "🚀 Release completed!" echo "Release completed!"
echo "Version: ${{ env.RELEASE_VERSION }}" echo "Version: ${{ env.RELEASE_VERSION }}"
echo "Docker Image: ${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}" echo "Docker Image: ${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}"
if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
echo "Trigger: Manual workflow dispatch" - name: Fetch Release Notes
else run: |
echo "Trigger: GitHub release (tag: ${{ github.ref_name }})" RESPONSE=$(curl -sf \
-H "Authorization: token ${{ github.token }}" \
"https://api.github.com/repos/${{ github.repository }}/releases/tags/${{ env.RELEASE_VERSION }}" 2>/dev/null) || true
if [[ -n "$RESPONSE" ]]; then
NOTES=$(echo "$RESPONSE" | jq -r '.body // empty')
fi fi
printf '%s' "${NOTES:-No release notes available.}" > /tmp/release_notes.txt
- name: Notify Slack
run: |
jq -n \
--arg version "${{ env.RELEASE_VERSION }}" \
--arg docker "${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}" \
--rawfile notes /tmp/release_notes.txt \
--arg url "https://github.com/${{ github.repository }}/releases/tag/${{ env.RELEASE_VERSION }}" \
'{
text: (":rocket: New :runpod-new-whiteonpurple: Runpod worker-vllm release: *" + $version + "*"),
blocks: [
{
type: "section",
text: {
type: "mrkdwn",
text: (":banana-dance: *New Release — worker-vllm " + $version + "*\n*Docker:* `" + $docker + "`\n<" + $url + "|View release on GitHub>")
}
},
{
type: "section",
text: {
type: "mrkdwn",
text: ("*Release Notes:*\n" + $notes)
}
}
]
}' | curl -sf -X POST "${{ secrets.SLACK_WEBHOOK_URL }}" \
-H "Content-Type: application/json" \
-d @-
@@ -0,0 +1,41 @@
name: Slack PR Notifications
on:
pull_request:
types: [opened]
issues:
types: [opened]
permissions:
contents: read
jobs:
notify:
runs-on: ubuntu-latest
steps:
- name: Notify Slack - New PR
if: github.event_name == 'pull_request'
run: |
curl -sf -X POST "${{ secrets.SLACK_WEBHOOK_URL }}" \
-H "Content-Type: application/json" \
-d '{
"text": ":rocket: New PR in worker-vllm: *${{ github.event.pull_request.title }}*",
"blocks": [
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": ":rocket: *New Pull Request — worker-vllm*\n*<${{ github.event.pull_request.html_url }}|${{ github.event.pull_request.title }}>*\nOpened by *${{ github.event.pull_request.user.login }}*"
}
},
{
"type": "context",
"elements": [
{
"type": "mrkdwn",
"text": "${{ github.event.pull_request.base.ref }} ← ${{ github.event.pull_request.head.ref }}"
}
]
}
]
}'
+73
View File
@@ -0,0 +1,73 @@
name: Monitor vLLM Releases
on:
schedule:
- cron: '0 0 * * *' # Every day at midnight
workflow_dispatch:
permissions:
contents: read
jobs:
check-vllm-release:
runs-on: ubuntu-latest
steps:
- name: Restore last known vLLM tag
uses: actions/cache/restore@v4
with:
path: .vllm-last-tag
key: vllm-tag-${{ github.run_id }}
restore-keys: vllm-tag-
- name: Get latest vLLM release
id: vllm
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
response=$(curl -sf https://api.github.com/repos/vllm-project/vllm/releases/latest \
-H "Authorization: Bearer $GH_TOKEN")
echo "tag=$(echo "$response" | jq -r '.tag_name')" >> $GITHUB_OUTPUT
echo "url=$(echo "$response" | jq -r '.html_url')" >> $GITHUB_OUTPUT
echo "name=$(echo "$response" | jq -r '.name')" >> $GITHUB_OUTPUT
- name: Check if new release
id: check
run: |
last=$(cat .vllm-last-tag 2>/dev/null || echo "")
current="${{ steps.vllm.outputs.tag }}"
echo "Last: $last Current: $current"
if [ -n "$current" ] && [ "$last" != "$current" ]; then
echo "is_new=true" >> $GITHUB_OUTPUT
else
echo "is_new=false" >> $GITHUB_OUTPUT
fi
- name: Notify Slack
if: steps.check.outputs.is_new == 'true'
run: |
curl -sf -X POST "${{ secrets.SLACK_WEBHOOK_URL }}" \
-H "Content-Type: application/json" \
-d '{
"text": ":rocket: New vLLM release: *${{ steps.vllm.outputs.tag }}*",
"blocks": [
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": ":rocket: *New vLLM Release: ${{ steps.vllm.outputs.tag }}*\n<${{ steps.vllm.outputs.url }}|View on GitHub>"
}
}
]
}'
- name: Save new tag
if: steps.check.outputs.is_new == 'true'
run: echo "${{ steps.vllm.outputs.tag }}" > .vllm-last-tag
- name: Update cache
if: steps.check.outputs.is_new == 'true'
uses: actions/cache/save@v4
with:
path: .vllm-last-tag
key: vllm-tag-${{ steps.vllm.outputs.tag }}
+37 -2
View File
@@ -1,4 +1,4 @@
![vLLM worker banner](https://cpjrphpz3t5wbwfe.public.blob.vercel-storage.com/worker-vllm_banner.jpeg) ![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png)
Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
@@ -6,6 +6,8 @@ Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
[![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm) [![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
Current vLLM version: [0.20.2](https://github.com/vllm-project/vllm/releases/tag/v0.20.2)
--- ---
## Endpoint Configuration ## Endpoint Configuration
@@ -27,11 +29,15 @@ All behaviour is controlled through environment variables:
| `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" | | `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String | | `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer | | `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
| `ENFORCE_EAGER` | If True, we will disable CUDA graph and always execute the model in eager mode. If False, we will use CUDA graph and eager execution in hybrid for maximal performance and flexibility. | true | boolean (true or false) |
**Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options. **Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options.
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md). For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
### Specify Transformers Version
To change the version of the [Transformers library](https://github.com/huggingface/transformers) use the `TRANSFORMERS_VERSION` environment variable to specify the version you want to use. Note this might break the handler, so use for development purposes.
## API Usage ## API Usage
This worker supports two API formats: **RunPod native** and **OpenAI-compatible**. This worker supports two API formats: **RunPod native** and **OpenAI-compatible**.
@@ -157,6 +163,35 @@ For external clients and SDKs, use the `/openai/v1` path prefix with your RunPod
{} {}
``` ```
#### OpenAI Responses API
**Path:** `/openai/v1/responses`
Supports the [OpenAI Responses API](https://platform.openai.com/docs/api-reference/responses) format. Note: this route bypasses the RunPod queue and is served directly — use `/openai/` prefixed paths rather than the RunPod job queue for these endpoints.
```json
{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"input": "Tell me a joke."
}
```
#### Anthropic Messages API
**Path:** `/openai/v1/messages`
Supports the [Anthropic Messages API](https://docs.anthropic.com/en/api/messages) format. Served directly, bypassing the RunPod queue.
```json
{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"max_tokens": 256,
"messages": [
{"role": "user", "content": "Hello!"}
]
}
```
#### Response Format #### Response Format
Both APIs return the same response format: Both APIs return the same response format:
@@ -190,7 +225,7 @@ Minimal Python example using the official `openai` SDK:
from openai import OpenAI from openai import OpenAI
import os import os
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL # Initialize the OpenAI Client with your Runpod API Key and Endpoint URL
client = OpenAI( client = OpenAI(
api_key=os.getenv("RUNPOD_API_KEY"), api_key=os.getenv("RUNPOD_API_KEY"),
base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1", base_url=f"https://api.runpod.ai/v2/<ENDPOINT_ID>/openai/v1",
+22 -2
View File
@@ -9,7 +9,7 @@
"containerDiskInGb": 150, "containerDiskInGb": 150,
"gpuIds": "ADA_80_PRO,AMPERE_80", "gpuIds": "ADA_80_PRO,AMPERE_80",
"gpuCount": 1, "gpuCount": 1,
"allowedCudaVersions": ["12.9", "12.8"], "allowedCudaVersions": ["13.0"],
"presets": [ "presets": [
{ {
"name": "deepseek-ai/deepseek-r1-distill-llama-8b", "name": "deepseek-ai/deepseek-r1-distill-llama-8b",
@@ -621,7 +621,7 @@
"name": "Enforce Eager", "name": "Enforce Eager",
"type": "boolean", "type": "boolean",
"description": "Always use eager-mode PyTorch. If False (0), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility", "description": "Always use eager-mode PyTorch. If False (0), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility",
"default": false, "default": true,
"advanced": true "advanced": true
} }
}, },
@@ -795,6 +795,26 @@
"default": "", "default": "",
"advanced": true "advanced": true
} }
},
{
"key": "PYTORCH_ALLOC_CONF",
"input": {
"name": "PyTorch Alloc Config",
"type": "string",
"description": "PyTorch allocation configuration, remove this if you want to use the default configuration",
"default": "expandable_segments:True",
"advanced": true
}
},
{
"key": "VLLM_USE_DEEP_GEMM",
"input": {
"name": "Use DeepGEMM",
"type": "string",
"description": "Enable DeepGEMM FP8 kernels (MoE and MQA logits). Set to 1 to enable, 0 to disable. Required for DeepSeek V4 models. Disabled by default — enable on H100/H200 for potential throughput gains. Some GPUs (e.g. H20) may perform better with this off.",
"default": "0",
"advanced": true
}
} }
] ]
} }
+18 -13
View File
@@ -1,20 +1,22 @@
FROM nvidia/cuda:12.9.1-base-ubuntu22.04 FROM nvidia/cuda:13.0.2-devel-ubuntu22.04
RUN apt-get update -y \ RUN apt-get update -y \
&& apt-get install -y python3-pip && apt-get install -y python3-pip curl git \
&& curl -LsSf https://astral.sh/uv/install.sh | sh
RUN ldconfig /usr/local/cuda-12.9/compat/ ENV PATH="/root/.local/bin:$PATH"
# Install vLLM with FlashInfer from the CUDA 12.9 wheel index.
RUN python3 -m pip install --upgrade pip && \
python3 -m pip install "vllm[flashinfer]==0.17.0" --extra-index-url https://download.pytorch.org/whl/cu129
RUN ldconfig /usr/local/cuda-13.0/compat/
# Install vLLM with FlashInfer - use CUDA 130 PyTorch wheels
RUN uv pip install --system "packaging>=24.2" && \
uv pip install --system "vllm[flashinfer]==0.20.2" && \
uv pip install --system git+https://github.com/deepseek-ai/DeepGEMM.git@714dd1a4a980f7937a74343d19a8eba4fe321480 --no-build-isolation
# Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts) # Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts)
COPY builder/requirements.txt /requirements.txt COPY builder/requirements.txt /requirements.txt
RUN --mount=type=cache,target=/root/.cache/pip \ RUN --mount=type=cache,target=/root/.cache/uv \
python3 -m pip install --upgrade -r /requirements.txt uv pip install --system -r /requirements.txt
# Setup for Option 2: Building the Image with the Model included # Setup for Option 2: Building the Image with the Model included
ARG MODEL_NAME="" ARG MODEL_NAME=""
@@ -41,17 +43,20 @@ ENV MODEL_NAME=$MODEL_NAME \
# Prevent rayon thread pool panic in containers where ulimit -u < nproc # Prevent rayon thread pool panic in containers where ulimit -u < nproc
# (tokenizers uses Rust's rayon which tries to spawn threads = CPU cores) # (tokenizers uses Rust's rayon which tries to spawn threads = CPU cores)
TOKENIZERS_PARALLELISM=false \ TOKENIZERS_PARALLELISM=false \
RAYON_NUM_THREADS=4 RAYON_NUM_THREADS=4 \
# Disable DeepGEMM MoE kernels by default; override with VLLM_USE_DEEP_GEMM=1 to enable
VLLM_USE_DEEP_GEMM=0
ENV PYTHONPATH="/:/vllm-workspace" ENV PYTHONPATH="/:/vllm-workspace"
RUN if [ "${VLLM_NIGHTLY}" = "true" ]; then \ RUN if [ "${VLLM_NIGHTLY}" = "true" ]; then \
pip install -U vllm --pre --index-url https://pypi.org/simple --extra-index-url https://wheels.vllm.ai/nightly && \ uv pip install --system -U vllm --pre --index-url https://pypi.org/simple --extra-index-url https://wheels.vllm.ai/nightly && \
apt-get update && apt-get install -y git && rm -rf /var/lib/apt/lists/* && \ apt-get update && apt-get install -y git && rm -rf /var/lib/apt/lists/* && \
pip install git+https://github.com/huggingface/transformers.git; \ uv pip install --system git+https://github.com/huggingface/transformers.git; \
fi fi
COPY src /src COPY src /src
RUN chmod +x /src/start.sh
RUN --mount=type=secret,id=HF_TOKEN,required=false \ RUN --mount=type=secret,id=HF_TOKEN,required=false \
if [ -f /run/secrets/HF_TOKEN ]; then \ if [ -f /run/secrets/HF_TOKEN ]; then \
export HF_TOKEN=$(cat /run/secrets/HF_TOKEN); \ export HF_TOKEN=$(cat /run/secrets/HF_TOKEN); \
@@ -61,4 +66,4 @@ RUN --mount=type=secret,id=HF_TOKEN,required=false \
fi fi
# Start the handler # Start the handler
CMD ["python3", "/src/handler.py"] CMD ["/bin/bash", "/src/start.sh"]
+86 -17
View File
@@ -2,10 +2,17 @@
# OpenAI-Compatible vLLM Serverless Endpoint Worker # OpenAI-Compatible vLLM Serverless Endpoint Worker
Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https://github.com/vllm-project/vllm) Inference Engine on RunPod Serverless with just a few clicks. Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https://github.com/vllm-project/vllm) Inference Engine on Runpod Serverless with just a few clicks.
</div> </div>
![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png)
Current vLLM version: [0.20.2](https://github.com/vllm-project/vllm/releases/tag/v0.20.2)
> Check out our Load Balancer implementation here: [vLLM Load Balancer](https://github.com/runpod-workers/vllm-loadbalancer-ep)
## Table of Contents ## Table of Contents
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker) - [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
@@ -21,9 +28,11 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
- [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker) - [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker)
- [OpenAI Request Input Parameters](#openai-request-input-parameters) - [OpenAI Request Input Parameters](#openai-request-input-parameters)
- [Chat Completions [RECOMMENDED]](#chat-completions-recommended) - [Chat Completions [RECOMMENDED]](#chat-completions-recommended)
- [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai) - [Examples: Using your Runpod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai)
- [Chat Completions](#chat-completions) - [Chat Completions](#chat-completions)
- [Getting a list of names for available models](#getting-a-list-of-names-for-available-models) - [Getting a list of names for available models](#getting-a-list-of-names-for-available-models)
- [OpenAI Responses API](#openai-responses-api)
- [Anthropic Messages API](#anthropic-messages-api)
- [Usage: Standard (Non-OpenAI)](#usage-standard-non-openai) - [Usage: Standard (Non-OpenAI)](#usage-standard-non-openai)
- [Request Input Parameters](#request-input-parameters) - [Request Input Parameters](#request-input-parameters)
- [Sampling Parameters](#sampling-parameters) - [Sampling Parameters](#sampling-parameters)
@@ -33,12 +42,12 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
## Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended] ## Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended]
**🚀 Deploy Guide**: Follow our [step-by-step deployment guide](https://docs.runpod.io/serverless/vllm/get-started) to deploy using the RunPod Console. **🚀 Deploy Guide**: Follow our [step-by-step deployment guide](https://docs.runpod.io/serverless/vllm/get-started) to deploy using the Runpod Console.
**📦 Docker Image**: `runpod/worker-v1-vllm:<version>` **📦 Docker Image**: `runpod/worker-v1-vllm:<version>`
- **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases) - **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases)
- **CUDA Compatibility**: Requires CUDA >= 12.1 - **CUDA Compatibility**: Requires CUDA >= 13.0
### Configuration ### Configuration
@@ -71,6 +80,10 @@ Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is a
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)** For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
### Specify Transformers Version
To change the version of the [Transformers library](https://github.com/huggingface/transformers) use the `TRANSFORMERS_VERSION` environment variable to specify the version you want to use. Note this might break the handler, so use for development purposes.
## Option 2: Build Docker Image with Model Inside ## Option 2: Build Docker Image with Model Inside
To build an image with the model baked in, you must specify the following docker arguments when building the image. To build an image with the model baked in, you must specify the following docker arguments when building the image.
@@ -142,13 +155,13 @@ You can deploy **any model on Hugging Face** that is supported by vLLM. For the
# Usage: OpenAI Compatibility # Usage: OpenAI Compatibility
The vLLM Worker is fully compatible with OpenAI's API, and you can use it with any OpenAI Codebase by changing only 3 lines in total. The supported routes are <ins>Chat Completions</ins> and <ins>Models</ins> - with both streaming and non-streaming. The vLLM Worker is fully compatible with OpenAI's API, and you can use it with any OpenAI Codebase by changing only 3 lines in total. The supported routes are <ins>Chat Completions</ins>, <ins>Models</ins>, <ins>Responses</ins>, and <ins>Messages</ins> - with both streaming and non-streaming.
## Modifying your OpenAI Codebase to use your deployed vLLM Worker ## Modifying your OpenAI Codebase to use your deployed vLLM Worker
**Python** (similar to Node.js, etc.): **Python** (similar to Node.js, etc.):
1. When initializing the OpenAI Client in your code, change the `api_key` to your RunPod API Key and the `base_url` to your RunPod Serverless Endpoint URL in the following format: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1`, filling in your deployed endpoint ID. For example, if your Endpoint ID is `abc1234`, the URL would be `https://api.runpod.ai/v2/abc1234/openai/v1`. 1. When initializing the OpenAI Client in your code, change the `api_key` to your Runpod API Key and the `base_url` to your Runpod Serverless Endpoint URL in the following format: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1`, filling in your deployed endpoint ID. For example, if your Endpoint ID is `abc1234`, the URL would be `https://api.runpod.ai/v2/abc1234/openai/v1`.
- Before: - Before:
@@ -174,7 +187,7 @@ The vLLM Worker is fully compatible with OpenAI's API, and you can use it with a
```python ```python
response = client.chat.completions.create( response = client.chat.completions.create(
model="gpt-3.5-turbo", model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}], messages=[{"role": "user", "content": "Why is Runpod the best platform?"}],
temperature=0, temperature=0,
max_tokens=100, max_tokens=100,
) )
@@ -183,7 +196,7 @@ The vLLM Worker is fully compatible with OpenAI's API, and you can use it with a
```python ```python
response = client.chat.completions.create( response = client.chat.completions.create(
model="<YOUR DEPLOYED MODEL REPO/NAME>", model="<YOUR DEPLOYED MODEL REPO/NAME>",
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}], messages=[{"role": "user", "content": "Why is Runpod the best platform?"}],
temperature=0, temperature=0,
max_tokens=100, max_tokens=100,
) )
@@ -191,7 +204,7 @@ The vLLM Worker is fully compatible with OpenAI's API, and you can use it with a
**Using http requests**: **Using http requests**:
1. Change the `Authorization` header to your RunPod API Key and the `url` to your RunPod Serverless Endpoint URL in the following format: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1` 1. Change the `Authorization` header to your Runpod API Key and the `url` to your Runpod Serverless Endpoint URL in the following format: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1`
- Before: - Before:
```bash ```bash
curl https://api.openai.com/v1/chat/completions \ curl https://api.openai.com/v1/chat/completions \
@@ -202,7 +215,7 @@ The vLLM Worker is fully compatible with OpenAI's API, and you can use it with a
"messages": [ "messages": [
{ {
"role": "user", "role": "user",
"content": "Why is RunPod the best platform?" "content": "Why is Runpod the best platform?"
} }
], ],
"temperature": 0, "temperature": 0,
@@ -219,7 +232,7 @@ The vLLM Worker is fully compatible with OpenAI's API, and you can use it with a
"messages": [ "messages": [
{ {
"role": "user", "role": "user",
"content": "Why is RunPod the best platform?" "content": "Why is Runpod the best platform?"
} }
], ],
"temperature": 0, "temperature": 0,
@@ -239,7 +252,7 @@ When using the chat completion feature of the vLLM Serverless Endpoint Worker, y
| Parameter | Type | Default Value | Description | | Parameter | Type | Default Value | Description |
| ------------------- | -------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | ------------------- | -------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `messages` | Union[str, List[Dict[str, str]]] | | List of messages, where each message is a dictionary with a `role` and `content`. The model's chat template will be applied to the messages automatically, so the model must have one or it should be specified as `CUSTOM_CHAT_TEMPLATE` env var. | | `messages` | Union[str, List[Dict[str, str]]] | | List of messages, where each message is a dictionary with a `role` and `content`. The model's chat template will be applied to the messages automatically, so the model must have one or it should be specified as `CUSTOM_CHAT_TEMPLATE` env var. |
| `model` | str | | The model repo that you've deployed on your RunPod Serverless Endpoint. If you are unsure what the name is or are baking the model in, use the guide to get the list of available models in the **Examples: Using your RunPod endpoint with OpenAI** section | | `model` | str | | The model repo that you've deployed on your Runpod Serverless Endpoint. If you are unsure what the name is or are baking the model in, use the guide to get the list of available models in the **Examples: Using your Runpod endpoint with OpenAI** section |
| `temperature` | Optional[float] | 0.7 | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. | | `temperature` | Optional[float] | 0.7 | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. |
| `top_p` | Optional[float] | 1.0 | Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. | | `top_p` | Optional[float] | 1.0 | Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. |
| `n` | Optional[int] | 1 | Number of output sequences to return for the given prompt. | | `n` | Optional[int] | 1 | Number of output sequences to return for the given prompt. |
@@ -269,15 +282,15 @@ Additional parameters supported by vLLM:
</details> </details>
### Examples: Using your RunPod endpoint with OpenAI ### Examples: Using your Runpod endpoint with OpenAI
First, initialize the OpenAI Client with your RunPod API Key and Endpoint URL: First, initialize the OpenAI Client with your Runpod API Key and Endpoint URL:
```python ```python
from openai import OpenAI from openai import OpenAI
import os import os
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL # Initialize the OpenAI Client with your Runpod API Key and Endpoint URL
client = OpenAI( client = OpenAI(
api_key=os.environ.get("RUNPOD_API_KEY"), api_key=os.environ.get("RUNPOD_API_KEY"),
base_url="https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1", base_url="https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1",
@@ -293,7 +306,7 @@ This is the format used for GPT-4 and focused on instruction-following and chat.
# Create a chat completion stream # Create a chat completion stream
response_stream = client.chat.completions.create( response_stream = client.chat.completions.create(
model="<YOUR DEPLOYED MODEL REPO/NAME>", model="<YOUR DEPLOYED MODEL REPO/NAME>",
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}], messages=[{"role": "user", "content": "Why is Runpod the best platform?"}],
temperature=0, temperature=0,
max_tokens=100, max_tokens=100,
stream=True, stream=True,
@@ -307,7 +320,7 @@ This is the format used for GPT-4 and focused on instruction-following and chat.
# Create a chat completion # Create a chat completion
response = client.chat.completions.create( response = client.chat.completions.create(
model="<YOUR DEPLOYED MODEL REPO/NAME>", model="<YOUR DEPLOYED MODEL REPO/NAME>",
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}], messages=[{"role": "user", "content": "Why is Runpod the best platform?"}],
temperature=0, temperature=0,
max_tokens=100, max_tokens=100,
) )
@@ -325,6 +338,62 @@ list_of_models = [model.id for model in models_response]
print(list_of_models) print(list_of_models)
``` ```
### OpenAI Responses API
**Path:** `/openai/v1/responses` (full URL: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1/responses`)
Supports the [OpenAI Responses API](https://platform.openai.com/docs/api-reference/responses) request shape. Like other `/openai/` routes, this is served directly—use the `/openai/` prefix rather than the RunPod native job queue for these calls.
```json
{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"input": "Tell me a joke."
}
```
**Using HTTP requests:**
```bash
curl https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <YOUR RUNPOD API KEY>" \
-d '{
"model": "<YOUR DEPLOYED MODEL REPO/NAME>",
"input": "Tell me a joke."
}'
```
### Anthropic Messages API
**Path:** `/openai/v1/messages` (full URL: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1/messages`)
Supports the [Anthropic Messages API](https://docs.anthropic.com/en/api/messages) format. Served directly, bypassing the RunPod queue.
```json
{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"max_tokens": 256,
"messages": [
{"role": "user", "content": "Hello!"}
]
}
```
**Using HTTP requests:**
```bash
curl https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1/messages \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <YOUR RUNPOD API KEY>" \
-d '{
"model": "<YOUR DEPLOYED MODEL REPO/NAME>",
"max_tokens": 256,
"messages": [
{"role": "user", "content": "Hello!"}
]
}'
```
# Usage: Standard (Non-OpenAI) # Usage: Standard (Non-OpenAI)
## Request Input Parameters ## Request Input Parameters
+4 -3
View File
@@ -1,14 +1,15 @@
ray ray
pandas pandas
pyarrow pyarrow
runpod runpod==1.9.0
huggingface-hub huggingface-hub
packaging lmcache==0.4.5
packaging>=24.2
typing-extensions>=4.8.0 typing-extensions>=4.8.0
pydantic pydantic
pydantic-settings pydantic-settings
hf-transfer hf-transfer
transformers>=4.57.0 transformers>=5
bitsandbytes>=0.45.0 bitsandbytes>=0.45.0
kernels kernels
torch-c-dlpack-ext torch-c-dlpack-ext
+3
View File
@@ -97,10 +97,13 @@ If `SPECULATIVE_CONFIG` is set, it takes priority over individual env vars. When
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. | | `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. | | `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models. | | `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models. |
| `VLLM_USE_DEEP_GEMM` | `0` | `str` (`0`/`1`) | Enable DeepGEMM FP8 kernels for MoE and MQA logits computation. Disabled by default. Must be `"0"` or `"1"` — not `true`/`false`. See note below. |
| `ATTENTION_BACKEND` | `None` | `str` | Attention backend to use (e.g., `FLASH_ATTN`, `FLASHINFER`, `TRITON_FLASH_ATTN`). Replaces deprecated `VLLM_ATTENTION_BACKEND`. | | `ATTENTION_BACKEND` | `None` | `str` | Attention backend to use (e.g., `FLASH_ATTN`, `FLASHINFER`, `TRITON_FLASH_ATTN`). Replaces deprecated `VLLM_ATTENTION_BACKEND`. |
| `ASYNC_SCHEDULING` | `None` | `bool` | Enable async scheduling (overlaps engine scheduling with GPU execution). Default: enabled in vLLM 0.14.0+. Set to `false` to disable. | | `ASYNC_SCHEDULING` | `None` | `bool` | Enable async scheduling (overlaps engine scheduling with GPU execution). Default: enabled in vLLM 0.14.0+. Set to `false` to disable. |
| `STREAM_INTERVAL` | `1` | `int` | Controls how often to yield streaming results. Lower = more frequent updates. | | `STREAM_INTERVAL` | `1` | `int` | Controls how often to yield streaming results. Lower = more frequent updates. |
> **Note (`VLLM_USE_DEEP_GEMM`):** DeepGEMM is used in two places: MoE weight computation and MQA logits computation. It is necessary for MQA logits computation on supported hardware — required for DeepSeek V4 models. Set `VLLM_USE_DEEP_GEMM=1` to enable. Set `VLLM_USE_DEEP_GEMM=0` to disable the MoE part and fall back to flashinfer/cutlass FP8 kernels. **Value must be `"0"` or `"1"` — not `"true"`/`"false"`.** Some users report better performance with `VLLM_USE_DEEP_GEMM=0`, particularly on H20 GPUs. Disabling it also skips the DeepGEMM warmup phase, reducing cold-start time. Requires CUDA 13.0+ and SM90+ (H100/H200) to use; the library is installed but inactive by default.
## Tokenizer Settings ## Tokenizer Settings
| Variable | Default | Type/Choices | Description | | Variable | Default | Type/Choices | Description |
+253 -26
View File
@@ -1,4 +1,5 @@
import asyncio import asyncio
import inspect
import json import json
import logging import logging
import os import os
@@ -7,7 +8,10 @@ from typing import AsyncGenerator, Optional
from dotenv import load_dotenv from dotenv import load_dotenv
from vllm import AsyncLLMEngine from vllm import AsyncLLMEngine
from vllm.inputs import TextPrompt
from vllm.entrypoints.logger import RequestLogger from vllm.entrypoints.logger import RequestLogger
from vllm.entrypoints.anthropic.protocol import AnthropicMessagesRequest, AnthropicMessagesResponse, AnthropicError, AnthropicErrorResponse
from vllm.entrypoints.anthropic.serving import AnthropicServingMessages
from vllm.entrypoints.openai.chat_completion.protocol import ChatCompletionRequest from vllm.entrypoints.openai.chat_completion.protocol import ChatCompletionRequest
from vllm.entrypoints.openai.chat_completion.serving import OpenAIServingChat from vllm.entrypoints.openai.chat_completion.serving import OpenAIServingChat
from vllm.entrypoints.openai.completion.protocol import CompletionRequest from vllm.entrypoints.openai.completion.protocol import CompletionRequest
@@ -15,6 +19,9 @@ from vllm.entrypoints.openai.completion.serving import OpenAIServingCompletion
from vllm.entrypoints.openai.engine.protocol import ErrorResponse from vllm.entrypoints.openai.engine.protocol import ErrorResponse
from vllm.entrypoints.openai.models.protocol import BaseModelPath, LoRAModulePath from vllm.entrypoints.openai.models.protocol import BaseModelPath, LoRAModulePath
from vllm.entrypoints.openai.models.serving import OpenAIServingModels from vllm.entrypoints.openai.models.serving import OpenAIServingModels
from vllm.entrypoints.openai.responses.protocol import ResponsesRequest, ResponsesResponse
from vllm.entrypoints.openai.responses.serving import OpenAIServingResponses
from vllm.entrypoints.serve.render.serving import OpenAIServingRender
from constants import DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MAX_CONCURRENCY, DEFAULT_MIN_BATCH_SIZE from constants import DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MAX_CONCURRENCY, DEFAULT_MIN_BATCH_SIZE
from engine_args import get_engine_args from engine_args import get_engine_args
@@ -25,20 +32,33 @@ class vLLMEngine:
def __init__(self, engine = None): def __init__(self, engine = None):
load_dotenv() # For local development load_dotenv() # For local development
self.engine_args = get_engine_args() self.engine_args = get_engine_args()
logging.info(f"Engine args: {self.engine_args}")
# Initialize vLLM engine first if engine is None:
self.llm = self._initialize_llm() if engine is None else engine.llm ea = self.engine_args
summary = {
"model": ea.model,
"dtype": ea.dtype,
"quantization": ea.quantization,
"max_model_len": ea.max_model_len,
"tensor_parallel_size": ea.tensor_parallel_size,
"gpu_memory_utilization": ea.gpu_memory_utilization,
}
if ea.tokenizer and ea.tokenizer != ea.model:
summary["tokenizer"] = ea.tokenizer
logging.info("Engine config: %s", summary)
logging.debug("Full engine args: %s", ea)
# Only create custom tokenizer wrapper if not using mistral tokenizer mode self.llm = self._initialize_llm()
# For mistral models, let vLLM handle tokenizer initialization
if self.engine_args.tokenizer_mode != 'mistral': if self.engine_args.tokenizer_mode != 'mistral':
self.tokenizer = TokenizerWrapper(self.engine_args.tokenizer or self.engine_args.model, self.tokenizer = TokenizerWrapper(self.engine_args.tokenizer or self.engine_args.model,
self.engine_args.tokenizer_revision, self.engine_args.tokenizer_revision,
self.engine_args.trust_remote_code) self.engine_args.trust_remote_code)
else:
self.tokenizer = None
else: else:
# For mistral models, we'll get the tokenizer from vLLM later self.llm = engine.llm
self.tokenizer = None self.tokenizer = engine.tokenizer
self.max_concurrency = int(os.getenv("MAX_CONCURRENCY", DEFAULT_MAX_CONCURRENCY)) self.max_concurrency = int(os.getenv("MAX_CONCURRENCY", DEFAULT_MAX_CONCURRENCY))
self.default_batch_size = int(os.getenv("DEFAULT_BATCH_SIZE", DEFAULT_BATCH_SIZE)) self.default_batch_size = int(os.getenv("DEFAULT_BATCH_SIZE", DEFAULT_BATCH_SIZE))
@@ -111,7 +131,7 @@ class vLLMEngine:
if apply_chat_template or isinstance(llm_input, list): if apply_chat_template or isinstance(llm_input, list):
tokenizer_wrapper = self._get_tokenizer_for_chat_template() tokenizer_wrapper = self._get_tokenizer_for_chat_template()
llm_input = tokenizer_wrapper.apply_chat_template(llm_input) llm_input = tokenizer_wrapper.apply_chat_template(llm_input)
results_generator = self.llm.generate(llm_input, validated_sampling_params, request_id) results_generator = self.llm.generate(TextPrompt(prompt=llm_input), validated_sampling_params, request_id)
n_responses, n_input_tokens, is_first_output = validated_sampling_params.n, 0, True n_responses, n_input_tokens, is_first_output = validated_sampling_params.n, 0, True
last_output_texts, token_counters = ["" for _ in range(n_responses)], {"batch": 0, "total": 0} last_output_texts, token_counters = ["" for _ in range(n_responses)], {"batch": 0, "total": 0}
@@ -201,19 +221,48 @@ class OpenAIvLLMEngine(vLLMEngine):
self.raw_openai_output = bool(int(raw_output_env)) self.raw_openai_output = bool(int(raw_output_env))
def _load_lora_adapters(self): def _load_lora_adapters(self):
adapters = [] lora_modules_env = os.getenv("LORA_MODULES", "")
try: if not lora_modules_env:
adapters = json.loads(os.getenv("LORA_MODULES", '[]')) return []
except Exception as e:
logging.info(f"---Initialized adapter json load error: {e}")
for i, adapter in enumerate(adapters): try:
parsed = json.loads(lora_modules_env)
except json.JSONDecodeError as e:
logging.error(
"LORA_MODULES could not be parsed as JSON: %s — no LoRA adapters loaded. Value: %r",
e, lora_modules_env,
)
return []
# Accept a single adapter dict as well as an array
if isinstance(parsed, dict):
parsed = [parsed]
if not isinstance(parsed, list):
logging.error(
"LORA_MODULES must be a JSON array of adapter objects, got %s — no LoRA adapters loaded.",
type(parsed).__name__,
)
return []
adapters = []
for i, adapter in enumerate(parsed):
try: try:
adapters[i] = LoRAModulePath(**adapter) adapters.append(LoRAModulePath(**adapter))
logging.info(f"---Initialized adapter: {adapter}") logging.info("Loaded LoRA adapter config [%d]: %s", i, adapter)
except Exception as e: except Exception as e:
logging.info(f"---Initialized adapter not worked: {e}") logging.error(
continue "Failed to parse LoRA adapter at index %d: %s. Config: %r",
i, e, adapter,
)
if parsed and not adapters:
logging.error(
"LORA_MODULES specified %d adapter(s) but none could be loaded — "
"OpenAI model name lookups for LoRA adapters will fail.",
len(parsed),
)
return adapters return adapters
async def _ensure_engines_initialized(self): async def _ensure_engines_initialized(self):
@@ -248,10 +297,26 @@ class OpenAIvLLMEngine(vLLMEngine):
if self.tokenizer and hasattr(self.tokenizer, 'tokenizer'): if self.tokenizer and hasattr(self.tokenizer, 'tokenizer'):
chat_template = self.tokenizer.tokenizer.chat_template chat_template = self.tokenizer.tokenizer.chat_template
self.openai_serving_render = OpenAIServingRender(
model_config=self.llm.model_config,
renderer=self.llm.renderer,
model_registry=self.serving_models.registry,
request_logger=None,
chat_template=chat_template,
chat_template_content_format="auto",
trust_request_chat_template=os.getenv('TRUST_REQUEST_CHAT_TEMPLATE', 'false').lower() == 'true',
enable_auto_tools=os.getenv('ENABLE_AUTO_TOOL_CHOICE', 'false').lower() == 'true',
exclude_tools_when_tool_choice_none=os.getenv('EXCLUDE_TOOLS_WHEN_TOOL_CHOICE_NONE', 'false').lower() == 'true',
tool_parser=os.getenv('TOOL_CALL_PARSER', "") or None,
reasoning_parser=os.getenv('REASONING_PARSER', "") or None,
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
)
self.chat_engine = OpenAIServingChat( self.chat_engine = OpenAIServingChat(
engine_client=self.llm, engine_client=self.llm,
models=self.serving_models, models=self.serving_models,
response_role=self.response_role, response_role=self.response_role,
openai_serving_render=self.openai_serving_render,
request_logger=None, request_logger=None,
chat_template=chat_template, chat_template=chat_template,
chat_template_content_format="auto", chat_template_content_format="auto",
@@ -264,20 +329,53 @@ class OpenAIvLLMEngine(vLLMEngine):
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true', enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true', enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
enable_log_outputs=os.getenv('ENABLE_LOG_OUTPUTS', 'false').lower() == 'true', enable_log_outputs=os.getenv('ENABLE_LOG_OUTPUTS', 'false').lower() == 'true',
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
) )
self.completion_engine = OpenAIServingCompletion( self.completion_engine = OpenAIServingCompletion(
engine_client=self.llm, engine_client=self.llm,
models=self.serving_models, models=self.serving_models,
openai_serving_render=self.openai_serving_render,
request_logger=None, request_logger=None,
return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true', return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true', enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true', enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true', )
self.responses_engine = OpenAIServingResponses(
engine_client=self.llm,
models=self.serving_models,
openai_serving_render=self.openai_serving_render,
request_logger=None,
chat_template=chat_template,
chat_template_content_format="auto",
return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
reasoning_parser=os.getenv('REASONING_PARSER', "") or "",
enable_auto_tools=os.getenv('ENABLE_AUTO_TOOL_CHOICE', 'false').lower() == 'true',
tool_parser=os.getenv('TOOL_CALL_PARSER', "") or None,
tool_server=None,
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
enable_log_outputs=os.getenv('ENABLE_LOG_OUTPUTS', 'false').lower() == 'true',
)
self.messages_engine = AnthropicServingMessages(
engine_client=self.llm,
models=self.serving_models,
response_role=self.response_role,
openai_serving_render=self.openai_serving_render,
request_logger=None,
chat_template=chat_template,
chat_template_content_format="auto",
return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
reasoning_parser=os.getenv('REASONING_PARSER', "") or "",
enable_auto_tools=os.getenv('ENABLE_AUTO_TOOL_CHOICE', 'false').lower() == 'true',
tool_parser=os.getenv('TOOL_CALL_PARSER', "") or None,
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
) )
if hasattr(self.chat_engine, 'warmup'): warmup = getattr(self.chat_engine, 'warmup', None)
await self.chat_engine.warmup() if callable(warmup):
result = warmup()
if inspect.isawaitable(result):
await result
async def generate(self, openai_request: JobInput): async def generate(self, openai_request: JobInput):
# Ensure engines are ready (no-op if already initialized at startup) # Ensure engines are ready (no-op if already initialized at startup)
@@ -288,6 +386,12 @@ class OpenAIvLLMEngine(vLLMEngine):
elif openai_request.openai_route in ["/v1/chat/completions", "/v1/completions"]: elif openai_request.openai_route in ["/v1/chat/completions", "/v1/completions"]:
async for response in self._handle_chat_or_completion_request(openai_request): async for response in self._handle_chat_or_completion_request(openai_request):
yield response yield response
elif openai_request.openai_route == "/v1/responses":
async for response in self._handle_responses_request(openai_request):
yield response
elif openai_request.openai_route == "/v1/messages":
async for response in self._handle_messages_request(openai_request):
yield response
else: else:
yield create_error_response("Invalid route").model_dump() yield create_error_response("Invalid route").model_dump()
@@ -343,3 +447,126 @@ class OpenAIvLLMEngine(vLLMEngine):
batch = "".join(batch) batch = "".join(batch)
yield batch yield batch
async def _handle_responses_request(self, openai_request: JobInput):
request_id = getattr(openai_request, "request_id", "unknown")
try:
request = ResponsesRequest(**openai_request.openai_input)
except Exception as e:
logging.error(
"Invalid ResponsesRequest JSON: %s",
e,
extra={"request_id": request_id}
)
yield create_error_response(
"Invalid request format",
err_type="BadRequestError"
).model_dump()
return
dummy_request = DummyRequest()
try:
response = await self.responses_engine.create_responses(request, raw_request=dummy_request)
except Exception as e:
logging.error(
"Failed to create Responses: %s",
e,
extra={"request_id": request_id},
exc_info=True
)
yield create_error_response(
"Internal server error during response generation",
err_type="InternalServerError"
).model_dump()
return
if isinstance(response, (ErrorResponse, ResponsesResponse)):
yield response.model_dump()
return
try:
async for event in response:
if not hasattr(event, "type"):
continue
event_type = getattr(event, "type", "unknown")
yield f"event: {event_type}\ndata: {event.model_dump_json(indent=None)}\n\n"
except Exception as e:
logging.error(
"Error processing responses stream: %s",
e,
extra={"request_id": request_id},
exc_info=True
)
error_payload = create_error_response(
"Streaming response failed",
err_type="InternalServerError"
).model_dump_json()
yield f"event: error\ndata: {error_payload}\n\n"
async def _handle_messages_request(self, openai_request: JobInput):
request_id = getattr(openai_request, "request_id", "unknown")
try:
request = AnthropicMessagesRequest(**openai_request.openai_input)
except Exception as e:
logging.error(
"Invalid AnthropicMessagesRequest: %s",
e,
extra={"request_id": request_id}
)
yield AnthropicErrorResponse(
error=AnthropicError(
type="invalid_request_error",
message="Invalid request format"
)
).model_dump()
return
dummy_request = DummyRequest()
try:
response = await self.messages_engine.create_messages(request, raw_request=dummy_request)
except Exception as e:
logging.error(
"Failed to create messages: %s",
e,
extra={"request_id": request_id},
exc_info=True
)
yield AnthropicErrorResponse(
error=AnthropicError(
type="internal_error",
message="Failed to generate messages"
)
).model_dump()
return
if isinstance(response, ErrorResponse):
error_type = getattr(response, "type", "internal_error")
error_message = getattr(response, "message", "Unknown error")
yield AnthropicErrorResponse(
error=AnthropicError(type=error_type, message=error_message)
).model_dump()
return
if isinstance(response, AnthropicMessagesResponse):
yield response.model_dump(exclude_none=True)
return
try:
async for chunk in response:
yield chunk
except Exception as e:
logging.error(
"Error streaming messages: %s",
e,
extra={"request_id": request_id},
exc_info=True
)
error_payload = AnthropicErrorResponse(
error=AnthropicError(
type="internal_error",
message="Error while streaming messages"
)
).model_dump_json()
yield f"event: error\ndata: {error_payload}\n\n"
+125
View File
@@ -1,3 +1,4 @@
import ast
import os import os
import json import json
import logging import logging
@@ -81,6 +82,16 @@ def _convert_env_value_to_field_type(value: str, field_name: str, field_type: ty
if type(None) in (args or ()): if type(None) in (args or ()):
return None return None
raise ValueError("empty value not allowed for non-optional field") raise ValueError("empty value not allowed for non-optional field")
# Union[bool, str, ...]: only coerce to bool for unambiguous literals;
# otherwise preserve the string (e.g. hf_token="hf_abc..." must stay a str).
if get_origin(field_type) is not None:
union_types = [a for a in (get_args(field_type) or ()) if a is not type(None)]
if bool in union_types and str in union_types:
if str(val).lower() in ("true", "false", "1", "0", "yes", "no", "on", "off"):
return str(val).lower() in ("true", "1", "yes", "on")
return str(val)
effective_type = _resolve_field_type(field_type) effective_type = _resolve_field_type(field_type)
# bool # bool
if effective_type is bool: if effective_type is bool:
@@ -113,6 +124,17 @@ def _convert_env_value_to_field_type(value: str, field_name: str, field_type: ty
except (json.JSONDecodeError, TypeError): except (json.JSONDecodeError, TypeError):
pass pass
return tuple(elem_type(x.strip()) for x in str(val).split(",") if x.strip()) return tuple(elem_type(x.strip()) for x in str(val).split(",") if x.strip())
# For dataclass/complex types, try JSON then Python literal parsing to dict
try:
return json.loads(val)
except (json.JSONDecodeError, TypeError):
pass
try:
parsed = ast.literal_eval(val)
if isinstance(parsed, (dict, list)):
return parsed
except (ValueError, SyntaxError):
pass
# Fallback: try int, float, then str # Fallback: try int, float, then str
try: try:
return int(val) return int(val)
@@ -330,6 +352,58 @@ def _sanitize_hf_overrides(hf_overrides: dict) -> dict | None:
return result or None return result or None
def _resolve_cached_model_path(model_name: str) -> str:
"""Return a local snapshot path when the HF cache was stored with lowercase names.
Some model stores (e.g. RunPod pre-cached volumes) normalize repo IDs to
lowercase. HuggingFace Hub stores caches as
``models--{org}--{model}/snapshots/{hash}/`` preserving the original casing,
so MODEL_NAME=Qwen/Qwen2.5-Coder-32B-Instruct-AWQ will miss a cache stored
as ``models--qwen--qwen2.5-coder-32b-instruct-awq/``.
If the exact-case cache directory is absent but a lowercase variant exists,
the latest snapshot path is returned so vLLM loads from disk rather than
attempting a redundant download.
"""
if os.path.isabs(model_name):
return model_name
cache_dir = (
os.getenv("HUGGINGFACE_HUB_CACHE")
or os.getenv("HF_HOME")
or os.path.expanduser("~/.cache/huggingface/hub")
)
folder_name = f"models--{model_name.replace('/', '--')}"
if os.path.isdir(os.path.join(cache_dir, folder_name)):
return model_name
lower_dir = os.path.join(cache_dir, folder_name.lower())
if not os.path.isdir(lower_dir):
return model_name
snapshots_dir = os.path.join(lower_dir, "snapshots")
if not os.path.isdir(snapshots_dir):
return model_name
try:
snapshots = sorted(os.listdir(snapshots_dir))
except OSError:
return model_name
if not snapshots:
return model_name
resolved = os.path.join(snapshots_dir, snapshots[-1])
logging.info(
"MODEL_NAME %r not found at original casing in HF cache; "
"resolved to lowercase cached snapshot at %r",
model_name, resolved,
)
return resolved
def get_local_args(): def get_local_args():
""" """
Retrieve local arguments from a JSON file. Retrieve local arguments from a JSON file.
@@ -401,6 +475,53 @@ def get_engine_args():
if os.getenv("MAX_PARALLEL_LOADING_WORKERS"): if os.getenv("MAX_PARALLEL_LOADING_WORKERS"):
logging.warning("Overriding MAX_PARALLEL_LOADING_WORKERS with None because more than 1 GPU is available.") logging.warning("Overriding MAX_PARALLEL_LOADING_WORKERS with None because more than 1 GPU is available.")
# LMCache requires HMA to be disabled
try:
_kv_transfer = args.get("kv_transfer_config")
if isinstance(_kv_transfer, str):
parsed = None
try:
parsed = json.loads(_kv_transfer)
except (json.JSONDecodeError, TypeError):
pass
if parsed is None:
try:
result = ast.literal_eval(_kv_transfer)
if isinstance(result, dict):
parsed = result
except (ValueError, SyntaxError):
pass
if parsed is not None:
_kv_transfer = parsed
args["kv_transfer_config"] = _kv_transfer
_kv_offload = args.get("kv_offloading_backend")
lmcache_via_offload = _kv_offload == "lmcache"
lmcache_via_transfer = (
isinstance(_kv_transfer, dict)
and isinstance(_kv_transfer.get("kv_connector"), str)
and "lmcache" in _kv_transfer.get("kv_connector", "").lower()
)
lmcache_detected = lmcache_via_offload or lmcache_via_transfer
if lmcache_detected:
current = args.get("disable_hybrid_kv_cache_manager")
if current is False:
logging.warning(
"disable_hybrid_kv_cache_manager=False conflicts with LMCache; "
"overriding to True (HMA must be disabled when using LMCache)"
)
args["disable_hybrid_kv_cache_manager"] = True
elif current is None:
args["disable_hybrid_kv_cache_manager"] = True
logging.info("LMCache detected: automatically setting disable_hybrid_kv_cache_manager=True")
except Exception as e:
logging.error(
"Failed to check LMCache configuration: %s",
e,
exc_info=True
)
# Deprecated env args backwards compatibility # Deprecated env args backwards compatibility
if args.get("kv_cache_dtype") == "fp8_e5m2": if args.get("kv_cache_dtype") == "fp8_e5m2":
args["kv_cache_dtype"] = "fp8" args["kv_cache_dtype"] = "fp8"
@@ -458,4 +579,8 @@ def get_engine_args():
if speculative_config: if speculative_config:
args["speculative_config"] = speculative_config args["speculative_config"] = speculative_config
# Resolve lowercase HF cache paths (FDE-174)
if args.get("model"):
args["model"] = _resolve_cached_model_path(args["model"])
return AsyncEngineArgs(**args) return AsyncEngineArgs(**args)
+9
View File
@@ -0,0 +1,9 @@
#!/bin/bash
set -e
if [ -n "${TRANSFORMERS_VERSION}" ]; then
echo "Installing transformers==${TRANSFORMERS_VERSION}"
uv pip install --system "transformers==${TRANSFORMERS_VERSION}"
fi
exec python3 /src/handler.py
+4 -2
View File
@@ -1,10 +1,12 @@
from transformers import AutoTokenizer import logging
import os import os
from typing import Union from typing import Union
from transformers import AutoTokenizer
class TokenizerWrapper: class TokenizerWrapper:
def __init__(self, tokenizer_name_or_path, tokenizer_revision, trust_remote_code): def __init__(self, tokenizer_name_or_path, tokenizer_revision, trust_remote_code):
print(f"tokenizer_name_or_path: {tokenizer_name_or_path}, tokenizer_revision: {tokenizer_revision}, trust_remote_code: {trust_remote_code}") logging.debug("tokenizer_name_or_path: %s, tokenizer_revision: %s, trust_remote_code: %s", tokenizer_name_or_path, tokenizer_revision, trust_remote_code)
self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name_or_path, revision=tokenizer_revision or "main", trust_remote_code=trust_remote_code) self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name_or_path, revision=tokenizer_revision or "main", trust_remote_code=trust_remote_code)
self.custom_chat_template = os.getenv("CUSTOM_CHAT_TEMPLATE") self.custom_chat_template = os.getenv("CUSTOM_CHAT_TEMPLATE")
self.has_chat_template = bool(self.tokenizer.chat_template) or bool(self.custom_chat_template) self.has_chat_template = bool(self.tokenizer.chat_template) or bool(self.custom_chat_template)