Compare commits

...
79 Commits
Author SHA1 Message Date
chrisvelaandGitHub 9e1c483136 Merge pull request #311 from runpod-workers/t3code/cd56c69d
Release / release (push) Waiting to run
fix: serve original model name when HF cache dir is lowercased (#310)
2026-06-26 13:48:13 -05:00
chrisvelaandGitHub d71ea9939d Merge pull request #308 from runpod-workers/docs/sync-vllm-version-0.20.2
docs: sync vLLM version to 0.20.2 in READMEs
2026-06-26 13:47:59 -05:00
chrisvelaandGitHub 7e2b4e2288 Merge pull request #314 from runpod-workers/runpod-package-update
chore: update runpod to 1.10.0
2026-06-26 13:45:16 -05:00
deanqandgithub-actions[bot] 015f8f3c4d chore: update runpod to 1.10.0 2026-06-26 17:20:41 +00:00
Hailong YangandGitHub 75ffcf73f2 Merge pull request #309 from adithyaJRunpod/feature/tuned-configs-round2
added configs for  Gemma 4 31B and GPT-OSS 120B
2026-06-22 17:41:50 -04:00
Tim Pietrusky d7ba3b6ab7 test: install pyyaml and isolate vllm config file in tests
after rebasing onto main, get_engine_args() loads a vllm-style config via
PyYAML (a transitive vllm dep). vllm is stubbed in tests, so add pyyaml
explicitly and point VLLM_CONFIG_FILE at a nonexistent path so no stray
config is picked up.
2026-06-19 19:05:24 +02:00
Tim Pietrusky b11c91722c test: complete vllm stub so src.utils imports under py<3.14
src.utils uses ErrorResponse (a vllm import) as a module-level return
annotation, evaluated eagerly on python <3.14. the vllm stub lacked it,
so collection failed with NameError on ci (py3.11) while passing locally
(py3.14, lazy annotations). add the missing vllm.utils / protocol /
SamplingParams symbols to the stub.
2026-06-19 19:04:45 +02:00
Tim Pietrusky fcdc799e0d fix: serve original model name when hf cache dir is lowercased (#310)
the fde-174 cache resolver rewrites engine_args.model to an on-disk
snapshot path when the model is found only under a lowercased hf cache
dir. the openai served model name is derived from engine_args.model, so
it silently became the filesystem path and requests using the real repo
id returned 404.

set served_model_name to the original repo id whenever the model is
rewritten to a path, unless an explicit served name (or
OPENAI_SERVED_MODEL_NAME_OVERRIDE) is provided.

add the first python tests in the repo (tests/) covering the cache-path
resolution and served-name decoupling, plus a Tests github workflow that
runs pytest on prs and pushes to main. vllm/torch are stubbed when absent
so the suite runs on a plain cpu runner.
2026-06-19 19:04:45 +02:00
AdithyaJob 84ec446493 added configs for Gemma 4 31B and GPT-OSS 120B 2026-06-17 22:19:58 -07:00
velaraptor-runpodandgithub-actions[bot] db246653a2 docs: sync vLLM version to 0.20.2 in READMEs 2026-06-12 20:51:06 +00:00
chrisvelaandGitHub 1b3228a2dc Merge pull request #307 from runpod-workers/fix/revert-0.20.0
Release / release (push) Waiting to run
revert to 0.20.2
2026-06-12 15:50:52 -05:00
velaraptor-runpod 4817d4a8e7 revert to 0.20.2 2026-06-12 15:49:23 -05:00
chrisvelaandGitHub 0378382a92 Merge pull request #306 from runpod-workers/revert/v2.20.1
Release / release (push) Waiting to run
chore: carry non-vllm changes from main (tests GPU + configs)
2026-06-12 15:21:22 -05:00
velaraptor-runpodandClaude Sonnet 4.6 08580e7ccf chore: carry non-vllm changes from main (tests GPU + configs)
Brings forward the L40 GPU type in tests.json and the new llama/qwen
tuned config files, while keeping Dockerfile pinned at vllm 0.20.2
(v2.20.1 state).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-12 15:18:36 -05:00
chrisvelaandGitHub 8868aae6b1 Merge pull request #305 from runpod-workers/fix/typo-tests-l40
Release / release (push) Waiting to run
fix: fix typo in tests
2026-06-12 14:00:54 -05:00
velaraptor-runpod fb8adc5c06 fix: fix typo in tests 2026-06-12 13:59:29 -05:00
chrisvelaandGitHub 7351da512b Merge pull request #304 from runpod-workers/fix/fix-tests
Release / release (push) Waiting to run
fix: change test gpu to L40
2026-06-12 13:49:15 -05:00
velaraptor-runpod 1e78043b2c fix: change test gpu to L40 2026-06-12 13:46:33 -05:00
chrisvelaandGitHub 352c64f4c1 Merge pull request #303 from runpod-workers/chore/check-vllm-versions-readmes
chore: check readmes vllm version in sync with version in Dockerfile
2026-06-12 09:49:07 -05:00
velaraptor-runpod 0922f5b435 chore: check readmes vllm version in sync with version in Dockerfile 2026-06-11 17:33:16 -05:00
chrisvelaandGitHub 5d9a48fc70 Merge pull request #302 from runpod-workers/feat/0.22.1
Release / release (push) Waiting to run
feat: upgrade vllm to 0.22.1
2026-06-11 17:27:38 -05:00
velaraptor-runpod 5d1579e361 fix: update readme links 2026-06-11 17:00:31 -05:00
velaraptor-runpod 0488b77d89 feat: upgrade vllm to 0.22.1 2026-06-11 16:29:02 -05:00
chrisvelaandGitHub d3a962c33b Merge pull request #301 from adithyaJRunpod/feature/tuned-configs
Release / release (push) Waiting to run
Add tuned configs,  CON-239
2026-06-11 14:26:00 -05:00
chrisvelaandGitHub 105c125698 Merge pull request #300 from runpod-workers/feat/0.21.0
feat: upgrade vllm to 0.21.0
2026-06-11 12:47:08 -05:00
AdithyaJob 0a0ccfcb60 Add tuned configs for Llama 3.1 8B and Qwen3 8B 2026-06-10 21:07:41 -07:00
velaraptor-runpod c8ce53c72c fix: add kenels, and fix for cuda 2026-06-10 16:01:53 -05:00
velaraptor-runpod 9618e799ba chore: fix cuda libraries 2026-06-04 15:35:58 -05:00
velaraptor-runpod cb3f077dba feat: upgrade vllm to 0.21.0 2026-06-03 14:43:56 -05:00
chrisvelaandGitHub 69646b9e99 Merge pull request #294 from runpod-workers/feat/allow-config
Release / release (push) Waiting to run
feat: allow config.yaml like vllm serve
2026-06-02 20:28:02 -05:00
Jacob CiparandGitHub 8b991a7ad7 Merge pull request #297 from runpod-workers/jhcipar/bump-runpod-python-version
feat: bump runpod-python version
2026-06-02 11:17:48 -04:00
Tim PietruskyandGitHub dac05b62b3 fix: make .runpod/tests.json hub tests pass on CUDA 13.0 (#299)
Release / release (push) Waiting to run
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (https://github.com/huggingface/kernels/pull/544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
2026-06-02 17:08:54 +02:00
jhcipar d356c31675 feat: bump runpod-python version 2026-06-01 20:28:40 -04:00
Tim PietruskyandGitHub 14b74a4989 chore: re-enable .runpod/tests.json hub tests (#295)
Release / release (push) Waiting to run
Rename tests_json back to tests.json to re-enable the automated hub
tests that were temporarily disabled in #253.

Refs: DR-1161
2026-06-01 17:52:23 +02:00
velaraptor-runpod 80072047ab feat: allow config.yaml like vllm serve 2026-05-29 15:50:22 -05:00
chrisvelaandGitHub 50aba8fb57 Merge pull request #293 from runpod-workers/fix/update-deep-gemm-hub-value
Release / release (push) Waiting to run
fix: update VLLM_USE_DEEP_GEMM hub to default to 0
2026-05-27 11:51:48 -05:00
velaraptor-runpod 9edc5715ce fix: update configuration.md 2026-05-27 11:33:37 -05:00
velaraptor-runpod 4c91f2c5b5 fix: update VLLM_USE_DEEP_GEMM hub to default to 0 2026-05-27 11:26:52 -05:00
chrisvelaandGitHub 6265b99348 Merge pull request #288 from runpod-workers/feat/0.20.0
feat: update to 0.20.2
2026-05-26 17:17:45 -05:00
velaraptor-runpod 026f8d700b fix: specify deepgemm commit version 2026-05-21 18:46:54 -05:00
chrisvelaandGitHub 146bdb0252 Merge branch 'main' into feat/0.20.0 2026-05-21 16:04:39 -05:00
velaraptor-runpod da01193a3d chore: fix readme 2026-05-21 15:57:40 -05:00
velaraptor-runpod c2e6cc9f61 chore: update readme with correct vllm version 2026-05-21 15:39:19 -05:00
velaraptor-runpod 69968a6b39 chore: fix logging, warning for text prompt 2026-05-21 15:34:09 -05:00
velaraptor-runpod 32b29d4c6c fix: add deepgemm, update base image and hub for cuda 13.0 2026-05-20 17:43:42 -05:00
velaraptor-runpod dcea4fc4f9 fix dockerfile 2026-05-15 12:08:31 -04:00
velaraptor-runpod 9c139e8ceb update: update to 0.20.1, update dockerfile to cuda 13 2026-05-15 11:35:53 -04:00
velaraptor-runpod 678bb4be8f feat: update to 0.20.1 for patch fixes 2026-05-07 11:58:22 -05:00
chrisvelaandGitHub 87d7365126 Merge pull request #292 from runpod-workers/fix/open-ai
Release / release (push) Waiting to run
fix: fix warmup
2026-05-01 18:06:11 -05:00
velaraptor-runpod 0e83616f93 fix: fix warmup 2026-05-01 17:39:51 -05:00
chrisvelaandGitHub ab6d39dcf8 Merge pull request #291 from runpod-workers/bug/281-hf-token
Release / release (push) Waiting to run
bug: fix hf-token being passed in engineargs
2026-05-01 15:32:04 -05:00
velaraptor-runpod ed315a175e merge main 2026-05-01 15:28:05 -05:00
velaraptor-runpod 73f030ae5e Merge branch 'main' into bug/281-hf-token 2026-05-01 14:33:51 -05:00
chrisvelaandGitHub 8a099c1723 Merge pull request #287 from runpod-workers/feat/0.19.1
feat: update to 0.19.1
2026-05-01 14:29:39 -05:00
velaraptor-runpod ff87840a58 fix: add enforce_eager as true, add pytorch_alloc_conf to expandle_segments to True for OOM, for hub defaults 2026-05-01 14:13:03 -05:00
velaraptor-runpod 7dc853b1fe chore: remove release to trigger on release publish, just use tags. duplicate 2026-05-01 10:13:32 -05:00
velaraptor-runpod a8b754b92a merge main 2026-05-01 10:12:43 -05:00
chrisvelaandGitHub 6357aeda51 Merge pull request #290 from runpod-workers/fix/fix-old-actions
Release / release (push) Waiting to run
fix: fix old github actions, trigger release on publish release
2026-05-01 10:01:22 -05:00
velaraptor-runpod 0140b29c44 chore: add release notes to slack notification 2026-05-01 09:48:43 -05:00
velaraptor-runpod 0cb8aeae77 chore: add specific runpod version 2026-05-01 09:47:23 -05:00
chrisvelaandGitHub cd8f9e9560 Merge branch 'main' into fix/fix-old-actions 2026-04-30 21:57:45 -05:00
chrisvelaandGitHub 895fd25fac Merge pull request #286 from runpod-workers/feat/0.18.1
feat: update vllm to 0.18.1
2026-04-30 21:56:07 -05:00
velaraptor-runpod 7bb8df73af bug: fix hf-token being passed in engineargs 2026-04-30 20:37:30 -05:00
velaraptor-runpod 747cdf5891 chore: update readme vllm version 2026-04-30 20:26:23 -05:00
velaraptor-runpod 3d4af5df9b chore: update readme vllm version 2026-04-30 20:25:35 -05:00
velaraptor-runpod 72547aa3bb chore: update readme vllm version 2026-04-30 20:25:02 -05:00
velaraptor-runpod 577fd8c3c3 fix: fix old github actions, trigger release on publish release 2026-04-30 20:22:24 -05:00
chrisvelaandGitHub f49f35456e Merge pull request #285 from runpod-workers/feat/0.17.1
feat: update vllm to 0.17.1
2026-04-30 20:05:45 -05:00
chrisvelaandGitHub cff7b09ef2 Merge pull request #289 from runpod-workers/feat/add-notifications
feat: add notifications for new prs, issues, and new releases of vllm
2026-04-30 20:03:48 -05:00
velaraptor-runpod 04b342c675 fix permissions 2026-04-30 20:00:48 -05:00
velaraptor-runpod 5b29643799 feat: add notifications for new prs, issues, and new releases of vllm 2026-04-30 19:55:29 -05:00
velaraptor-runpod 4f8a16df5d fix: update transformers to >=5 2026-04-30 19:41:09 -05:00
velaraptor-runpodandClaude Sonnet 4.6 22356ee2b3 feat: upgrade vLLM to 0.20.0
- Bump vllm[flashinfer] to 0.20.0 in Dockerfile
- Remove io_processor param from OpenAIServingRender (dropped in 0.20.0)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 18:39:07 -05:00
velaraptor-runpodandClaude Sonnet 4.6 fa42ecd79a fix: resolve lowercase HF cache paths when MODEL_NAME uses original casing
Fixes FDE-174. Some model stores (e.g. RunPod pre-cached network volumes)
normalize repo IDs to lowercase. HuggingFace Hub caches using the original
casing, so MODEL_NAME=Qwen/Qwen2.5-Coder-32B-Instruct-AWQ would miss a
cache stored as models--qwen--qwen2.5-coder-32b-instruct-awq/ and attempt
a redundant download that fails on limited container storage.

If the exact-case HF cache directory is absent but a lowercase variant
exists, the latest snapshot path is returned directly so vLLM loads from
disk. Absolute paths and models with no lowercase cache are unchanged.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 17:54:14 -05:00
velaraptor-runpodandClaude Sonnet 4.6 178c72238e fix: surface LORA_MODULES parse failures instead of silently loading zero adapters
Fixes FDE-194. Previously a malformed LORA_MODULES value was swallowed at
info level and the engine would start with no LoRA adapters, causing 500s
on any request using an adapter model name (e.g. npc-sim-*).

Changes:
- Log at error level when LORA_MODULES cannot be parsed as JSON
- Log at error level when individual adapter dicts fail LoRAModulePath validation
- Log a final error when all adapters fail to load so the cause is obvious
- Accept a single adapter dict (not just an array) for convenience
- Return early when LORA_MODULES is unset to skip unnecessary parsing

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 17:49:09 -05:00
velaraptor-runpodandClaude Sonnet 4.6 e6950bdebd feat: upgrade vLLM to 0.19.1
- Bump vllm[flashinfer] to 0.19.1 in Dockerfile
- Add OpenAIServingRender (new required dependency in 0.19.x serving layer)
- Pass openai_serving_render to all four serving class constructors
- Remove log_error_stack param (removed upstream in 0.19.x)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 17:41:11 -05:00
velaraptor-runpod a774cefe85 fix: clean workspace file 2026-04-30 17:29:00 -05:00
velaraptor-runpod 9de17d49b7 feat: update vllm to 0.18.1 2026-04-30 17:25:23 -05:00
velaraptor-runpod 296556a6f7 feat: update vllm to 0.17.1 2026-04-30 17:21:03 -05:00
23 changed files with 836 additions and 96 deletions
@@ -0,0 +1,71 @@
name: CI | Sync vLLM version in READMEs
on:
push:
branches: ["main"]
paths:
- "Dockerfile"
workflow_dispatch:
permissions:
contents: write
pull-requests: write
jobs:
sync_version:
runs-on: ubuntu-latest
name: Check README version matches Dockerfile and update if needed
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Extract vLLM version from Dockerfile and sync READMEs
run: |
echo "Extracting vLLM version from Dockerfile..."
dockerfile_version=$(grep -oP 'vllm(?:\[[\w,]+\])?==\K[\d.]+' Dockerfile | head -1)
if [ -z "$dockerfile_version" ]; then
echo "ERROR: Could not extract vLLM version from Dockerfile."
exit 1
fi
echo "Dockerfile vLLM version: $dockerfile_version"
echo "VLLM_VERSION=$dockerfile_version" >> $GITHUB_ENV
updated=0
for readme in README.md .runpod/README.md; do
if [ ! -f "$readme" ]; then
echo "Skipping $readme (not found)"
continue
fi
readme_version=$(grep -oP 'Current vLLM version: \[\K[\d.]+' "$readme" || echo "")
echo "$readme current version: ${readme_version:-not found}"
if [ "$readme_version" = "$dockerfile_version" ]; then
echo "$readme is already up to date."
continue
fi
echo "Updating $readme from $readme_version to $dockerfile_version..."
sed -i "s|Current vLLM version: \[${readme_version}\](https://github.com/vllm-project/vllm/releases/tag/v${readme_version})|Current vLLM version: [${dockerfile_version}](https://github.com/vllm-project/vllm/releases/tag/v${dockerfile_version})|g" "$readme"
updated=1
done
echo "UPDATED=$updated" >> $GITHUB_ENV
- name: Create Pull Request
if: env.UPDATED == '1'
uses: peter-evans/create-pull-request@v7
with:
token: ${{ secrets.GITHUB_TOKEN }}
commit-message: "docs: sync vLLM version to ${{ env.VLLM_VERSION }} in READMEs"
title: "docs: sync vLLM version to ${{ env.VLLM_VERSION }} in READMEs"
body: |
The vLLM version in the Dockerfile has been updated to `${{ env.VLLM_VERSION }}`.
This PR syncs the version badge/link in:
- `README.md`
- `.runpod/README.md`
branch: docs/sync-vllm-version-${{ env.VLLM_VERSION }}
labels: documentation
+32 -31
View File
@@ -9,59 +9,60 @@ on:
workflow_dispatch:
permissions:
contents: write
pull-requests: write
jobs:
check_dep:
runs-on: ubuntu-latest
name: Check python requirements file and update
steps:
- name: Checkout
uses: actions/checkout@v2
uses: actions/checkout@v4
- name: Check for new package version and update
run: |
echo "Fetching the current runpod version from requirements.txt..."
# Get current version, allowing both == and ~= in the search pattern
current_version=$(grep -oP 'runpod[~=]{1,2}\K[^"]+' ./builder/requirements.txt)
echo "Current version: $current_version"
echo "Fetching current runpod version from requirements.txt..."
# Extract major and minor from current version
current_major_minor=$(echo $current_version | cut -d. -f1,2)
echo "Current major.minor: $current_major_minor"
# Match runpod with any version specifier or no specifier at all
current_version=$(grep -oP '^runpod([~>=!<]{1,2}\K[\d.]+)?' ./builder/requirements.txt | grep -oP '[\d.]+' || echo "")
echo "Current version: ${current_version:-unset}"
echo "Fetching the latest runpod version from PyPI..."
# Get new version from PyPI
new_version=$(curl -s https://pypi.org/pypi/runpod/json | jq -r .info.version)
echo "Fetching latest runpod version from PyPI..."
new_version=$(curl -sf https://pypi.org/pypi/runpod/json | jq -r .info.version)
echo "NEW_VERSION_ENV=$new_version" >> $GITHUB_ENV
echo "New version: $new_version"
# Extract major and minor from new version
new_major_minor=$(echo $new_version | cut -d. -f1,2)
echo "New major.minor: $new_major_minor"
if [ -z "$new_version" ]; then
echo "ERROR: Failed to fetch the new version from PyPI."
exit 1
echo "ERROR: Failed to fetch new version from PyPI."
exit 1
fi
# Check if the major or minor version is different
if [ "$current_major_minor" = "$new_major_minor" ]; then
echo "No update needed. The new version ($new_major_minor) is within the allowed range (~= $current_major_minor)."
if [ -z "$current_version" ]; then
echo "No version pin found — pinning to $new_version."
else
current_major_minor=$(echo "$current_version" | cut -d. -f1,2)
new_major_minor=$(echo "$new_version" | cut -d. -f1,2)
echo "Current major.minor: $current_major_minor New major.minor: $new_major_minor"
if [ "$current_major_minor" = "$new_major_minor" ]; then
echo "No update needed. New version ($new_version) is within ~= $current_major_minor range."
exit 0
fi
echo "New major/minor detected ($new_major_minor). Updating requirements.txt..."
fi
echo "New major/minor detected ($new_major_minor). Updating requirements.txt..."
# Update requirements.txt, preserving the existing constraint type (~= or ==)
sed -i "s/runpod[~=][^ ]*/runpod~=$new_version/" ./builder/requirements.txt
echo "requirements.txt has been updated."
# Replace any `runpod`, `runpod==x`, `runpod~=x`, etc. with pinned version
sed -i "s|^runpod.*|runpod~=$new_version|" ./builder/requirements.txt
echo "requirements.txt updated."
- name: Create Pull Request
uses: peter-evans/create-pull-request@v3
uses: peter-evans/create-pull-request@v7
with:
token: ${{ secrets.GITHUB_TOKEN }}
commit-message: Update runpod package version
title: Update runpod package version
body: The package version has been updated to ${{ env.NEW_VERSION_ENV }}
commit-message: "chore: update runpod to ${{ env.NEW_VERSION_ENV }}"
title: "chore: update runpod to ${{ env.NEW_VERSION_ENV }}"
body: The `runpod` package has been updated to `${{ env.NEW_VERSION_ENV }}`.
branch: runpod-package-update
+43 -12
View File
@@ -3,7 +3,7 @@ name: Release
on:
push:
tags:
- "v[0-9]+.[0-9]+.[0-9]+*" # Trigger on version tags like v1.0.0, v2.1.0, etc.
- "v[0-9]+.[0-9]+.[0-9]+*"
workflow_dispatch:
inputs:
version:
@@ -53,16 +53,13 @@ jobs:
# Determine version based on trigger type
if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
# Manual trigger: use input version
VERSION="${{ github.event.inputs.version }}"
echo "RELEASE_VERSION=${VERSION}" >> $GITHUB_ENV
echo "IS_MANUAL_RELEASE=true" >> $GITHUB_ENV
elif [[ "${{ github.event_name }}" == "release" ]]; then
VERSION="${{ github.event.release.tag_name }}"
else
# Tag trigger: use tag name (remove refs/tags/ prefix)
VERSION=${GITHUB_REF#refs/tags/}
echo "RELEASE_VERSION=${VERSION}" >> $GITHUB_ENV
echo "IS_MANUAL_RELEASE=false" >> $GITHUB_ENV
fi
echo "RELEASE_VERSION=${VERSION}" >> $GITHUB_ENV
- name: Build and push the images to Docker Hub
uses: docker/bake-action@v2
@@ -76,11 +73,45 @@ jobs:
- name: Release Summary
run: |
echo "🚀 Release completed!"
echo "Release completed!"
echo "Version: ${{ env.RELEASE_VERSION }}"
echo "Docker Image: ${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}"
if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
echo "Trigger: Manual workflow dispatch"
else
echo "Trigger: GitHub release (tag: ${{ github.ref_name }})"
- name: Fetch Release Notes
run: |
RESPONSE=$(curl -sf \
-H "Authorization: token ${{ github.token }}" \
"https://api.github.com/repos/${{ github.repository }}/releases/tags/${{ env.RELEASE_VERSION }}" 2>/dev/null) || true
if [[ -n "$RESPONSE" ]]; then
NOTES=$(echo "$RESPONSE" | jq -r '.body // empty')
fi
printf '%s' "${NOTES:-No release notes available.}" > /tmp/release_notes.txt
- name: Notify Slack
run: |
jq -n \
--arg version "${{ env.RELEASE_VERSION }}" \
--arg docker "${{ env.DOCKERHUB_REPO }}/${{ env.DOCKERHUB_IMG }}:${{ env.RELEASE_VERSION }}" \
--rawfile notes /tmp/release_notes.txt \
--arg url "https://github.com/${{ github.repository }}/releases/tag/${{ env.RELEASE_VERSION }}" \
'{
text: (":rocket: New :runpod-new-whiteonpurple: Runpod worker-vllm release: *" + $version + "*"),
blocks: [
{
type: "section",
text: {
type: "mrkdwn",
text: (":banana-dance: *New Release — worker-vllm " + $version + "*\n*Docker:* `" + $docker + "`\n<" + $url + "|View release on GitHub>")
}
},
{
type: "section",
text: {
type: "mrkdwn",
text: ("*Release Notes:*\n" + $notes)
}
}
]
}' | curl -sf -X POST "${{ secrets.SLACK_WEBHOOK_URL }}" \
-H "Content-Type: application/json" \
-d @-
@@ -0,0 +1,41 @@
name: Slack PR Notifications
on:
pull_request:
types: [opened]
issues:
types: [opened]
permissions:
contents: read
jobs:
notify:
runs-on: ubuntu-latest
steps:
- name: Notify Slack - New PR
if: github.event_name == 'pull_request'
run: |
curl -sf -X POST "${{ secrets.SLACK_WEBHOOK_URL }}" \
-H "Content-Type: application/json" \
-d '{
"text": ":rocket: New PR in worker-vllm: *${{ github.event.pull_request.title }}*",
"blocks": [
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": ":rocket: *New Pull Request — worker-vllm*\n*<${{ github.event.pull_request.html_url }}|${{ github.event.pull_request.title }}>*\nOpened by *${{ github.event.pull_request.user.login }}*"
}
},
{
"type": "context",
"elements": [
{
"type": "mrkdwn",
"text": "${{ github.event.pull_request.base.ref }} ← ${{ github.event.pull_request.head.ref }}"
}
]
}
]
}'
+73
View File
@@ -0,0 +1,73 @@
name: Monitor vLLM Releases
on:
schedule:
- cron: '0 0 * * *' # Every day at midnight
workflow_dispatch:
permissions:
contents: read
jobs:
check-vllm-release:
runs-on: ubuntu-latest
steps:
- name: Restore last known vLLM tag
uses: actions/cache/restore@v4
with:
path: .vllm-last-tag
key: vllm-tag-${{ github.run_id }}
restore-keys: vllm-tag-
- name: Get latest vLLM release
id: vllm
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
response=$(curl -sf https://api.github.com/repos/vllm-project/vllm/releases/latest \
-H "Authorization: Bearer $GH_TOKEN")
echo "tag=$(echo "$response" | jq -r '.tag_name')" >> $GITHUB_OUTPUT
echo "url=$(echo "$response" | jq -r '.html_url')" >> $GITHUB_OUTPUT
echo "name=$(echo "$response" | jq -r '.name')" >> $GITHUB_OUTPUT
- name: Check if new release
id: check
run: |
last=$(cat .vllm-last-tag 2>/dev/null || echo "")
current="${{ steps.vllm.outputs.tag }}"
echo "Last: $last Current: $current"
if [ -n "$current" ] && [ "$last" != "$current" ]; then
echo "is_new=true" >> $GITHUB_OUTPUT
else
echo "is_new=false" >> $GITHUB_OUTPUT
fi
- name: Notify Slack
if: steps.check.outputs.is_new == 'true'
run: |
curl -sf -X POST "${{ secrets.SLACK_WEBHOOK_URL }}" \
-H "Content-Type: application/json" \
-d '{
"text": ":rocket: New vLLM release: *${{ steps.vllm.outputs.tag }}*",
"blocks": [
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": ":rocket: *New vLLM Release: ${{ steps.vllm.outputs.tag }}*\n<${{ steps.vllm.outputs.url }}|View on GitHub>"
}
}
]
}'
- name: Save new tag
if: steps.check.outputs.is_new == 'true'
run: echo "${{ steps.vllm.outputs.tag }}" > .vllm-last-tag
- name: Update cache
if: steps.check.outputs.is_new == 'true'
uses: actions/cache/save@v4
with:
path: .vllm-last-tag
key: vllm-tag-${{ steps.vllm.outputs.tag }}
+32
View File
@@ -0,0 +1,32 @@
name: Tests
on:
pull_request:
branches:
- "**"
push:
branches:
- "main"
permissions:
contents: read
jobs:
pytest:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install test dependencies
run: |
python -m pip install --upgrade pip
pip install -r tests/requirements.txt
- name: Run unit tests
run: python -m pytest tests -v
+14
View File
@@ -6,6 +6,8 @@ Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API
[![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm)
Current vLLM version: [0.20.2](https://github.com/vllm-project/vllm/releases/tag/v0.20.2)
---
## Endpoint Configuration
@@ -27,9 +29,21 @@ All behaviour is controlled through environment variables:
| `REASONING_PARSER` | Parser for reasoning-capable models | | "deepseek_r1", "qwen3", "granite", "hunyuan_a13b" |
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | Override served model name in API | | String |
| `MAX_CONCURRENCY` | Maximum concurrent requests | 300 | Integer |
| `ENFORCE_EAGER` | If True, we will disable CUDA graph and always execute the model in eager mode. If False, we will use CUDA graph and eager execution in hybrid for maximal performance and flexibility. | true | boolean (true or false) |
**Pass any vLLM engine arg** not listed above by setting an env var with the **UPPERCASED** field name (e.g. `MAX_MODEL_LEN=4096`, `ENABLE_CHUNKED_PREFILL=true`). The worker auto-discovers all `AsyncEngineArgs` fields from env. See the [vLLM engine args docs](https://docs.vllm.ai/en/latest/configuration/engine_args) for all available options.
**Configuration file:** You can also supply a `config.yaml` instead of (or alongside) env vars. Mount it at `/vllm_config.yaml` in the container, or set `VLLM_CONFIG_FILE` to a custom path. Use the same key names as `vllm serve` — hyphens and underscores both work:
```yaml
model: meta-llama/Llama-3.1-8B-Instruct
max-model-len: 8192
gpu-memory-utilization: 0.90
quantization: awq
```
Environment variables always override config file values.
For complete configuration options, see the [full configuration documentation](https://github.com/runpod-workers/worker-vllm/blob/main/docs/configuration.md).
### Specify Transformers Version
+22 -2
View File
@@ -9,7 +9,7 @@
"containerDiskInGb": 150,
"gpuIds": "ADA_80_PRO,AMPERE_80",
"gpuCount": 1,
"allowedCudaVersions": ["12.9", "12.8"],
"allowedCudaVersions": ["13.0"],
"presets": [
{
"name": "deepseek-ai/deepseek-r1-distill-llama-8b",
@@ -621,7 +621,7 @@
"name": "Enforce Eager",
"type": "boolean",
"description": "Always use eager-mode PyTorch. If False (0), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility",
"default": false,
"default": true,
"advanced": true
}
},
@@ -795,6 +795,26 @@
"default": "",
"advanced": true
}
},
{
"key": "PYTORCH_ALLOC_CONF",
"input": {
"name": "PyTorch Alloc Config",
"type": "string",
"description": "PyTorch allocation configuration, remove this if you want to use the default configuration",
"default": "expandable_segments:True",
"advanced": true
}
},
{
"key": "VLLM_USE_DEEP_GEMM",
"input": {
"name": "Use DeepGEMM",
"type": "string",
"description": "Enable DeepGEMM FP8 kernels (MoE and MQA logits). Set to 1 to enable, 0 to disable. Required for DeepSeek V4 models. Disabled by default — enable on H100/H200 for potential throughput gains. Some GPUs (e.g. H20) may perform better with this off.",
"default": "0",
"advanced": true
}
}
]
}
+4 -4
View File
@@ -5,7 +5,7 @@
"input": {
"prompt": "Write a short poem about artificial intelligence."
},
"timeout": 30000
"timeout": 300000
},
{
"name": "openai_messages_test",
@@ -26,11 +26,11 @@
"temperature": 0.1
}
},
"timeout": 30000
"timeout": 300000
}
],
"config": {
"gpuTypeId": "NVIDIA GeForce RTX 4090",
"gpuTypeId": "NVIDIA L40",
"gpuCount": 1,
"env": [
{
@@ -38,6 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
}
],
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"]
"allowedCudaVersions": ["13.0"]
}
}
+10 -7
View File
@@ -1,16 +1,17 @@
FROM nvidia/cuda:12.9.1-base-ubuntu22.04
FROM nvidia/cuda:13.0.2-devel-ubuntu22.04
RUN apt-get update -y \
&& apt-get install -y python3-pip curl \
&& curl -LsSf https://astral.sh/uv/0.10.9/install.sh | sh
&& apt-get install -y python3-pip curl git \
&& curl -LsSf https://astral.sh/uv/install.sh | sh
ENV PATH="/root/.local/bin:$PATH"
RUN ldconfig /usr/local/cuda-12.9/compat/
RUN ldconfig /usr/local/cuda-13.0/compat/
# Install vLLM with FlashInfer - use CUDA 12.9 PyTorch wheels
# Install vLLM with FlashInfer - use CUDA 130 PyTorch wheels
RUN uv pip install --system "packaging>=24.2" && \
uv pip install --system "vllm[flashinfer]==0.16.0" --extra-index-url https://download.pytorch.org/whl/cu129
uv pip install --system "vllm[flashinfer]==0.20.2" && \
uv pip install --system git+https://github.com/deepseek-ai/DeepGEMM.git@714dd1a4a980f7937a74343d19a8eba4fe321480 --no-build-isolation
# Install additional Python dependencies (after vLLM to avoid PyTorch version conflicts)
COPY builder/requirements.txt /requirements.txt
@@ -42,7 +43,9 @@ ENV MODEL_NAME=$MODEL_NAME \
# Prevent rayon thread pool panic in containers where ulimit -u < nproc
# (tokenizers uses Rust's rayon which tries to spawn threads = CPU cores)
TOKENIZERS_PARALLELISM=false \
RAYON_NUM_THREADS=4
RAYON_NUM_THREADS=4 \
# Disable DeepGEMM MoE kernels by default; override with VLLM_USE_DEEP_GEMM=1 to enable
VLLM_USE_DEEP_GEMM=0
ENV PYTHONPATH="/:/vllm-workspace"
+17 -2
View File
@@ -8,7 +8,8 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png)
Current vLLM version: [0.16.0](https://github.com/vllm-project/vllm/releases/tag/v0.16.0)
Current vLLM version: [0.20.2](https://github.com/vllm-project/vllm/releases/tag/v0.20.2)
> Check out our Load Balancer implementation here: [vLLM Load Balancer](https://github.com/runpod-workers/vllm-loadbalancer-ep)
@@ -46,7 +47,7 @@ Current vLLM version: [0.16.0](https://github.com/vllm-project/vllm/releases/tag
**📦 Docker Image**: `runpod/worker-v1-vllm:<version>`
- **Available Versions**: See [GitHub Releases](https://github.com/runpod-workers/worker-vllm/releases)
- **CUDA Compatibility**: Requires CUDA >= 12.1
- **CUDA Compatibility**: Requires CUDA >= 13.0
### Configuration
@@ -77,6 +78,20 @@ Configure worker-vllm using environment variables:
Any env var whose name matches a valid `AsyncEngineArgs` field (uppercased) is applied automatically. Backward-compat aliases: `MODEL_NAME`, `TOKENIZER_NAME`, `MAX_CONTEXT_LEN_TO_CAPTURE`. This lets you configure any vLLM option without waiting for explicit worker support.
### Configuration File (config.yaml)
As an alternative to environment variables, you can supply a `config.yaml` file using the same key names as `vllm serve` (hyphens or underscores both work):
```yaml
model: meta-llama/Llama-3.1-8B-Instruct
max-model-len: 8192
gpu-memory-utilization: 0.90
quantization: awq
tensor-parallel-size: 2
```
Mount the file into the container at `/vllm_config.yaml`, or point to a custom path with the `VLLM_CONFIG_FILE` env var. Environment variables always take precedence over config file values.
For the complete list of all available environment variables, examples, and detailed descriptions: **[Configuration](docs/configuration.md)**
### Specify Transformers Version
+4 -4
View File
@@ -1,15 +1,15 @@
ray
pandas
pyarrow
runpod
runpod~=1.10.0
huggingface-hub
lmcache==0.4.1
lmcache==0.4.5
packaging>=24.2
typing-extensions>=4.8.0
pydantic
pydantic-settings
hf-transfer
transformers>=4.57.0,<5
transformers>=5
bitsandbytes>=0.45.0
kernels
kernels<0.15
torch-c-dlpack-ext
+11
View File
@@ -0,0 +1,11 @@
model: google/gemma-4-31b-it
gpu-memory-utilization: 0.95
max-model-len: 8192
dtype: auto
trust-remote-code: true
quantization: fp8
kv-cache-dtype: fp8
enforce-eager: false
enable-prefix-caching: true
enable-chunked-prefill: true
speculative-config: '{"model":"RedHatAI/gemma-4-31B-it-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
+9
View File
@@ -0,0 +1,9 @@
model: openai/gpt-oss-120b
gpu-memory-utilization: 0.95
max-model-len: 8192
dtype: auto
trust-remote-code: true
enforce-eager: false
enable-prefix-caching: true
enable-chunked-prefill: true
speculative-config: '{"model":"RedHatAI/gpt-oss-120b-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
+10
View File
@@ -0,0 +1,10 @@
model: meta-llama/Llama-3.1-8B-Instruct
gpu-memory-utilization: 0.95
max-model-len: 8192
dtype: auto
trust-remote-code: true
quantization: fp8
kv-cache-dtype: fp8
enforce-eager: false
enable-prefix-caching: true
speculative-config: '{"model":"RedHatAI/Llama-3.1-8B-Instruct-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
+10
View File
@@ -0,0 +1,10 @@
model: Qwen/Qwen3-8B
gpu-memory-utilization: 0.95
max-model-len: 8192
dtype: auto
trust-remote-code: true
quantization: fp8
kv-cache-dtype: fp8
enforce-eager: false
enable-prefix-caching: true
speculative-config: '{"model":"RedHatAI/Qwen3-8B-speculator.eagle3","method":"eagle3","num_speculative_tokens":3}'
+3
View File
@@ -97,10 +97,13 @@ If `SPECULATIVE_CONFIG` is set, it takes priority over individual env vars. When
| `MAX_SEQ_LEN_TO_CAPTURE` | `8192` | `int` | Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode. |
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` | Enables or disables custom all reduce. |
| `ENABLE_EXPERT_PARALLEL` | `False` | `bool` | Enable Expert Parallel for MoE models. |
| `VLLM_USE_DEEP_GEMM` | `0` | `str` (`0`/`1`) | Enable DeepGEMM FP8 kernels for MoE and MQA logits computation. Disabled by default. Must be `"0"` or `"1"` — not `true`/`false`. See note below. |
| `ATTENTION_BACKEND` | `None` | `str` | Attention backend to use (e.g., `FLASH_ATTN`, `FLASHINFER`, `TRITON_FLASH_ATTN`). Replaces deprecated `VLLM_ATTENTION_BACKEND`. |
| `ASYNC_SCHEDULING` | `None` | `bool` | Enable async scheduling (overlaps engine scheduling with GPU execution). Default: enabled in vLLM 0.14.0+. Set to `false` to disable. |
| `STREAM_INTERVAL` | `1` | `int` | Controls how often to yield streaming results. Lower = more frequent updates. |
> **Note (`VLLM_USE_DEEP_GEMM`):** DeepGEMM is used in two places: MoE weight computation and MQA logits computation. It is necessary for MQA logits computation on supported hardware — required for DeepSeek V4 models. Set `VLLM_USE_DEEP_GEMM=1` to enable. Set `VLLM_USE_DEEP_GEMM=0` to disable the MoE part and fall back to flashinfer/cutlass FP8 kernels. **Value must be `"0"` or `"1"` — not `"true"`/`"false"`.** Some users report better performance with `VLLM_USE_DEEP_GEMM=0`, particularly on H20 GPUs. Disabling it also skips the DeepGEMM warmup phase, reducing cold-start time. Requires CUDA 13.0+ and SM90+ (H100/H200) to use; the library is installed but inactive by default.
## Tokenizer Settings
| Variable | Default | Type/Choices | Description |
+96 -32
View File
@@ -1,4 +1,5 @@
import asyncio
import inspect
import json
import logging
import os
@@ -7,6 +8,7 @@ from typing import AsyncGenerator, Optional
from dotenv import load_dotenv
from vllm import AsyncLLMEngine
from vllm.inputs import TextPrompt
from vllm.entrypoints.logger import RequestLogger
from vllm.entrypoints.anthropic.protocol import AnthropicMessagesRequest, AnthropicMessagesResponse, AnthropicError, AnthropicErrorResponse
from vllm.entrypoints.anthropic.serving import AnthropicServingMessages
@@ -19,6 +21,7 @@ from vllm.entrypoints.openai.models.protocol import BaseModelPath, LoRAModulePat
from vllm.entrypoints.openai.models.serving import OpenAIServingModels
from vllm.entrypoints.openai.responses.protocol import ResponsesRequest, ResponsesResponse
from vllm.entrypoints.openai.responses.serving import OpenAIServingResponses
from vllm.entrypoints.serve.render.serving import OpenAIServingRender
from constants import DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MAX_CONCURRENCY, DEFAULT_MIN_BATCH_SIZE
from engine_args import get_engine_args
@@ -29,20 +32,33 @@ class vLLMEngine:
def __init__(self, engine = None):
load_dotenv() # For local development
self.engine_args = get_engine_args()
logging.info(f"Engine args: {self.engine_args}")
# Initialize vLLM engine first
self.llm = self._initialize_llm() if engine is None else engine.llm
# Only create custom tokenizer wrapper if not using mistral tokenizer mode
# For mistral models, let vLLM handle tokenizer initialization
if self.engine_args.tokenizer_mode != 'mistral':
self.tokenizer = TokenizerWrapper(self.engine_args.tokenizer or self.engine_args.model,
self.engine_args.tokenizer_revision,
self.engine_args.trust_remote_code)
if engine is None:
ea = self.engine_args
summary = {
"model": ea.model,
"dtype": ea.dtype,
"quantization": ea.quantization,
"max_model_len": ea.max_model_len,
"tensor_parallel_size": ea.tensor_parallel_size,
"gpu_memory_utilization": ea.gpu_memory_utilization,
}
if ea.tokenizer and ea.tokenizer != ea.model:
summary["tokenizer"] = ea.tokenizer
logging.info("Engine config: %s", summary)
logging.debug("Full engine args: %s", ea)
self.llm = self._initialize_llm()
if self.engine_args.tokenizer_mode != 'mistral':
self.tokenizer = TokenizerWrapper(self.engine_args.tokenizer or self.engine_args.model,
self.engine_args.tokenizer_revision,
self.engine_args.trust_remote_code)
else:
self.tokenizer = None
else:
# For mistral models, we'll get the tokenizer from vLLM later
self.tokenizer = None
self.llm = engine.llm
self.tokenizer = engine.tokenizer
self.max_concurrency = int(os.getenv("MAX_CONCURRENCY", DEFAULT_MAX_CONCURRENCY))
self.default_batch_size = int(os.getenv("DEFAULT_BATCH_SIZE", DEFAULT_BATCH_SIZE))
@@ -115,7 +131,7 @@ class vLLMEngine:
if apply_chat_template or isinstance(llm_input, list):
tokenizer_wrapper = self._get_tokenizer_for_chat_template()
llm_input = tokenizer_wrapper.apply_chat_template(llm_input)
results_generator = self.llm.generate(llm_input, validated_sampling_params, request_id)
results_generator = self.llm.generate(TextPrompt(prompt=llm_input), validated_sampling_params, request_id)
n_responses, n_input_tokens, is_first_output = validated_sampling_params.n, 0, True
last_output_texts, token_counters = ["" for _ in range(n_responses)], {"batch": 0, "total": 0}
@@ -205,19 +221,48 @@ class OpenAIvLLMEngine(vLLMEngine):
self.raw_openai_output = bool(int(raw_output_env))
def _load_lora_adapters(self):
adapters = []
try:
adapters = json.loads(os.getenv("LORA_MODULES", '[]'))
except Exception as e:
logging.info(f"---Initialized adapter json load error: {e}")
lora_modules_env = os.getenv("LORA_MODULES", "")
if not lora_modules_env:
return []
for i, adapter in enumerate(adapters):
try:
parsed = json.loads(lora_modules_env)
except json.JSONDecodeError as e:
logging.error(
"LORA_MODULES could not be parsed as JSON: %s — no LoRA adapters loaded. Value: %r",
e, lora_modules_env,
)
return []
# Accept a single adapter dict as well as an array
if isinstance(parsed, dict):
parsed = [parsed]
if not isinstance(parsed, list):
logging.error(
"LORA_MODULES must be a JSON array of adapter objects, got %s — no LoRA adapters loaded.",
type(parsed).__name__,
)
return []
adapters = []
for i, adapter in enumerate(parsed):
try:
adapters[i] = LoRAModulePath(**adapter)
logging.info(f"---Initialized adapter: {adapter}")
adapters.append(LoRAModulePath(**adapter))
logging.info("Loaded LoRA adapter config [%d]: %s", i, adapter)
except Exception as e:
logging.info(f"---Initialized adapter not worked: {e}")
continue
logging.error(
"Failed to parse LoRA adapter at index %d: %s. Config: %r",
i, e, adapter,
)
if parsed and not adapters:
logging.error(
"LORA_MODULES specified %d adapter(s) but none could be loaded — "
"OpenAI model name lookups for LoRA adapters will fail.",
len(parsed),
)
return adapters
async def _ensure_engines_initialized(self):
@@ -246,16 +291,32 @@ class OpenAIvLLMEngine(vLLMEngine):
lora_modules=self.lora_adapters,
)
await self.serving_models.init_static_loras()
# Get chat template from vLLM tokenizer if available
chat_template = None
if self.tokenizer and hasattr(self.tokenizer, 'tokenizer'):
chat_template = self.tokenizer.tokenizer.chat_template
self.openai_serving_render = OpenAIServingRender(
model_config=self.llm.model_config,
renderer=self.llm.renderer,
model_registry=self.serving_models.registry,
request_logger=None,
chat_template=chat_template,
chat_template_content_format="auto",
trust_request_chat_template=os.getenv('TRUST_REQUEST_CHAT_TEMPLATE', 'false').lower() == 'true',
enable_auto_tools=os.getenv('ENABLE_AUTO_TOOL_CHOICE', 'false').lower() == 'true',
exclude_tools_when_tool_choice_none=os.getenv('EXCLUDE_TOOLS_WHEN_TOOL_CHOICE_NONE', 'false').lower() == 'true',
tool_parser=os.getenv('TOOL_CALL_PARSER', "") or None,
reasoning_parser=os.getenv('REASONING_PARSER', "") or None,
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
)
self.chat_engine = OpenAIServingChat(
engine_client=self.llm,
engine_client=self.llm,
models=self.serving_models,
response_role=self.response_role,
openai_serving_render=self.openai_serving_render,
request_logger=None,
chat_template=chat_template,
chat_template_content_format="auto",
@@ -268,20 +329,20 @@ class OpenAIvLLMEngine(vLLMEngine):
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
enable_log_outputs=os.getenv('ENABLE_LOG_OUTPUTS', 'false').lower() == 'true',
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
)
self.completion_engine = OpenAIServingCompletion(
engine_client=self.llm,
models=self.serving_models,
openai_serving_render=self.openai_serving_render,
request_logger=None,
return_tokens_as_token_ids=os.getenv('RETURN_TOKENS_AS_TOKEN_IDS', 'false').lower() == 'true',
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
)
self.responses_engine = OpenAIServingResponses(
engine_client=self.llm,
models=self.serving_models,
openai_serving_render=self.openai_serving_render,
request_logger=None,
chat_template=chat_template,
chat_template_content_format="auto",
@@ -293,12 +354,12 @@ class OpenAIvLLMEngine(vLLMEngine):
enable_prompt_tokens_details=os.getenv('ENABLE_PROMPT_TOKENS_DETAILS', 'false').lower() == 'true',
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
enable_log_outputs=os.getenv('ENABLE_LOG_OUTPUTS', 'false').lower() == 'true',
log_error_stack=os.getenv('LOG_ERROR_STACK', 'false').lower() == 'true',
)
self.messages_engine = AnthropicServingMessages(
engine_client=self.llm,
models=self.serving_models,
response_role=self.response_role,
openai_serving_render=self.openai_serving_render,
request_logger=None,
chat_template=chat_template,
chat_template_content_format="auto",
@@ -310,8 +371,11 @@ class OpenAIvLLMEngine(vLLMEngine):
enable_force_include_usage=os.getenv('ENABLE_FORCE_INCLUDE_USAGE', 'false').lower() == 'true',
)
if hasattr(self.chat_engine, 'warmup'):
await self.chat_engine.warmup()
warmup = getattr(self.chat_engine, 'warmup', None)
if callable(warmup):
result = warmup()
if inspect.isawaitable(result):
await result
async def generate(self, openai_request: JobInput):
# Ensure engines are ready (no-op if already initialized at startup)
+93
View File
@@ -82,6 +82,16 @@ def _convert_env_value_to_field_type(value: str, field_name: str, field_type: ty
if type(None) in (args or ()):
return None
raise ValueError("empty value not allowed for non-optional field")
# Union[bool, str, ...]: only coerce to bool for unambiguous literals;
# otherwise preserve the string (e.g. hf_token="hf_abc..." must stay a str).
if get_origin(field_type) is not None:
union_types = [a for a in (get_args(field_type) or ()) if a is not type(None)]
if bool in union_types and str in union_types:
if str(val).lower() in ("true", "false", "1", "0", "yes", "no", "on", "off"):
return str(val).lower() in ("true", "1", "yes", "on")
return str(val)
effective_type = _resolve_field_type(field_type)
# bool
if effective_type is bool:
@@ -342,6 +352,75 @@ def _sanitize_hf_overrides(hf_overrides: dict) -> dict | None:
return result or None
def _resolve_cached_model_path(model_name: str) -> str:
"""Return a local snapshot path when the HF cache was stored with lowercase names.
Some model stores (e.g. RunPod pre-cached volumes) normalize repo IDs to
lowercase. HuggingFace Hub stores caches as
``models--{org}--{model}/snapshots/{hash}/`` preserving the original casing,
so MODEL_NAME=Qwen/Qwen2.5-Coder-32B-Instruct-AWQ will miss a cache stored
as ``models--qwen--qwen2.5-coder-32b-instruct-awq/``.
If the exact-case cache directory is absent but a lowercase variant exists,
the latest snapshot path is returned so vLLM loads from disk rather than
attempting a redundant download.
"""
if os.path.isabs(model_name):
return model_name
cache_dir = (
os.getenv("HUGGINGFACE_HUB_CACHE")
or os.getenv("HF_HOME")
or os.path.expanduser("~/.cache/huggingface/hub")
)
folder_name = f"models--{model_name.replace('/', '--')}"
if os.path.isdir(os.path.join(cache_dir, folder_name)):
return model_name
lower_dir = os.path.join(cache_dir, folder_name.lower())
if not os.path.isdir(lower_dir):
return model_name
snapshots_dir = os.path.join(lower_dir, "snapshots")
if not os.path.isdir(snapshots_dir):
return model_name
try:
snapshots = sorted(os.listdir(snapshots_dir))
except OSError:
return model_name
if not snapshots:
return model_name
resolved = os.path.join(snapshots_dir, snapshots[-1])
logging.info(
"MODEL_NAME %r not found at original casing in HF cache; "
"resolved to lowercase cached snapshot at %r",
model_name, resolved,
)
return resolved
def _get_args_from_config_file() -> dict:
"""Load engine args from a vLLM-style config.yaml.
Checks VLLM_CONFIG_FILE env var, then falls back to /vllm_config.yaml.
Keys use the same long-form names as vllm serve (hyphens converted to underscores).
"""
import yaml
path = os.getenv("VLLM_CONFIG_FILE", "/vllm_config.yaml")
if not os.path.exists(path):
return {}
with open(path) as f:
raw = yaml.safe_load(f) or {}
normalized = {k.replace("-", "_"): v for k, v in raw.items()}
logging.info("Loaded engine args from config file %s: %s", path, list(normalized.keys()))
return normalized
def get_local_args():
"""
Retrieve local arguments from a JSON file.
@@ -367,6 +446,9 @@ def get_engine_args():
# Start with worker custom defaults (only where we differ from vLLM)
args = dict(DEFAULT_ARGS)
# Config file values sit above defaults but below env vars
args.update(_get_args_from_config_file())
# Auto-discover: every AsyncEngineArgs field from env UPPERCASED (e.g. MAX_MODEL_LEN)
args.update(_get_args_from_env_auto_discover())
@@ -517,4 +599,15 @@ def get_engine_args():
if speculative_config:
args["speculative_config"] = speculative_config
# Resolve lowercase HF cache paths (FDE-174)
if args.get("model"):
original_model = args["model"]
args["model"] = _resolve_cached_model_path(original_model)
# When the model was rewritten to an on-disk snapshot path, keep serving
# under the original repo id so the OpenAI API model name does not become
# a filesystem path (issue #310). An explicit served_model_name (or the
# OPENAI_SERVED_MODEL_NAME_OVERRIDE handled downstream) still wins.
if args["model"] != original_model and not args.get("served_model_name"):
args["served_model_name"] = original_model
return AsyncEngineArgs(**args)
+4 -2
View File
@@ -1,10 +1,12 @@
from transformers import AutoTokenizer
import logging
import os
from typing import Union
from transformers import AutoTokenizer
class TokenizerWrapper:
def __init__(self, tokenizer_name_or_path, tokenizer_revision, trust_remote_code):
print(f"tokenizer_name_or_path: {tokenizer_name_or_path}, tokenizer_revision: {tokenizer_revision}, trust_remote_code: {trust_remote_code}")
logging.debug("tokenizer_name_or_path: %s, tokenizer_revision: %s, trust_remote_code: %s", tokenizer_name_or_path, tokenizer_revision, trust_remote_code)
self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name_or_path, revision=tokenizer_revision or "main", trust_remote_code=trust_remote_code)
self.custom_chat_template = os.getenv("CUSTOM_CHAT_TEMPLATE")
self.has_chat_template = bool(self.tokenizer.chat_template) or bool(self.custom_chat_template)
+105
View File
@@ -0,0 +1,105 @@
"""Shared test fixtures.
``src/engine_args.py`` hard-imports ``vllm`` (and a tensorizer submodule) and
``torch.cuda``. Both are only installed inside the GPU Docker image, so when the
tests run on a machine without them we install lightweight stubs. When the real
packages *are* available (e.g. CI inside the worker image) the stubs are skipped
and the real ones are used instead.
"""
import sys
import types
from dataclasses import dataclass
from typing import Optional, Union, List
def _install_torch_stub():
try:
import torch # noqa: F401
return # real torch present, nothing to stub
except Exception:
pass
torch = types.ModuleType("torch")
cuda = types.ModuleType("torch.cuda")
# No GPU in the test environment -> 0 devices (skips tensor-parallel setup).
cuda.device_count = lambda: 0
torch.cuda = cuda
sys.modules["torch"] = torch
sys.modules["torch.cuda"] = cuda
def _install_vllm_stub():
try:
import vllm # noqa: F401
return # real vLLM present, nothing to stub
except Exception:
pass
vllm = types.ModuleType("vllm")
@dataclass
class AsyncEngineArgs:
# Only the fields the worker actually sets/reads need to exist here;
# get_engine_args() filters args down to AsyncEngineArgs.__dataclass_fields__
# before construction, so unknown keys are dropped rather than passed.
model: Optional[str] = None
served_model_name: Optional[Union[str, List[str]]] = None
revision: Optional[str] = None
tokenizer: Optional[str] = None
trust_remote_code: bool = False
max_model_len: Optional[int] = None
max_num_batched_tokens: Optional[int] = None
disable_log_stats: bool = False
gpu_memory_utilization: float = 0.9
tensor_parallel_size: int = 1
max_parallel_loading_workers: Optional[int] = None
kv_cache_dtype: Optional[str] = None
class _Stub: # pragma: no cover - placeholder for vllm symbols
def __init__(self, *args, **kwargs):
pass
vllm.AsyncEngineArgs = AsyncEngineArgs
vllm.SamplingParams = _Stub
sys.modules["vllm"] = vllm
# src.utils imports these at module load and uses ErrorResponse as a return
# annotation, which Python evaluates eagerly on <3.14 -> must be defined.
vllm_utils = types.ModuleType("vllm.utils")
vllm_utils.random_uuid = lambda: "stub-uuid"
vllm.utils = vllm_utils
sys.modules["vllm.utils"] = vllm_utils
protocol = types.ModuleType("vllm.entrypoints.openai.engine.protocol")
protocol.ErrorResponse = _Stub
protocol.ErrorInfo = _Stub
protocol.RequestResponseMetadata = _Stub
for name in (
"vllm.entrypoints",
"vllm.entrypoints.openai",
"vllm.entrypoints.openai.engine",
):
sys.modules.setdefault(name, types.ModuleType(name))
sys.modules["vllm.entrypoints.openai.engine.protocol"] = protocol
# vllm.model_executor.model_loader.tensorizer.TensorizerConfig
model_executor = types.ModuleType("vllm.model_executor")
model_loader = types.ModuleType("vllm.model_executor.model_loader")
tensorizer = types.ModuleType("vllm.model_executor.model_loader.tensorizer")
class TensorizerConfig: # pragma: no cover - placeholder
def __init__(self, *args, **kwargs):
pass
tensorizer.TensorizerConfig = TensorizerConfig
model_loader.tensorizer = tensorizer
model_executor.model_loader = model_loader
vllm.model_executor = model_executor
sys.modules["vllm.model_executor"] = model_executor
sys.modules["vllm.model_executor.model_loader"] = model_loader
sys.modules["vllm.model_executor.model_loader.tensorizer"] = tensorizer
_install_torch_stub()
_install_vllm_stub()
+6
View File
@@ -0,0 +1,6 @@
# Test-only dependencies. vllm/torch are stubbed in conftest.py when absent,
# so the unit tests run on a plain CPU runner without the GPU image.
pytest>=8,<10
# get_engine_args() reads a vLLM-style config via PyYAML (a transitive vllm dep
# at runtime); install it explicitly here since vllm itself is stubbed.
pyyaml
+126
View File
@@ -0,0 +1,126 @@
"""Tests for HF cache path resolution and served-model-name decoupling.
Regression coverage for issue #310: when MODEL_NAME is served from a lowercased
HF cache dir, the cache resolver rewrites engine_args.model to a snapshot path.
The served model name must stay the original repo id, not the path.
"""
import os
import pytest
from src import engine_args
from src.engine_args import _resolve_cached_model_path, get_engine_args
MODEL = "Qwen/Qwen3.6-27B-FP8"
SNAPSHOT_HASH = "e89b16ebf1988b3d6befa7de50abc2d76f26eb09"
def _make_cache(root, folder_name, snapshot=SNAPSHOT_HASH):
"""Create a HF-style ``models--…/snapshots/<hash>/`` dir and return its path."""
snap_dir = os.path.join(root, folder_name, "snapshots", snapshot)
os.makedirs(snap_dir)
return snap_dir
def _is_case_sensitive_fs(path):
"""The lowercase-cache resolution only matters on case-sensitive filesystems.
On macOS (APFS, case-insensitive by default) ``models--Qwen--…`` and
``models--qwen--…`` collide, so the resolver always sees the exact-case dir
as present. Production runs on Linux (case-sensitive), which is what these
tests exercise.
"""
probe = os.path.join(path, "CaseProbe")
open(probe, "w").close()
try:
return not os.path.exists(os.path.join(path, "caseprobe"))
finally:
os.remove(probe)
requires_case_sensitive_fs = pytest.mark.skipif(
not _is_case_sensitive_fs(os.environ.get("TMPDIR", "/tmp")),
reason="lowercase HF cache resolution only applies on case-sensitive filesystems",
)
@pytest.fixture
def hf_cache(tmp_path, monkeypatch):
cache = tmp_path / "hub"
cache.mkdir()
monkeypatch.setenv("HUGGINGFACE_HUB_CACHE", str(cache))
# Make sure HF_HOME does not shadow the explicit cache dir during the test.
monkeypatch.delenv("HF_HOME", raising=False)
return cache
class TestResolveCachedModelPath:
def test_exact_case_dir_returns_repo_id(self, hf_cache):
_make_cache(str(hf_cache), "models--Qwen--Qwen3.6-27B-FP8")
assert _resolve_cached_model_path(MODEL) == MODEL
def test_no_cache_returns_repo_id(self, hf_cache):
assert _resolve_cached_model_path(MODEL) == MODEL
def test_absolute_path_passthrough(self, hf_cache):
path = "/runpod-volume/some/local/model"
assert _resolve_cached_model_path(path) == path
@requires_case_sensitive_fs
def test_lowercase_dir_returns_snapshot_path(self, hf_cache):
snap = _make_cache(str(hf_cache), "models--qwen--qwen3.6-27b-fp8")
assert _resolve_cached_model_path(MODEL) == snap
def test_lowercase_dir_without_snapshots_returns_repo_id(self, hf_cache):
# Dir exists but has no snapshots subdir -> nothing to resolve to.
os.makedirs(os.path.join(str(hf_cache), "models--qwen--qwen3.6-27b-fp8"))
assert _resolve_cached_model_path(MODEL) == MODEL
@requires_case_sensitive_fs
def test_lowercase_dir_picks_latest_snapshot(self, hf_cache):
folder = "models--qwen--qwen3.6-27b-fp8"
_make_cache(str(hf_cache), folder, snapshot="aaaa")
latest = _make_cache(str(hf_cache), folder, snapshot="zzzz")
assert _resolve_cached_model_path(MODEL) == latest
class TestGetEngineArgsServedName:
"""Issue #310: served name must be decoupled from the resolved on-disk path."""
@pytest.fixture(autouse=True)
def base_env(self, monkeypatch):
# Avoid the network branch in _resolve_max_model_len.
monkeypatch.setenv("MAX_NUM_BATCHED_TOKENS", "2048")
monkeypatch.delenv("SERVED_MODEL_NAME", raising=False)
# Don't pick up a stray vLLM config file from the environment.
monkeypatch.setenv("VLLM_CONFIG_FILE", "/nonexistent-vllm-config.yaml")
@requires_case_sensitive_fs
def test_served_name_is_repo_id_when_path_rewritten(self, hf_cache, monkeypatch):
snap = _make_cache(str(hf_cache), "models--qwen--qwen3.6-27b-fp8")
monkeypatch.setenv("MODEL_NAME", MODEL)
result = get_engine_args()
assert result.model == snap # weights load from the lowercase cache
assert result.served_model_name == MODEL # API still serves the repo id
def test_served_name_untouched_when_no_rewrite(self, hf_cache, monkeypatch):
_make_cache(str(hf_cache), "models--Qwen--Qwen3.6-27B-FP8")
monkeypatch.setenv("MODEL_NAME", MODEL)
result = get_engine_args()
assert result.model == MODEL
assert result.served_model_name is None
def test_explicit_served_name_not_overridden(self, hf_cache, monkeypatch):
_make_cache(str(hf_cache), "models--qwen--qwen3.6-27b-fp8")
monkeypatch.setenv("MODEL_NAME", MODEL)
monkeypatch.setenv("SERVED_MODEL_NAME", "custom-name")
result = get_engine_args()
assert result.served_model_name == "custom-name"