Commit Graph
607 Commits
Author SHA1 Message Date
chrisvelaandGitHub 76f345be08 Merge pull request #321 from runpod-workers/revert-318-feat/0.24.0
Release / release (push) Waiting to run
Revert "feat: update to 0.24.0"
v2.23.0
2026-07-24 15:23:50 -05:00
chrisvelaandGitHub 4dfda80fd0 Revert "feat: update to 0.24.0" 2026-07-24 15:16:55 -05:00
Dean QuiñanolaandGitHub 89f25b0d70 Merge pull request #319 from runpod-workers/runpod-package-update
chore: update runpod to 1.11.0
2026-07-24 12:13:02 -07:00
deanqandgithub-actions[bot] 4808160735 chore: update runpod to 1.11.0 2026-07-24 18:30:38 +00:00
Dean QuiñanolaandGitHub 51e1aee770 Merge pull request #318 from runpod-workers/feat/0.24.0
feat: update to 0.24.0
2026-07-24 11:30:29 -07:00
velaraptor-runpod 86b7951c39 feat: update to 0.24.0 2026-07-16 18:33:38 -05:00
chrisvelaandGitHub e786204cf5 Merge pull request #316 from runpod-workers/docs/sync-vllm-version-0.23.0
docs: sync vLLM version to 0.23.0 in READMEs
2026-07-10 11:10:53 -05:00
chrisvelaandGitHub 591f9d6531 Merge pull request #317 from runpod-workers/fix/test-sls-gh-action
fix: fix image to test with sls in gh actions
2026-07-10 09:55:53 -05:00
velaraptor-runpod 49967d0c09 fix: fix image to test with sls in gh actions 2026-07-09 22:25:40 -05:00
velaraptor-runpodandgithub-actions[bot] e97cab854f docs: sync vLLM version to 0.23.0 in READMEs 2026-07-10 02:18:51 +00:00
chrisvelaandGitHub 828e797498 Merge pull request #315 from runpod-workers/feat/0.23.0
feat: update vllm to 0.23.0 + add sls e2e model testing
2026-07-09 21:18:41 -05:00
velaraptor-runpod 373847d9f9 fix: pin models to version, except for gpt-oss 2026-07-09 19:21:24 -05:00
velaraptor-runpod 2000acc0a3 chore: add timeout for e2e serverless 2026-07-09 19:15:19 -05:00
velaraptor-runpod 89e3b2e920 fix: fix for huggingface, secret keys 2026-07-09 18:56:58 -05:00
velaraptor-runpod 85f8c81c1c add sls e2e tests 2026-07-09 18:22:57 -05:00
velaraptor-runpod 88490e9551 feat: update to 0.23.0 2026-07-09 18:22:40 -05:00
velaraptor-runpod 798f2b12eb feat: update to 0.23.0 2026-07-01 16:19:01 -05:00
chrisvelaandGitHub 9e1c483136 Merge pull request #311 from runpod-workers/t3code/cd56c69d
Release / release (push) Waiting to run
fix: serve original model name when HF cache dir is lowercased (#310)
v2.22.5
2026-06-26 13:48:13 -05:00
chrisvelaandGitHub d71ea9939d Merge pull request #308 from runpod-workers/docs/sync-vllm-version-0.20.2
docs: sync vLLM version to 0.20.2 in READMEs
2026-06-26 13:47:59 -05:00
chrisvelaandGitHub 7e2b4e2288 Merge pull request #314 from runpod-workers/runpod-package-update
chore: update runpod to 1.10.0
2026-06-26 13:45:16 -05:00
deanqandgithub-actions[bot] 015f8f3c4d chore: update runpod to 1.10.0 2026-06-26 17:20:41 +00:00
Hailong YangandGitHub 75ffcf73f2 Merge pull request #309 from adithyaJRunpod/feature/tuned-configs-round2
added configs for  Gemma 4 31B and GPT-OSS 120B
2026-06-22 17:41:50 -04:00
Tim Pietrusky d7ba3b6ab7 test: install pyyaml and isolate vllm config file in tests
after rebasing onto main, get_engine_args() loads a vllm-style config via
PyYAML (a transitive vllm dep). vllm is stubbed in tests, so add pyyaml
explicitly and point VLLM_CONFIG_FILE at a nonexistent path so no stray
config is picked up.
2026-06-19 19:05:24 +02:00
Tim Pietrusky b11c91722c test: complete vllm stub so src.utils imports under py<3.14
src.utils uses ErrorResponse (a vllm import) as a module-level return
annotation, evaluated eagerly on python <3.14. the vllm stub lacked it,
so collection failed with NameError on ci (py3.11) while passing locally
(py3.14, lazy annotations). add the missing vllm.utils / protocol /
SamplingParams symbols to the stub.
2026-06-19 19:04:45 +02:00
Tim Pietrusky fcdc799e0d fix: serve original model name when hf cache dir is lowercased (#310)
the fde-174 cache resolver rewrites engine_args.model to an on-disk
snapshot path when the model is found only under a lowercased hf cache
dir. the openai served model name is derived from engine_args.model, so
it silently became the filesystem path and requests using the real repo
id returned 404.

set served_model_name to the original repo id whenever the model is
rewritten to a path, unless an explicit served name (or
OPENAI_SERVED_MODEL_NAME_OVERRIDE) is provided.

add the first python tests in the repo (tests/) covering the cache-path
resolution and served-name decoupling, plus a Tests github workflow that
runs pytest on prs and pushes to main. vllm/torch are stubbed when absent
so the suite runs on a plain cpu runner.
2026-06-19 19:04:45 +02:00
AdithyaJob 84ec446493 added configs for Gemma 4 31B and GPT-OSS 120B 2026-06-17 22:19:58 -07:00
velaraptor-runpodandgithub-actions[bot] db246653a2 docs: sync vLLM version to 0.20.2 in READMEs 2026-06-12 20:51:06 +00:00
chrisvelaandGitHub 1b3228a2dc Merge pull request #307 from runpod-workers/fix/revert-0.20.0
Release / release (push) Waiting to run
revert to 0.20.2
v2.22.4
2026-06-12 15:50:52 -05:00
velaraptor-runpod 4817d4a8e7 revert to 0.20.2 2026-06-12 15:49:23 -05:00
chrisvelaandGitHub 0378382a92 Merge pull request #306 from runpod-workers/revert/v2.20.1
Release / release (push) Waiting to run
chore: carry non-vllm changes from main (tests GPU + configs)
v2.22.3
2026-06-12 15:21:22 -05:00
velaraptor-runpodandClaude Sonnet 4.6 08580e7ccf chore: carry non-vllm changes from main (tests GPU + configs)
Brings forward the L40 GPU type in tests.json and the new llama/qwen
tuned config files, while keeping Dockerfile pinned at vllm 0.20.2
(v2.20.1 state).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-12 15:18:36 -05:00
chrisvelaandGitHub 8868aae6b1 Merge pull request #305 from runpod-workers/fix/typo-tests-l40
Release / release (push) Waiting to run
fix: fix typo in tests
v2.22.2
2026-06-12 14:00:54 -05:00
velaraptor-runpod fb8adc5c06 fix: fix typo in tests 2026-06-12 13:59:29 -05:00
chrisvelaandGitHub 7351da512b Merge pull request #304 from runpod-workers/fix/fix-tests
Release / release (push) Waiting to run
fix: change test gpu to L40
v2.22.1
2026-06-12 13:49:15 -05:00
velaraptor-runpod 1e78043b2c fix: change test gpu to L40 2026-06-12 13:46:33 -05:00
chrisvelaandGitHub 352c64f4c1 Merge pull request #303 from runpod-workers/chore/check-vllm-versions-readmes
chore: check readmes vllm version in sync with version in Dockerfile
2026-06-12 09:49:07 -05:00
velaraptor-runpod 0922f5b435 chore: check readmes vllm version in sync with version in Dockerfile 2026-06-11 17:33:16 -05:00
chrisvelaandGitHub 5d9a48fc70 Merge pull request #302 from runpod-workers/feat/0.22.1
Release / release (push) Waiting to run
feat: upgrade vllm to 0.22.1
v2.22.0
2026-06-11 17:27:38 -05:00
velaraptor-runpod 5d1579e361 fix: update readme links 2026-06-11 17:00:31 -05:00
velaraptor-runpod 0488b77d89 feat: upgrade vllm to 0.22.1 2026-06-11 16:29:02 -05:00
chrisvelaandGitHub d3a962c33b Merge pull request #301 from adithyaJRunpod/feature/tuned-configs
Release / release (push) Waiting to run
Add tuned configs,  CON-239
v2.21.1
2026-06-11 14:26:00 -05:00
chrisvelaandGitHub 105c125698 Merge pull request #300 from runpod-workers/feat/0.21.0
feat: upgrade vllm to 0.21.0
2026-06-11 12:47:08 -05:00
AdithyaJob 0a0ccfcb60 Add tuned configs for Llama 3.1 8B and Qwen3 8B 2026-06-10 21:07:41 -07:00
velaraptor-runpod c8ce53c72c fix: add kenels, and fix for cuda 2026-06-10 16:01:53 -05:00
velaraptor-runpod 9618e799ba chore: fix cuda libraries 2026-06-04 15:35:58 -05:00
velaraptor-runpod cb3f077dba feat: upgrade vllm to 0.21.0 2026-06-03 14:43:56 -05:00
chrisvelaandGitHub 69646b9e99 Merge pull request #294 from runpod-workers/feat/allow-config
Release / release (push) Waiting to run
feat: allow config.yaml like vllm serve
v2.20.1
2026-06-02 20:28:02 -05:00
Jacob CiparandGitHub 8b991a7ad7 Merge pull request #297 from runpod-workers/jhcipar/bump-runpod-python-version
feat: bump runpod-python version
2026-06-02 11:17:48 -04:00
Tim PietruskyandGitHub dac05b62b3 fix: make .runpod/tests.json hub tests pass on CUDA 13.0 (#299)
Release / release (push) Waiting to run
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (https://github.com/huggingface/kernels/pull/544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
v2.20.0
2026-06-02 17:08:54 +02:00
jhcipar d356c31675 feat: bump runpod-python version 2026-06-01 20:28:40 -04:00