chrisvela and GitHub
8868aae6b1
Merge pull request #305 from runpod-workers/fix/typo-tests-l40
...
Release / release (push) Waiting to run
fix: fix typo in tests
v2.22.2
2026-06-12 14:00:54 -05:00
velaraptor-runpod
fb8adc5c06
fix: fix typo in tests
2026-06-12 13:59:29 -05:00
chrisvela and GitHub
7351da512b
Merge pull request #304 from runpod-workers/fix/fix-tests
...
Release / release (push) Waiting to run
fix: change test gpu to L40
v2.22.1
2026-06-12 13:49:15 -05:00
velaraptor-runpod
1e78043b2c
fix: change test gpu to L40
2026-06-12 13:46:33 -05:00
chrisvela and GitHub
352c64f4c1
Merge pull request #303 from runpod-workers/chore/check-vllm-versions-readmes
...
chore: check readmes vllm version in sync with version in Dockerfile
2026-06-12 09:49:07 -05:00
velaraptor-runpod
0922f5b435
chore: check readmes vllm version in sync with version in Dockerfile
2026-06-11 17:33:16 -05:00
chrisvela and GitHub
5d9a48fc70
Merge pull request #302 from runpod-workers/feat/0.22.1
...
Release / release (push) Waiting to run
feat: upgrade vllm to 0.22.1
v2.22.0
2026-06-11 17:27:38 -05:00
velaraptor-runpod
5d1579e361
fix: update readme links
2026-06-11 17:00:31 -05:00
velaraptor-runpod
0488b77d89
feat: upgrade vllm to 0.22.1
2026-06-11 16:29:02 -05:00
chrisvela and GitHub
d3a962c33b
Merge pull request #301 from adithyaJRunpod/feature/tuned-configs
...
Release / release (push) Waiting to run
Add tuned configs, CON-239
v2.21.1
2026-06-11 14:26:00 -05:00
chrisvela and GitHub
105c125698
Merge pull request #300 from runpod-workers/feat/0.21.0
...
feat: upgrade vllm to 0.21.0
2026-06-11 12:47:08 -05:00
AdithyaJob
0a0ccfcb60
Add tuned configs for Llama 3.1 8B and Qwen3 8B
2026-06-10 21:07:41 -07:00
velaraptor-runpod
c8ce53c72c
fix: add kenels, and fix for cuda
2026-06-10 16:01:53 -05:00
velaraptor-runpod
9618e799ba
chore: fix cuda libraries
2026-06-04 15:35:58 -05:00
velaraptor-runpod
cb3f077dba
feat: upgrade vllm to 0.21.0
2026-06-03 14:43:56 -05:00
chrisvela and GitHub
69646b9e99
Merge pull request #294 from runpod-workers/feat/allow-config
...
Release / release (push) Waiting to run
feat: allow config.yaml like vllm serve
v2.20.1
2026-06-02 20:28:02 -05:00
Jacob Cipar and GitHub
8b991a7ad7
Merge pull request #297 from runpod-workers/jhcipar/bump-runpod-python-version
...
feat: bump runpod-python version
2026-06-02 11:17:48 -04:00
Tim Pietrusky and GitHub
dac05b62b3
fix: make .runpod/tests.json hub tests pass on CUDA 13.0 ( #299 )
...
Release / release (push) Waiting to run
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):
1. tests.json allowedCudaVersions: 12.x → 13.0
The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
(#288 , #289 ), but tests.json was still pinned to 12.5–12.9, so
the test pod was scheduled on a GPU with driver < 13.0 and
container init failed at the nvidia-container-cli hook with
"unsatisfied condition: cuda>=13.0".
2. requirements.txt kernels<0.15
huggingface/kernels v0.15.1 tightened LayerRepository to require
a revision or version argument
(https://github.com/huggingface/kernels/pull/544 ). transformers
>=5 still constructs LayerRepository(repo_id=..., layer_name=...)
without either, so worker import raised ValueError during
`from transformers import ...`. 0.14.1 is the last safe release.
3. tests.json timeout 30000 → 300000
vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
for SmolLM2-135M takes ~60–70s before the first request can be
served. The previous 30s per-test timeout fired before the
worker came up, producing "context cancelled or timed out:
context deadline exceeded" for every test even when the worker
was healthy. 300s gives enough headroom for cold start + the
actual inference call.
Refs: DR-1161
v2.20.0
2026-06-02 17:08:54 +02:00
jhcipar
d356c31675
feat: bump runpod-python version
2026-06-01 20:28:40 -04:00
Tim Pietrusky and GitHub
14b74a4989
chore: re-enable .runpod/tests.json hub tests ( #295 )
...
Release / release (push) Waiting to run
Rename tests_json back to tests.json to re-enable the automated hub
tests that were temporarily disabled in #253 .
Refs: DR-1161
v2.19.1
v2.19.2
2026-06-01 17:52:23 +02:00
velaraptor-runpod
80072047ab
feat: allow config.yaml like vllm serve
2026-05-29 15:50:22 -05:00
chrisvela and GitHub
50aba8fb57
Merge pull request #293 from runpod-workers/fix/update-deep-gemm-hub-value
...
Release / release (push) Waiting to run
fix: update VLLM_USE_DEEP_GEMM hub to default to 0
v2.19.0
2026-05-27 11:51:48 -05:00
velaraptor-runpod
9edc5715ce
fix: update configuration.md
2026-05-27 11:33:37 -05:00
velaraptor-runpod
4c91f2c5b5
fix: update VLLM_USE_DEEP_GEMM hub to default to 0
2026-05-27 11:26:52 -05:00
chrisvela and GitHub
6265b99348
Merge pull request #288 from runpod-workers/feat/0.20.0
...
feat: update to 0.20.2
2026-05-26 17:17:45 -05:00
velaraptor-runpod
026f8d700b
fix: specify deepgemm commit version
2026-05-21 18:46:54 -05:00
chrisvela and GitHub
146bdb0252
Merge branch 'main' into feat/0.20.0
2026-05-21 16:04:39 -05:00
velaraptor-runpod
da01193a3d
chore: fix readme
2026-05-21 15:57:40 -05:00
velaraptor-runpod
c2e6cc9f61
chore: update readme with correct vllm version
2026-05-21 15:39:19 -05:00
velaraptor-runpod
69968a6b39
chore: fix logging, warning for text prompt
2026-05-21 15:34:09 -05:00
velaraptor-runpod
32b29d4c6c
fix: add deepgemm, update base image and hub for cuda 13.0
2026-05-20 17:43:42 -05:00
velaraptor-runpod
dcea4fc4f9
fix dockerfile
2026-05-15 12:08:31 -04:00
velaraptor-runpod
9c139e8ceb
update: update to 0.20.1, update dockerfile to cuda 13
2026-05-15 11:35:53 -04:00
velaraptor-runpod
678bb4be8f
feat: update to 0.20.1 for patch fixes
2026-05-07 11:58:22 -05:00
chrisvela and GitHub
87d7365126
Merge pull request #292 from runpod-workers/fix/open-ai
...
Release / release (push) Waiting to run
fix: fix warmup
v2.18.1
2026-05-01 18:06:11 -05:00
velaraptor-runpod
0e83616f93
fix: fix warmup
2026-05-01 17:39:51 -05:00
chrisvela and GitHub
ab6d39dcf8
Merge pull request #291 from runpod-workers/bug/281-hf-token
...
Release / release (push) Waiting to run
bug: fix hf-token being passed in engineargs
v2.18.0
2026-05-01 15:32:04 -05:00
velaraptor-runpod
ed315a175e
merge main
2026-05-01 15:28:05 -05:00
velaraptor-runpod
73f030ae5e
Merge branch 'main' into bug/281-hf-token
2026-05-01 14:33:51 -05:00
chrisvela and GitHub
8a099c1723
Merge pull request #287 from runpod-workers/feat/0.19.1
...
feat: update to 0.19.1
2026-05-01 14:29:39 -05:00
velaraptor-runpod
ff87840a58
fix: add enforce_eager as true, add pytorch_alloc_conf to expandle_segments to True for OOM, for hub defaults
2026-05-01 14:13:03 -05:00
velaraptor-runpod
7dc853b1fe
chore: remove release to trigger on release publish, just use tags. duplicate
2026-05-01 10:13:32 -05:00
velaraptor-runpod
a8b754b92a
merge main
2026-05-01 10:12:43 -05:00
chrisvela and GitHub
6357aeda51
Merge pull request #290 from runpod-workers/fix/fix-old-actions
...
Release / release (push) Waiting to run
fix: fix old github actions, trigger release on publish release
v2.17.0
2026-05-01 10:01:22 -05:00
velaraptor-runpod
0140b29c44
chore: add release notes to slack notification
2026-05-01 09:48:43 -05:00
velaraptor-runpod
0cb8aeae77
chore: add specific runpod version
2026-05-01 09:47:23 -05:00
chrisvela and GitHub
cd8f9e9560
Merge branch 'main' into fix/fix-old-actions
2026-04-30 21:57:45 -05:00
chrisvela and GitHub
895fd25fac
Merge pull request #286 from runpod-workers/feat/0.18.1
...
feat: update vllm to 0.18.1
2026-04-30 21:56:07 -05:00
velaraptor-runpod
7bb8df73af
bug: fix hf-token being passed in engineargs
2026-04-30 20:37:30 -05:00
velaraptor-runpod
747cdf5891
chore: update readme vllm version
2026-04-30 20:26:23 -05:00