Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):
1. tests.json allowedCudaVersions: 12.x → 13.0
The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
(#288, #289), but tests.json was still pinned to 12.5–12.9, so
the test pod was scheduled on a GPU with driver < 13.0 and
container init failed at the nvidia-container-cli hook with
"unsatisfied condition: cuda>=13.0".
2. requirements.txt kernels<0.15
huggingface/kernels v0.15.1 tightened LayerRepository to require
a revision or version argument
(https://github.com/huggingface/kernels/pull/544). transformers
>=5 still constructs LayerRepository(repo_id=..., layer_name=...)
without either, so worker import raised ValueError during
`from transformers import ...`. 0.14.1 is the last safe release.
3. tests.json timeout 30000 → 300000
vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
for SmolLM2-135M takes ~60–70s before the first request can be
served. The previous 30s per-test timeout fired before the
worker came up, producing "context cancelled or timed out:
context deadline exceeded" for every test even when the worker
was healthy. 300s gives enough headroom for cold start + the
actual inference call.
Refs: DR-1161
* VLLM upgrade to 0.12.0 and compatibility fixes
* MAX_NUM_BATCHED_TOKENS fix and CUDA tester
* Sys kill worker instead of marking as failed
* upgrade to vllm 0.12.0
* Update to vllm 0.15.0 and lora fix
* Update for HUB and removal of deprected env variables
* reverted docker-bake changes
* removed leftovers
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Update src/utils.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Clean up of docs and comments in code
* nit: lowercase p
* nit: lowercase p
---------
Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment
- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions
refs: AE-1452
* feat: moved config into docs; added banner; auto detect "messages" in input
* docs: moved config into docs
* chore: added .DS_Store
* chore: get the original stuff working again
* chore: remove all changes
* docs: reduced toc and added small config table
---------
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
* docs: remove outdated video; remove old info; added missing config for tools
* ci: use proper release for dev (pr only) and production (release only)
* ci(hub): added openai example; use smollm2 as base model
* docs: added conventions to be able to work with ai ide's
* chore: remove outdated stuff
* chore: update copyright to 2025
* ci: added github permissions
* feat: added gpuIds, gputCount and allowedCudaVersions; removed default value for LOAD_FORMAT to check which influence this has on the ui
---------
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>