fix: make .runpod/tests.json hub tests pass on CUDA 13.0 (#299)
Release / release (push) Waiting to run

Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (https://github.com/huggingface/kernels/pull/544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
This commit is contained in:
Tim Pietrusky
2026-06-02 17:08:54 +02:00
committed by GitHub
parent 14b74a4989
commit dac05b62b3
2 changed files with 4 additions and 4 deletions
+3 -3
View File
@@ -5,7 +5,7 @@
"input": { "input": {
"prompt": "Write a short poem about artificial intelligence." "prompt": "Write a short poem about artificial intelligence."
}, },
"timeout": 30000 "timeout": 300000
}, },
{ {
"name": "openai_messages_test", "name": "openai_messages_test",
@@ -26,7 +26,7 @@
"temperature": 0.1 "temperature": 0.1
} }
}, },
"timeout": 30000 "timeout": 300000
} }
], ],
"config": { "config": {
@@ -38,6 +38,6 @@
"value": "HuggingFaceTB/SmolLM2-135M-Instruct" "value": "HuggingFaceTB/SmolLM2-135M-Instruct"
} }
], ],
"allowedCudaVersions": ["12.9", "12.8", "12.7", "12.6", "12.5"] "allowedCudaVersions": ["13.0"]
} }
} }
+1 -1
View File
@@ -11,5 +11,5 @@ pydantic-settings
hf-transfer hf-transfer
transformers>=5 transformers>=5
bitsandbytes>=0.45.0 bitsandbytes>=0.45.0
kernels kernels<0.15
torch-c-dlpack-ext torch-c-dlpack-ext