Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):
1. tests.json allowedCudaVersions: 12.x → 13.0
The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
(#288, #289), but tests.json was still pinned to 12.5–12.9, so
the test pod was scheduled on a GPU with driver < 13.0 and
container init failed at the nvidia-container-cli hook with
"unsatisfied condition: cuda>=13.0".
2. requirements.txt kernels<0.15
huggingface/kernels v0.15.1 tightened LayerRepository to require
a revision or version argument
(https://github.com/huggingface/kernels/pull/544). transformers
>=5 still constructs LayerRepository(repo_id=..., layer_name=...)
without either, so worker import raised ValueError during
`from transformers import ...`. 0.14.1 is the last safe release.
3. tests.json timeout 30000 → 300000
vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
for SmolLM2-135M takes ~60–70s before the first request can be
served. The previous 30s per-test timeout fired before the
worker came up, producing "context cancelled or timed out:
context deadline exceeded" for every test even when the worker
was healthy. 300s gives enough headroom for cold start + the
actual inference call.
Refs: DR-1161
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
except blocks of _handle_responses_request and _handle_messages_request;
emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
* fix: update CUDA to 12.4.1 for Blackwell GPU support
- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json
This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.
Fixes: DR-1118
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* revert: remove NVIDIA B200 from default gpuIds
The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* fix: remove FlashInfer to avoid JIT compilation errors
FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.
vLLM will use its built-in fallback sampling methods instead.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment
- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions
refs: AE-1452