after rebasing onto main, get_engine_args() loads a vllm-style config via
PyYAML (a transitive vllm dep). vllm is stubbed in tests, so add pyyaml
explicitly and point VLLM_CONFIG_FILE at a nonexistent path so no stray
config is picked up.
src.utils uses ErrorResponse (a vllm import) as a module-level return
annotation, evaluated eagerly on python <3.14. the vllm stub lacked it,
so collection failed with NameError on ci (py3.11) while passing locally
(py3.14, lazy annotations). add the missing vllm.utils / protocol /
SamplingParams symbols to the stub.
the fde-174 cache resolver rewrites engine_args.model to an on-disk
snapshot path when the model is found only under a lowercased hf cache
dir. the openai served model name is derived from engine_args.model, so
it silently became the filesystem path and requests using the real repo
id returned 404.
set served_model_name to the original repo id whenever the model is
rewritten to a path, unless an explicit served name (or
OPENAI_SERVED_MODEL_NAME_OVERRIDE) is provided.
add the first python tests in the repo (tests/) covering the cache-path
resolution and served-name decoupling, plus a Tests github workflow that
runs pytest on prs and pushes to main. vllm/torch are stubbed when absent
so the suite runs on a plain cpu runner.
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):
1. tests.json allowedCudaVersions: 12.x → 13.0
The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
(#288, #289), but tests.json was still pinned to 12.5–12.9, so
the test pod was scheduled on a GPU with driver < 13.0 and
container init failed at the nvidia-container-cli hook with
"unsatisfied condition: cuda>=13.0".
2. requirements.txt kernels<0.15
huggingface/kernels v0.15.1 tightened LayerRepository to require
a revision or version argument
(https://github.com/huggingface/kernels/pull/544). transformers
>=5 still constructs LayerRepository(repo_id=..., layer_name=...)
without either, so worker import raised ValueError during
`from transformers import ...`. 0.14.1 is the last safe release.
3. tests.json timeout 30000 → 300000
vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
for SmolLM2-135M takes ~60–70s before the first request can be
served. The previous 30s per-test timeout fired before the
worker came up, producing "context cancelled or timed out:
context deadline exceeded" for every test even when the worker
was healthy. 300s gives enough headroom for cold start + the
actual inference call.
Refs: DR-1161
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
except blocks of _handle_responses_request and _handle_messages_request;
emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
* fix: update CUDA to 12.4.1 for Blackwell GPU support
- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json
This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.
Fixes: DR-1118
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* revert: remove NVIDIA B200 from default gpuIds
The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* fix: remove FlashInfer to avoid JIT compilation errors
FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.
vLLM will use its built-in fallback sampling methods instead.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment
- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions
refs: AE-1452