src.utils uses ErrorResponse (a vllm import) as a module-level return
annotation, evaluated eagerly on python <3.14. the vllm stub lacked it,
so collection failed with NameError on ci (py3.11) while passing locally
(py3.14, lazy annotations). add the missing vllm.utils / protocol /
SamplingParams symbols to the stub.
the fde-174 cache resolver rewrites engine_args.model to an on-disk
snapshot path when the model is found only under a lowercased hf cache
dir. the openai served model name is derived from engine_args.model, so
it silently became the filesystem path and requests using the real repo
id returned 404.
set served_model_name to the original repo id whenever the model is
rewritten to a path, unless an explicit served name (or
OPENAI_SERVED_MODEL_NAME_OVERRIDE) is provided.
add the first python tests in the repo (tests/) covering the cache-path
resolution and served-name decoupling, plus a Tests github workflow that
runs pytest on prs and pushes to main. vllm/torch are stubbed when absent
so the suite runs on a plain cpu runner.
Brings forward the L40 GPU type in tests.json and the new llama/qwen
tuned config files, while keeping Dockerfile pinned at vllm 0.20.2
(v2.20.1 state).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):
1. tests.json allowedCudaVersions: 12.x → 13.0
The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
(#288, #289), but tests.json was still pinned to 12.5–12.9, so
the test pod was scheduled on a GPU with driver < 13.0 and
container init failed at the nvidia-container-cli hook with
"unsatisfied condition: cuda>=13.0".
2. requirements.txt kernels<0.15
huggingface/kernels v0.15.1 tightened LayerRepository to require
a revision or version argument
(https://github.com/huggingface/kernels/pull/544). transformers
>=5 still constructs LayerRepository(repo_id=..., layer_name=...)
without either, so worker import raised ValueError during
`from transformers import ...`. 0.14.1 is the last safe release.
3. tests.json timeout 30000 → 300000
vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
for SmolLM2-135M takes ~60–70s before the first request can be
served. The previous 30s per-test timeout fired before the
worker came up, producing "context cancelled or timed out:
context deadline exceeded" for every test even when the worker
was healthy. 300s gives enough headroom for cold start + the
actual inference call.
Refs: DR-1161