Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):
1. tests.json allowedCudaVersions: 12.x → 13.0
The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
(#288, #289), but tests.json was still pinned to 12.5–12.9, so
the test pod was scheduled on a GPU with driver < 13.0 and
container init failed at the nvidia-container-cli hook with
"unsatisfied condition: cuda>=13.0".
2. requirements.txt kernels<0.15
huggingface/kernels v0.15.1 tightened LayerRepository to require
a revision or version argument
(https://github.com/huggingface/kernels/pull/544). transformers
>=5 still constructs LayerRepository(repo_id=..., layer_name=...)
without either, so worker import raised ValueError during
`from transformers import ...`. 0.14.1 is the last safe release.
3. tests.json timeout 30000 → 300000
vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
for SmolLM2-135M takes ~60–70s before the first request can be
served. The previous 30s per-test timeout fired before the
worker came up, producing "context cancelled or timed out:
context deadline exceeded" for every test even when the worker
was healthy. 300s gives enough headroom for cold start + the
actual inference call.
Refs: DR-1161
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
except blocks of _handle_responses_request and _handle_messages_request;
emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
* VLLM upgrade to 0.12.0 and compatibility fixes
* MAX_NUM_BATCHED_TOKENS fix and CUDA tester
* Sys kill worker instead of marking as failed
* upgrade to vllm 0.12.0
* Update to vllm 0.15.0 and lora fix
* Update for HUB and removal of deprected env variables
* reverted docker-bake changes
* removed leftovers
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Update src/utils.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Clean up of docs and comments in code
* nit: lowercase p
* nit: lowercase p
---------
Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7