• fix: make .runpod/tests.json hub tests pass on CUDA 13.0 (#299)

    owen released this 2026-06-02 11:08:54 -04:00

    Three coupled fixes verified end-to-end on a private fork
    (TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

    1. tests.json allowedCudaVersions: 12.x → 13.0
      The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
      (#288, #289), but tests.json was still pinned to 12.5–12.9, so
      the test pod was scheduled on a GPU with driver < 13.0 and
      container init failed at the nvidia-container-cli hook with
      "unsatisfied condition: cuda>=13.0".

    2. requirements.txt kernels<0.15
      huggingface/kernels v0.15.1 tightened LayerRepository to require
      a revision or version argument
      (https://github.com/huggingface/kernels/pull/544). transformers

      =5 still constructs LayerRepository(repo_id=..., layer_name=...)
      without either, so worker import raised ValueError during
      from transformers import .... 0.14.1 is the last safe release.

    3. tests.json timeout 30000 → 300000
      vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
      for SmolLM2-135M takes ~60–70s before the first request can be
      served. The previous 30s per-test timeout fired before the
      worker came up, producing "context cancelled or timed out:
      context deadline exceeded" for every test even when the worker
      was healthy. 300s gives enough headroom for cold start + the
      actual inference call.

    Refs: DR-1161

    Downloads