Commit Graph
34 Commits
Author SHA1 Message Date
Tim Pietrusky d7ba3b6ab7 test: install pyyaml and isolate vllm config file in tests
after rebasing onto main, get_engine_args() loads a vllm-style config via
PyYAML (a transitive vllm dep). vllm is stubbed in tests, so add pyyaml
explicitly and point VLLM_CONFIG_FILE at a nonexistent path so no stray
config is picked up.
2026-06-19 19:05:24 +02:00
Tim Pietrusky b11c91722c test: complete vllm stub so src.utils imports under py<3.14
src.utils uses ErrorResponse (a vllm import) as a module-level return
annotation, evaluated eagerly on python <3.14. the vllm stub lacked it,
so collection failed with NameError on ci (py3.11) while passing locally
(py3.14, lazy annotations). add the missing vllm.utils / protocol /
SamplingParams symbols to the stub.
2026-06-19 19:04:45 +02:00
Tim Pietrusky fcdc799e0d fix: serve original model name when hf cache dir is lowercased (#310)
the fde-174 cache resolver rewrites engine_args.model to an on-disk
snapshot path when the model is found only under a lowercased hf cache
dir. the openai served model name is derived from engine_args.model, so
it silently became the filesystem path and requests using the real repo
id returned 404.

set served_model_name to the original repo id whenever the model is
rewritten to a path, unless an explicit served name (or
OPENAI_SERVED_MODEL_NAME_OVERRIDE) is provided.

add the first python tests in the repo (tests/) covering the cache-path
resolution and served-name decoupling, plus a Tests github workflow that
runs pytest on prs and pushes to main. vllm/torch are stubbed when absent
so the suite runs on a plain cpu runner.
2026-06-19 19:04:45 +02:00
Tim PietruskyandGitHub dac05b62b3 fix: make .runpod/tests.json hub tests pass on CUDA 13.0 (#299)
Release / release (push) Waiting to run
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (https://github.com/huggingface/kernels/pull/544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
2026-06-02 17:08:54 +02:00
Tim PietruskyandGitHub 14b74a4989 chore: re-enable .runpod/tests.json hub tests (#295)
Release / release (push) Waiting to run
Rename tests_json back to tests.json to re-enable the automated hub
tests that were temporarily disabled in #253.

Refs: DR-1161
2026-06-01 17:52:23 +02:00
Tim Pietrusky 3403889528 fix: address review comments on responses/messages handlers and lmcache guard
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
  except blocks of _handle_responses_request and _handle_messages_request;
  emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
  blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
  actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
  and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
2026-04-23 11:28:50 +02:00
Tim Pietrusky f299204770 docs: fix anthropic messages path and missing comma in routes list 2026-04-23 10:54:24 +02:00
Tim PietruskyandGitHub 6d6cbe7095 fix: deactivate RunPod tests to fix hub release (#253)
Release / release (push) Waiting to run
Rename tests.json to tests_json to temporarily disable automated
tests while fixing the release on the hub.
2026-01-22 18:06:36 +01:00
90c16b472d fix: update CUDA to 12.4.1 for Blackwell GPU support (#251)
Release / release (push) Waiting to run
* fix: update CUDA to 12.4.1 for Blackwell GPU support

- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json

This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.

Fixes: DR-1118

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* revert: remove NVIDIA B200 from default gpuIds

The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove FlashInfer to avoid JIT compilation errors

FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.

vLLM will use its built-in fallback sampling methods instead.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-13 22:01:36 +01:00
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
Tim PietruskyandGitHub 912892f94e fix: remove space from gpuIds (#234) 2025-11-14 17:23:09 +01:00
Tim PietruskyandGitHub f8bf82469c fix(config): update allowed cuda versions in hub and tests config (#236)
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment

- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions

refs: AE-1452
2025-11-14 17:22:43 +01:00
Tim PietruskyandGitHub 6337a6673a fix: allow also CUDA 12.8 & 12.9 (#228)
Release / release (push) Waiting to run
2025-10-24 18:48:26 +02:00
Tim PietruskyandGitHub 66e1b1605b Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
2025-10-22 22:54:21 +02:00
Tim PietruskyandGitHub 60c8f257a8 Merge pull request #227 from runpod-workers/chore/vllm-0.11.0
chore: update vllm to 0.11.0
2025-10-22 22:53:51 +02:00
Tim Pietrusky fae16e7ee1 chore: update vllm to 0.11.0 2025-10-22 13:40:53 -07:00
Tim Pietrusky ecd562e112 fix: remove "access token" as this is handled by the platform
Release / release (push) Waiting to run
2025-09-19 21:00:47 +02:00
Tim Pietrusky 0e0d6df859 docs: updated example-tag for dev and release 2025-08-28 14:28:19 +02:00
Tim Pietrusky d1718aec00 ci: removed "github release" step as that is not needed 2025-08-28 14:27:50 +02:00
Tim Pietrusky f7514dea4b refactor: moved MODEL_NAME & HF_TOKEN out of advanced into the top section 2025-08-18 16:05:43 +02:00
Tim Pietrusky 121a3dd44b fix: parse value for RAW_OPENAI_OUTPUT correctly 2025-08-13 12:20:44 +02:00
Tim Pietrusky e2e111b942 fix: allow "None" as string for setting env variables (like quantization) 2025-08-13 10:27:00 +02:00
Tim Pietrusky 0133c23be8 ci: added manual workflow trigger for releases 2025-08-04 09:58:23 +02:00
Tim Pietrusky a129cff47d docs: use "version" instead of actual version, so that people can check the releases 2025-07-31 12:03:17 +02:00
Tim Pietrusky 30f2c4630e refactor: use correct version 2025-07-31 12:02:46 +02:00
Tim Pietrusky b98636e432 feat: added "dev" and "release" workflows; removed "vllm-base-image" as it's not needed 2025-07-28 16:40:20 +02:00
Tim Pietrusky 46a3300cbb chore: reverted changeds to only focus on vllm update 2025-06-20 14:56:39 +02:00
Tim Pietrusky 1e9a731380 Fix Mistral tokenizer initialization: let vLLM handle tokenizer for mistral models 2025-06-14 15:31:56 +02:00
Tim Pietrusky c47e649a24 Add CONFIG_FORMAT environment variable support 2025-06-14 15:12:46 +02:00
Tim Pietrusky 7192bcaeef Trigger build automatically on feat/0.9.1 branch 2025-06-14 13:49:22 +02:00
Tim Pietrusky 8665ffb78d Revert workflow back to original configuration 2025-06-14 13:47:16 +02:00
Tim Pietrusky 57431b30ad Fix workflow: use standard GitHub runners and actions 2025-06-14 13:42:46 +02:00
Tim Pietrusky 437a84c77a ci: added workflow to build the image 2025-06-14 13:31:36 +02:00
Tim Pietrusky a4062fc488 feat: update to 0.9.1 2025-06-14 13:31:27 +02:00