Commit Graph
559 Commits
Author SHA1 Message Date
Tim Pietrusky f299204770 docs: fix anthropic messages path and missing comma in routes list 2026-04-23 10:54:24 +02:00
velaraptor-runpod 1ed25eea20 update readme with responses and messages routes 2026-04-16 15:30:30 -05:00
velaraptor-runpod dc4ad7ddeb fix: kv_transfer_config is dataclass not json, fix for env var 2026-04-09 14:33:11 -05:00
velaraptor-runpod 9035b0e07f fix: lmcache version 2026-04-09 14:32:19 -05:00
velaraptor-runpod c979f0020f update readmes on TRANSFORMERS_VERSION 2026-04-06 16:54:08 -05:00
velaraptor-runpod 6fbd480a26 fix: allow for transformers_version 2026-04-06 16:49:32 -05:00
velaraptor-runpod 3ef1fb8e7b requested changes 2026-04-03 15:53:40 -05:00
velaraptor-runpod 30e8514d63 Runpod not RunPod 2026-03-17 20:59:52 -05:00
velaraptor-runpod d9808815ee feat: update requirements.txt 2026-03-17 19:56:39 -05:00
velaraptor-runpod d8ed3b5353 feat: update 0.16.0, add lmcache 2026-03-17 19:54:11 -05:00
velaraptor-runpod 4c4e039565 feat: add messages route for anthropic/claude 2026-03-17 19:51:50 -05:00
chrisvelaandGitHub 9d1686960d Merge pull request #273 from runpod-workers/bug/hf-overides-rope-scaling
bug: fix rope scaling to be forward compatible from hf_overrides
2026-03-10 11:21:44 -05:00
velaraptor-runpod 45d1eeee47 bug: fix rope scaling to be forward compatible from hf_overrides 2026-03-06 15:34:11 -06:00
chrisvelaandGitHub 17efb0e7d0 Merge pull request #272 from runpod-workers/feat/vllm-0.16.0
Release / release (push) Waiting to run
feat: Update to 0.16.0
v2.14.0
2026-03-05 13:06:45 -06:00
velaraptor-runpod 2b5f07df63 feat: Update to 0.16.0, remove NUM_GPU_BLOCKS_OVERRIDE in hub default since 0 will break 2026-03-04 16:38:40 -06:00
chrisvelaandGitHub 13fa71878e Merge pull request #269 from runpod-workers/feat/allow-engine-args-env
Release / release (push) Waiting to run
feat: allow all AsyncEngineArgs as env vars
v2.13.1
2026-02-27 15:34:23 -06:00
velaraptor-runpod 8a9365bed4 remove DEFAULT_ARGS that are none, fix MAX_CONTEXT_LEN_TO_CAPTURE 2026-02-27 14:04:15 -06:00
velaraptor-runpod cd485a1af1 update readme 2026-02-25 22:57:17 -06:00
velaraptor-runpod b9043639e9 requested changes/refactor 2026-02-25 16:07:38 -06:00
chrisvelaandGitHub 407dbd7773 Merge pull request #270 from runpod-workers/feat/update-vllm-v0.15.1
feat: update vllm to 0.15.1
2026-02-25 15:35:58 -06:00
velaraptor-runpod f103c142c1 feat: update vllm to 0.15.1 2026-02-24 17:44:49 -06:00
velaraptor-runpod efb093e198 add as VLLM_RUNPOD prefix and update readme 2026-02-24 17:37:57 -06:00
velaraptor-runpod 42443f735e feat: allow engine args through VLLM_ and checks the engine args 2026-02-24 16:05:18 -06:00
chrisvelaandGitHub b7c6d4f9a2 feat: update dockerfile to 12.9.1 (#267)
Release / release (push) Waiting to run
* feat: update dockerfile to 12.9.1

* update readme on VLLM_NIGHTLY build arg
v2.13.0
2026-02-19 10:13:14 +01:00
chrisvelaandGitHub d69cc021e8 Merge pull request #268 from runpod-workers/fix/spec-config-0-to-none
Release / release (push) Waiting to run
fix: spec config env vars should be none if zero
v2.12.3
2026-02-18 15:51:51 -06:00
velaraptor-runpod 61faa8f137 fix: spec config env vars should be none if zero 2026-02-18 15:41:19 -06:00
chrisvelaandGitHub 1606cff557 Merge pull request #265 from runpod-workers/fix/zero-max-model-num_batches
Release / release (push) Waiting to run
fix: check for zero param and set to None
v2.12.2
2026-02-13 15:26:06 -06:00
velaraptor-runpod e705c9494b fix: check for zero param and set to None 2026-02-13 15:23:54 -06:00
chrisvelaandGitHub b749aa5718 Merge pull request #264 from runpod-workers/fix/max_num_batched_tokens
Release / release (push) Waiting to run
fix: max num batched tokens
v2.12.1
2026-02-13 12:38:01 -06:00
velaraptor-runpod 4705ba8a7c fix: check max_num_batched_tokenz if max_model_len not set 2026-02-13 03:29:52 -06:00
velaraptor-runpod 767c66c301 make minimal changes 2026-02-13 03:23:44 -06:00
velaraptor-runpod fefdbe21a9 update changes 2026-02-13 03:16:43 -06:00
velaraptor-runpod ee961ad28d Update hub.json 2026-02-13 03:08:19 -06:00
velaraptor-runpod 2e8c251447 Merge branch 'main' into feat/update-vllm-v0.15.0 2026-02-13 03:01:05 -06:00
velaraptor-runpod c3cf43b228 Update hub.json 2026-02-13 00:22:16 -06:00
velaraptor-runpod 7ec10b98cd Update utils.py 2026-02-12 15:28:31 -06:00
c45ac42acd vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes

* MAX_NUM_BATCHED_TOKENS fix and CUDA tester

* Sys kill worker instead of marking as failed

* upgrade to vllm 0.12.0

* Update to vllm 0.15.0 and lora fix

* Update for HUB and removal of deprected env variables

* reverted docker-bake changes

* removed leftovers

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/utils.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Clean up of docs and comments in code

* nit: lowercase p

* nit: lowercase p

---------

Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
v2.12.0
2026-02-12 21:50:34 +01:00
velaraptor-runpod 340bc0b3c6 fix: served model name 2026-02-10 21:42:58 -06:00
velaraptor-runpod e1e9ef74ad add changes from pr 2026-02-06 18:10:09 -06:00
velaraptor-runpod 461f89cea6 add torch-c-dlpack-ext requirement 2026-02-06 17:03:39 -06:00
velaraptor-runpod 8eb55b90c1 add changes for v0.15.0 2026-02-05 17:24:16 -06:00
Tim PietruskyandGitHub 6d6cbe7095 fix: deactivate RunPod tests to fix hub release (#253)
Release / release (push) Waiting to run
Rename tests.json to tests_json to temporarily disable automated
tests while fixing the release on the hub.
v2.11.3
2026-01-22 18:06:36 +01:00
90c16b472d fix: update CUDA to 12.4.1 for Blackwell GPU support (#251)
Release / release (push) Waiting to run
* fix: update CUDA to 12.4.1 for Blackwell GPU support

- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json

This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.

Fixes: DR-1118

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* revert: remove NVIDIA B200 from default gpuIds

The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove FlashInfer to avoid JIT compilation errors

FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.

vLLM will use its built-in fallback sampling methods instead.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
v2.11.2
2026-01-13 22:01:36 +01:00
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
v2.11.1
2025-11-24 16:42:21 +01:00
chrisvelaandGitHub 3851d53f93 add ENABLE_EXPERT_PARALLEL engine arg for MoE models (#239)
Release / release (push) Waiting to run
* enable expert parallel arg for moe models

* add ENABLE_EXPERT_PARALLEL to hub config
v2.11.0
2025-11-17 19:25:19 +01:00
Witold WydmańskiandGitHub c896438f21 feat: bump transformers to allow Qwen3-VL (#225)
Release / release (push) Waiting to run
v2.10.0
2025-11-14 17:23:34 +01:00
Tim PietruskyandGitHub 912892f94e fix: remove space from gpuIds (#234) 2025-11-14 17:23:09 +01:00
Tim PietruskyandGitHub f8bf82469c fix(config): update allowed cuda versions in hub and tests config (#236)
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment

- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions

refs: AE-1452
2025-11-14 17:22:43 +01:00
Hailong YangandGitHub ec1664902b Merge pull request #230 from runpod-workers/feat/cse-853-vllm-template-params
Feat/cse 853 vllm template params
2025-10-31 13:32:34 -04:00
Eugene Klitenik d09122de4a remove un-needed 2025-10-29 13:06:33 -04:00