25 Commits
Author SHA1 Message Date
velaraptor-runpod 2b5f07df63 feat: Update to 0.16.0, remove NUM_GPU_BLOCKS_OVERRIDE in hub default since 0 will break 2026-03-04 16:38:40 -06:00
velaraptor-runpod 767c66c301 make minimal changes 2026-02-13 03:23:44 -06:00
velaraptor-runpod ee961ad28d Update hub.json 2026-02-13 03:08:19 -06:00
velaraptor-runpod 2e8c251447 Merge branch 'main' into feat/update-vllm-v0.15.0 2026-02-13 03:01:05 -06:00
velaraptor-runpod c3cf43b228 Update hub.json 2026-02-13 00:22:16 -06:00
c45ac42acd vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes

* MAX_NUM_BATCHED_TOKENS fix and CUDA tester

* Sys kill worker instead of marking as failed

* upgrade to vllm 0.12.0

* Update to vllm 0.15.0 and lora fix

* Update for HUB and removal of deprected env variables

* reverted docker-bake changes

* removed leftovers

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/utils.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Clean up of docs and comments in code

* nit: lowercase p

* nit: lowercase p

---------

Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
2026-02-12 21:50:34 +01:00
velaraptor-runpod e1e9ef74ad add changes from pr 2026-02-06 18:10:09 -06:00
chrisvelaandGitHub 3851d53f93 add ENABLE_EXPERT_PARALLEL engine arg for MoE models (#239)
Release / release (push) Waiting to run
* enable expert parallel arg for moe models

* add ENABLE_EXPERT_PARALLEL to hub config
2025-11-17 19:25:19 +01:00
Tim PietruskyandGitHub 912892f94e fix: remove space from gpuIds (#234) 2025-11-14 17:23:09 +01:00
Tim PietruskyandGitHub f8bf82469c fix(config): update allowed cuda versions in hub and tests config (#236)
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment

- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions

refs: AE-1452
2025-11-14 17:22:43 +01:00
Eugene Klitenik 205847471c reduce default container disk size to 150GB 2025-10-28 13:38:16 -04:00
Tim PietruskyandGitHub 66e1b1605b Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
2025-10-22 22:54:21 +02:00
max4c 2becd35345 Revert "fix: added back the HF_TOKEN (#219)"
Release / release (push) Waiting to run
This reverts commit 33d88df6c0.
2025-09-23 12:38:24 -07:00
33d88df6c0 fix: added back the HF_TOKEN (#219)
Release / release (push) Waiting to run
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-23 19:25:12 +02:00
Tim Pietrusky ecd562e112 fix: remove "access token" as this is handled by the platform
Release / release (push) Waiting to run
2025-09-19 21:00:47 +02:00
5cffaab8e8 docs: how to use the reasoning parser (#218)
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-09-17 07:49:37 +02:00
5f0fc69d75 feat: prepare worker-vllm for the hub (#214)
Release / release (push) Waiting to run
* docs: remove outdated video; remove old info; added missing config for tools

* ci: use proper release for dev (pr only) and production (release only)

* ci(hub): added openai example; use smollm2 as base model

* docs: added conventions to be able to work with ai ide's

* chore: remove outdated stuff

* chore: update copyright to 2025

* ci: added github permissions

* feat: added gpuIds, gputCount and allowedCudaVersions; removed default value for LOAD_FORMAT to check which influence this has on the ui

---------

Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
2025-08-28 10:09:27 +02:00
Tim Pietrusky f7514dea4b refactor: moved MODEL_NAME & HF_TOKEN out of advanced into the top section 2025-08-18 16:05:43 +02:00
Ezekiel WotringandGitHub 7cf7f3e4e6 Update hub.json 2025-04-03 09:04:28 -08:00
Marut PandyaandGitHub 4395d0c67b Update hub.json 2025-04-03 08:10:28 -07:00
Marut PandyaandGitHub 5864fa6843 Update hub.json 2025-03-31 15:55:35 -07:00
Marut PandyaandGitHub 19f25b17de Update hub.json 2025-03-31 15:14:40 -07:00
Marut PandyaandGitHub 70aa748c30 Update hub.json 2025-03-31 15:14:24 -07:00
Ezekiel WotringandGitHub 8005bcc1a8 fix description 2025-03-07 13:22:34 -09:00
Ezekiel WotringandGitHub aa7b00ddda Create hub.json 2025-03-07 13:20:52 -09:00