chrisvela and GitHub
895fd25fac
Merge pull request #286 from runpod-workers/feat/0.18.1
...
feat: update vllm to 0.18.1
2026-04-30 21:56:07 -05:00
velaraptor-runpod
72547aa3bb
chore: update readme vllm version
2026-04-30 20:25:02 -05:00
chrisvela and GitHub
f49f35456e
Merge pull request #285 from runpod-workers/feat/0.17.1
...
feat: update vllm to 0.17.1
v.2.16.0
2026-04-30 20:05:45 -05:00
chrisvela and GitHub
cff7b09ef2
Merge pull request #289 from runpod-workers/feat/add-notifications
...
feat: add notifications for new prs, issues, and new releases of vllm
2026-04-30 20:03:48 -05:00
velaraptor-runpod
04b342c675
fix permissions
2026-04-30 20:00:48 -05:00
velaraptor-runpod
5b29643799
feat: add notifications for new prs, issues, and new releases of vllm
2026-04-30 19:55:29 -05:00
velaraptor-runpod
a774cefe85
fix: clean workspace file
2026-04-30 17:29:00 -05:00
velaraptor-runpod
9de17d49b7
feat: update vllm to 0.18.1
2026-04-30 17:25:23 -05:00
velaraptor-runpod
296556a6f7
feat: update vllm to 0.17.1
2026-04-30 17:21:03 -05:00
chrisvela and GitHub
a1544ea70d
Merge pull request #277 from runpod-workers/feat/lmcache
...
feat: uv installer, LMCache support, add /v1/responses and /v1/messages endpoints
v.2.15.0
2026-04-28 19:44:00 -03:00
Tim Pietrusky
3403889528
fix: address review comments on responses/messages handlers and lmcache guard
...
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
except blocks of _handle_responses_request and _handle_messages_request;
emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
2026-04-23 11:28:50 +02:00
Tim Pietrusky
f299204770
docs: fix anthropic messages path and missing comma in routes list
2026-04-23 10:54:24 +02:00
velaraptor-runpod
1ed25eea20
update readme with responses and messages routes
2026-04-16 15:30:30 -05:00
velaraptor-runpod
dc4ad7ddeb
fix: kv_transfer_config is dataclass not json, fix for env var
2026-04-09 14:33:11 -05:00
velaraptor-runpod
9035b0e07f
fix: lmcache version
2026-04-09 14:32:19 -05:00
velaraptor-runpod
c979f0020f
update readmes on TRANSFORMERS_VERSION
2026-04-06 16:54:08 -05:00
velaraptor-runpod
6fbd480a26
fix: allow for transformers_version
2026-04-06 16:49:32 -05:00
velaraptor-runpod
3ef1fb8e7b
requested changes
2026-04-03 15:53:40 -05:00
velaraptor-runpod
30e8514d63
Runpod not RunPod
2026-03-17 20:59:52 -05:00
velaraptor-runpod
d9808815ee
feat: update requirements.txt
2026-03-17 19:56:39 -05:00
velaraptor-runpod
d8ed3b5353
feat: update 0.16.0, add lmcache
2026-03-17 19:54:11 -05:00
velaraptor-runpod
4c4e039565
feat: add messages route for anthropic/claude
2026-03-17 19:51:50 -05:00
chrisvela and GitHub
9d1686960d
Merge pull request #273 from runpod-workers/bug/hf-overides-rope-scaling
...
bug: fix rope scaling to be forward compatible from hf_overrides
2026-03-10 11:21:44 -05:00
velaraptor-runpod
45d1eeee47
bug: fix rope scaling to be forward compatible from hf_overrides
2026-03-06 15:34:11 -06:00
chrisvela and GitHub
17efb0e7d0
Merge pull request #272 from runpod-workers/feat/vllm-0.16.0
...
Release / release (push) Waiting to run
feat: Update to 0.16.0
v2.14.0
2026-03-05 13:06:45 -06:00
velaraptor-runpod
2b5f07df63
feat: Update to 0.16.0, remove NUM_GPU_BLOCKS_OVERRIDE in hub default since 0 will break
2026-03-04 16:38:40 -06:00
chrisvela and GitHub
13fa71878e
Merge pull request #269 from runpod-workers/feat/allow-engine-args-env
...
Release / release (push) Waiting to run
feat: allow all AsyncEngineArgs as env vars
v2.13.1
2026-02-27 15:34:23 -06:00
velaraptor-runpod
8a9365bed4
remove DEFAULT_ARGS that are none, fix MAX_CONTEXT_LEN_TO_CAPTURE
2026-02-27 14:04:15 -06:00
velaraptor-runpod
cd485a1af1
update readme
2026-02-25 22:57:17 -06:00
velaraptor-runpod
b9043639e9
requested changes/refactor
2026-02-25 16:07:38 -06:00
chrisvela and GitHub
407dbd7773
Merge pull request #270 from runpod-workers/feat/update-vllm-v0.15.1
...
feat: update vllm to 0.15.1
2026-02-25 15:35:58 -06:00
velaraptor-runpod
f103c142c1
feat: update vllm to 0.15.1
2026-02-24 17:44:49 -06:00
velaraptor-runpod
efb093e198
add as VLLM_RUNPOD prefix and update readme
2026-02-24 17:37:57 -06:00
velaraptor-runpod
42443f735e
feat: allow engine args through VLLM_ and checks the engine args
2026-02-24 16:05:18 -06:00
chrisvela and GitHub
b7c6d4f9a2
feat: update dockerfile to 12.9.1 ( #267 )
...
Release / release (push) Waiting to run
* feat: update dockerfile to 12.9.1
* update readme on VLLM_NIGHTLY build arg
v2.13.0
2026-02-19 10:13:14 +01:00
chrisvela and GitHub
d69cc021e8
Merge pull request #268 from runpod-workers/fix/spec-config-0-to-none
...
Release / release (push) Waiting to run
fix: spec config env vars should be none if zero
v2.12.3
2026-02-18 15:51:51 -06:00
velaraptor-runpod
61faa8f137
fix: spec config env vars should be none if zero
2026-02-18 15:41:19 -06:00
chrisvela and GitHub
1606cff557
Merge pull request #265 from runpod-workers/fix/zero-max-model-num_batches
...
Release / release (push) Waiting to run
fix: check for zero param and set to None
v2.12.2
2026-02-13 15:26:06 -06:00
velaraptor-runpod
e705c9494b
fix: check for zero param and set to None
2026-02-13 15:23:54 -06:00
chrisvela and GitHub
b749aa5718
Merge pull request #264 from runpod-workers/fix/max_num_batched_tokens
...
Release / release (push) Waiting to run
fix: max num batched tokens
v2.12.1
2026-02-13 12:38:01 -06:00
velaraptor-runpod
4705ba8a7c
fix: check max_num_batched_tokenz if max_model_len not set
2026-02-13 03:29:52 -06:00
velaraptor-runpod
767c66c301
make minimal changes
2026-02-13 03:23:44 -06:00
velaraptor-runpod
fefdbe21a9
update changes
2026-02-13 03:16:43 -06:00
velaraptor-runpod
ee961ad28d
Update hub.json
2026-02-13 03:08:19 -06:00
velaraptor-runpod
2e8c251447
Merge branch 'main' into feat/update-vllm-v0.15.0
2026-02-13 03:01:05 -06:00
velaraptor-runpod
c3cf43b228
Update hub.json
2026-02-13 00:22:16 -06:00
velaraptor-runpod
7ec10b98cd
Update utils.py
2026-02-12 15:28:31 -06:00
c45ac42acd
vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 ( #259 )
...
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes
* MAX_NUM_BATCHED_TOKENS fix and CUDA tester
* Sys kill worker instead of marking as failed
* upgrade to vllm 0.12.0
* Update to vllm 0.15.0 and lora fix
* Update for HUB and removal of deprected env variables
* reverted docker-bake changes
* removed leftovers
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Update src/utils.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Clean up of docs and comments in code
* nit: lowercase p
* nit: lowercase p
---------
Co-authored-by: Dj Isaac <contact@dejaydev.com >
Co-authored-by: chrisvela <chris.vela@runpod.io >
v2.12.0
2026-02-12 21:50:34 +01:00
velaraptor-runpod
340bc0b3c6
fix: served model name
2026-02-10 21:42:58 -06:00
velaraptor-runpod
e1e9ef74ad
add changes from pr
2026-02-06 18:10:09 -06:00