chrisvela and GitHub
591f9d6531
Merge pull request #317 from runpod-workers/fix/test-sls-gh-action
...
fix: fix image to test with sls in gh actions
2026-07-10 09:55:53 -05:00
velaraptor-runpod
49967d0c09
fix: fix image to test with sls in gh actions
2026-07-09 22:25:40 -05:00
chrisvela and GitHub
828e797498
Merge pull request #315 from runpod-workers/feat/0.23.0
...
feat: update vllm to 0.23.0 + add sls e2e model testing
2026-07-09 21:18:41 -05:00
velaraptor-runpod
373847d9f9
fix: pin models to version, except for gpt-oss
2026-07-09 19:21:24 -05:00
velaraptor-runpod
2000acc0a3
chore: add timeout for e2e serverless
2026-07-09 19:15:19 -05:00
velaraptor-runpod
89e3b2e920
fix: fix for huggingface, secret keys
2026-07-09 18:56:58 -05:00
velaraptor-runpod
85f8c81c1c
add sls e2e tests
2026-07-09 18:22:57 -05:00
velaraptor-runpod
88490e9551
feat: update to 0.23.0
2026-07-09 18:22:40 -05:00
velaraptor-runpod
798f2b12eb
feat: update to 0.23.0
2026-07-01 16:19:01 -05:00
chrisvela and GitHub
9e1c483136
Merge pull request #311 from runpod-workers/t3code/cd56c69d
...
Release / release (push) Waiting to run
fix: serve original model name when HF cache dir is lowercased (#310 )
2026-06-26 13:48:13 -05:00
chrisvela and GitHub
d71ea9939d
Merge pull request #308 from runpod-workers/docs/sync-vllm-version-0.20.2
...
docs: sync vLLM version to 0.20.2 in READMEs
2026-06-26 13:47:59 -05:00
chrisvela and GitHub
7e2b4e2288
Merge pull request #314 from runpod-workers/runpod-package-update
...
chore: update runpod to 1.10.0
2026-06-26 13:45:16 -05:00
chrisvela and GitHub
1b3228a2dc
Merge pull request #307 from runpod-workers/fix/revert-0.20.0
...
Release / release (push) Waiting to run
revert to 0.20.2
2026-06-12 15:50:52 -05:00
velaraptor-runpod
4817d4a8e7
revert to 0.20.2
2026-06-12 15:49:23 -05:00
chrisvela and GitHub
0378382a92
Merge pull request #306 from runpod-workers/revert/v2.20.1
...
Release / release (push) Waiting to run
chore: carry non-vllm changes from main (tests GPU + configs)
2026-06-12 15:21:22 -05:00
velaraptor-runpod and Claude Sonnet 4.6
08580e7ccf
chore: carry non-vllm changes from main (tests GPU + configs)
...
Brings forward the L40 GPU type in tests.json and the new llama/qwen
tuned config files, while keeping Dockerfile pinned at vllm 0.20.2
(v2.20.1 state).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-06-12 15:18:36 -05:00
chrisvela and GitHub
8868aae6b1
Merge pull request #305 from runpod-workers/fix/typo-tests-l40
...
Release / release (push) Waiting to run
fix: fix typo in tests
2026-06-12 14:00:54 -05:00
velaraptor-runpod
fb8adc5c06
fix: fix typo in tests
2026-06-12 13:59:29 -05:00
chrisvela and GitHub
7351da512b
Merge pull request #304 from runpod-workers/fix/fix-tests
...
Release / release (push) Waiting to run
fix: change test gpu to L40
2026-06-12 13:49:15 -05:00
velaraptor-runpod
1e78043b2c
fix: change test gpu to L40
2026-06-12 13:46:33 -05:00
chrisvela and GitHub
352c64f4c1
Merge pull request #303 from runpod-workers/chore/check-vllm-versions-readmes
...
chore: check readmes vllm version in sync with version in Dockerfile
2026-06-12 09:49:07 -05:00
velaraptor-runpod
0922f5b435
chore: check readmes vllm version in sync with version in Dockerfile
2026-06-11 17:33:16 -05:00
chrisvela and GitHub
5d9a48fc70
Merge pull request #302 from runpod-workers/feat/0.22.1
...
Release / release (push) Waiting to run
feat: upgrade vllm to 0.22.1
2026-06-11 17:27:38 -05:00
velaraptor-runpod
5d1579e361
fix: update readme links
2026-06-11 17:00:31 -05:00
velaraptor-runpod
0488b77d89
feat: upgrade vllm to 0.22.1
2026-06-11 16:29:02 -05:00
chrisvela and GitHub
d3a962c33b
Merge pull request #301 from adithyaJRunpod/feature/tuned-configs
...
Release / release (push) Waiting to run
Add tuned configs, CON-239
2026-06-11 14:26:00 -05:00
chrisvela and GitHub
105c125698
Merge pull request #300 from runpod-workers/feat/0.21.0
...
feat: upgrade vllm to 0.21.0
2026-06-11 12:47:08 -05:00
velaraptor-runpod
c8ce53c72c
fix: add kenels, and fix for cuda
2026-06-10 16:01:53 -05:00
velaraptor-runpod
9618e799ba
chore: fix cuda libraries
2026-06-04 15:35:58 -05:00
velaraptor-runpod
cb3f077dba
feat: upgrade vllm to 0.21.0
2026-06-03 14:43:56 -05:00
chrisvela and GitHub
69646b9e99
Merge pull request #294 from runpod-workers/feat/allow-config
...
Release / release (push) Waiting to run
feat: allow config.yaml like vllm serve
2026-06-02 20:28:02 -05:00
velaraptor-runpod
80072047ab
feat: allow config.yaml like vllm serve
2026-05-29 15:50:22 -05:00
chrisvela and GitHub
50aba8fb57
Merge pull request #293 from runpod-workers/fix/update-deep-gemm-hub-value
...
Release / release (push) Waiting to run
fix: update VLLM_USE_DEEP_GEMM hub to default to 0
2026-05-27 11:51:48 -05:00
velaraptor-runpod
9edc5715ce
fix: update configuration.md
2026-05-27 11:33:37 -05:00
velaraptor-runpod
4c91f2c5b5
fix: update VLLM_USE_DEEP_GEMM hub to default to 0
2026-05-27 11:26:52 -05:00
chrisvela and GitHub
6265b99348
Merge pull request #288 from runpod-workers/feat/0.20.0
...
feat: update to 0.20.2
2026-05-26 17:17:45 -05:00
velaraptor-runpod
026f8d700b
fix: specify deepgemm commit version
2026-05-21 18:46:54 -05:00
chrisvela and GitHub
146bdb0252
Merge branch 'main' into feat/0.20.0
2026-05-21 16:04:39 -05:00
velaraptor-runpod
da01193a3d
chore: fix readme
2026-05-21 15:57:40 -05:00
velaraptor-runpod
c2e6cc9f61
chore: update readme with correct vllm version
2026-05-21 15:39:19 -05:00
velaraptor-runpod
69968a6b39
chore: fix logging, warning for text prompt
2026-05-21 15:34:09 -05:00
velaraptor-runpod
32b29d4c6c
fix: add deepgemm, update base image and hub for cuda 13.0
2026-05-20 17:43:42 -05:00
velaraptor-runpod
dcea4fc4f9
fix dockerfile
2026-05-15 12:08:31 -04:00
velaraptor-runpod
9c139e8ceb
update: update to 0.20.1, update dockerfile to cuda 13
2026-05-15 11:35:53 -04:00
velaraptor-runpod
678bb4be8f
feat: update to 0.20.1 for patch fixes
2026-05-07 11:58:22 -05:00
chrisvela and GitHub
87d7365126
Merge pull request #292 from runpod-workers/fix/open-ai
...
Release / release (push) Waiting to run
fix: fix warmup
2026-05-01 18:06:11 -05:00
velaraptor-runpod
0e83616f93
fix: fix warmup
2026-05-01 17:39:51 -05:00
chrisvela and GitHub
ab6d39dcf8
Merge pull request #291 from runpod-workers/bug/281-hf-token
...
Release / release (push) Waiting to run
bug: fix hf-token being passed in engineargs
2026-05-01 15:32:04 -05:00
velaraptor-runpod
ed315a175e
merge main
2026-05-01 15:28:05 -05:00
velaraptor-runpod
73f030ae5e
Merge branch 'main' into bug/281-hf-token
2026-05-01 14:33:51 -05:00
chrisvela and GitHub
8a099c1723
Merge pull request #287 from runpod-workers/feat/0.19.1
...
feat: update to 0.19.1
2026-05-01 14:29:39 -05:00
velaraptor-runpod
ff87840a58
fix: add enforce_eager as true, add pytorch_alloc_conf to expandle_segments to True for OOM, for hub defaults
2026-05-01 14:13:03 -05:00
velaraptor-runpod
7dc853b1fe
chore: remove release to trigger on release publish, just use tags. duplicate
2026-05-01 10:13:32 -05:00
velaraptor-runpod
a8b754b92a
merge main
2026-05-01 10:12:43 -05:00
chrisvela and GitHub
6357aeda51
Merge pull request #290 from runpod-workers/fix/fix-old-actions
...
Release / release (push) Waiting to run
fix: fix old github actions, trigger release on publish release
2026-05-01 10:01:22 -05:00
velaraptor-runpod
0140b29c44
chore: add release notes to slack notification
2026-05-01 09:48:43 -05:00
velaraptor-runpod
0cb8aeae77
chore: add specific runpod version
2026-05-01 09:47:23 -05:00
chrisvela and GitHub
cd8f9e9560
Merge branch 'main' into fix/fix-old-actions
2026-04-30 21:57:45 -05:00
chrisvela and GitHub
895fd25fac
Merge pull request #286 from runpod-workers/feat/0.18.1
...
feat: update vllm to 0.18.1
2026-04-30 21:56:07 -05:00
velaraptor-runpod
7bb8df73af
bug: fix hf-token being passed in engineargs
2026-04-30 20:37:30 -05:00
velaraptor-runpod
747cdf5891
chore: update readme vllm version
2026-04-30 20:26:23 -05:00
velaraptor-runpod
3d4af5df9b
chore: update readme vllm version
2026-04-30 20:25:35 -05:00
velaraptor-runpod
72547aa3bb
chore: update readme vllm version
2026-04-30 20:25:02 -05:00
velaraptor-runpod
577fd8c3c3
fix: fix old github actions, trigger release on publish release
2026-04-30 20:22:24 -05:00
chrisvela and GitHub
f49f35456e
Merge pull request #285 from runpod-workers/feat/0.17.1
...
feat: update vllm to 0.17.1
2026-04-30 20:05:45 -05:00
chrisvela and GitHub
cff7b09ef2
Merge pull request #289 from runpod-workers/feat/add-notifications
...
feat: add notifications for new prs, issues, and new releases of vllm
2026-04-30 20:03:48 -05:00
velaraptor-runpod
04b342c675
fix permissions
2026-04-30 20:00:48 -05:00
velaraptor-runpod
5b29643799
feat: add notifications for new prs, issues, and new releases of vllm
2026-04-30 19:55:29 -05:00
velaraptor-runpod
4f8a16df5d
fix: update transformers to >=5
2026-04-30 19:41:09 -05:00
velaraptor-runpod and Claude Sonnet 4.6
22356ee2b3
feat: upgrade vLLM to 0.20.0
...
- Bump vllm[flashinfer] to 0.20.0 in Dockerfile
- Remove io_processor param from OpenAIServingRender (dropped in 0.20.0)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-04-30 18:39:07 -05:00
velaraptor-runpod and Claude Sonnet 4.6
fa42ecd79a
fix: resolve lowercase HF cache paths when MODEL_NAME uses original casing
...
Fixes FDE-174. Some model stores (e.g. RunPod pre-cached network volumes)
normalize repo IDs to lowercase. HuggingFace Hub caches using the original
casing, so MODEL_NAME=Qwen/Qwen2.5-Coder-32B-Instruct-AWQ would miss a
cache stored as models--qwen--qwen2.5-coder-32b-instruct-awq/ and attempt
a redundant download that fails on limited container storage.
If the exact-case HF cache directory is absent but a lowercase variant
exists, the latest snapshot path is returned directly so vLLM loads from
disk. Absolute paths and models with no lowercase cache are unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-04-30 17:54:14 -05:00
velaraptor-runpod and Claude Sonnet 4.6
178c72238e
fix: surface LORA_MODULES parse failures instead of silently loading zero adapters
...
Fixes FDE-194. Previously a malformed LORA_MODULES value was swallowed at
info level and the engine would start with no LoRA adapters, causing 500s
on any request using an adapter model name (e.g. npc-sim-*).
Changes:
- Log at error level when LORA_MODULES cannot be parsed as JSON
- Log at error level when individual adapter dicts fail LoRAModulePath validation
- Log a final error when all adapters fail to load so the cause is obvious
- Accept a single adapter dict (not just an array) for convenience
- Return early when LORA_MODULES is unset to skip unnecessary parsing
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-04-30 17:49:09 -05:00
velaraptor-runpod and Claude Sonnet 4.6
e6950bdebd
feat: upgrade vLLM to 0.19.1
...
- Bump vllm[flashinfer] to 0.19.1 in Dockerfile
- Add OpenAIServingRender (new required dependency in 0.19.x serving layer)
- Pass openai_serving_render to all four serving class constructors
- Remove log_error_stack param (removed upstream in 0.19.x)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-04-30 17:41:11 -05:00
velaraptor-runpod
a774cefe85
fix: clean workspace file
2026-04-30 17:29:00 -05:00
velaraptor-runpod
9de17d49b7
feat: update vllm to 0.18.1
2026-04-30 17:25:23 -05:00
velaraptor-runpod
296556a6f7
feat: update vllm to 0.17.1
2026-04-30 17:21:03 -05:00
chrisvela and GitHub
a1544ea70d
Merge pull request #277 from runpod-workers/feat/lmcache
...
feat: uv installer, LMCache support, add /v1/responses and /v1/messages endpoints
2026-04-28 19:44:00 -03:00
velaraptor-runpod
1ed25eea20
update readme with responses and messages routes
2026-04-16 15:30:30 -05:00
velaraptor-runpod
dc4ad7ddeb
fix: kv_transfer_config is dataclass not json, fix for env var
2026-04-09 14:33:11 -05:00
velaraptor-runpod
9035b0e07f
fix: lmcache version
2026-04-09 14:32:19 -05:00
velaraptor-runpod
c979f0020f
update readmes on TRANSFORMERS_VERSION
2026-04-06 16:54:08 -05:00
velaraptor-runpod
6fbd480a26
fix: allow for transformers_version
2026-04-06 16:49:32 -05:00
velaraptor-runpod
3ef1fb8e7b
requested changes
2026-04-03 15:53:40 -05:00
velaraptor-runpod
30e8514d63
Runpod not RunPod
2026-03-17 20:59:52 -05:00
velaraptor-runpod
d9808815ee
feat: update requirements.txt
2026-03-17 19:56:39 -05:00
velaraptor-runpod
d8ed3b5353
feat: update 0.16.0, add lmcache
2026-03-17 19:54:11 -05:00
velaraptor-runpod
4c4e039565
feat: add messages route for anthropic/claude
2026-03-17 19:51:50 -05:00
chrisvela and GitHub
9d1686960d
Merge pull request #273 from runpod-workers/bug/hf-overides-rope-scaling
...
bug: fix rope scaling to be forward compatible from hf_overrides
2026-03-10 11:21:44 -05:00
velaraptor-runpod
45d1eeee47
bug: fix rope scaling to be forward compatible from hf_overrides
2026-03-06 15:34:11 -06:00
chrisvela and GitHub
17efb0e7d0
Merge pull request #272 from runpod-workers/feat/vllm-0.16.0
...
Release / release (push) Waiting to run
feat: Update to 0.16.0
2026-03-05 13:06:45 -06:00
velaraptor-runpod
2b5f07df63
feat: Update to 0.16.0, remove NUM_GPU_BLOCKS_OVERRIDE in hub default since 0 will break
2026-03-04 16:38:40 -06:00
chrisvela and GitHub
13fa71878e
Merge pull request #269 from runpod-workers/feat/allow-engine-args-env
...
Release / release (push) Waiting to run
feat: allow all AsyncEngineArgs as env vars
2026-02-27 15:34:23 -06:00
velaraptor-runpod
8a9365bed4
remove DEFAULT_ARGS that are none, fix MAX_CONTEXT_LEN_TO_CAPTURE
2026-02-27 14:04:15 -06:00
velaraptor-runpod
cd485a1af1
update readme
2026-02-25 22:57:17 -06:00
velaraptor-runpod
b9043639e9
requested changes/refactor
2026-02-25 16:07:38 -06:00
chrisvela and GitHub
407dbd7773
Merge pull request #270 from runpod-workers/feat/update-vllm-v0.15.1
...
feat: update vllm to 0.15.1
2026-02-25 15:35:58 -06:00
velaraptor-runpod
f103c142c1
feat: update vllm to 0.15.1
2026-02-24 17:44:49 -06:00
velaraptor-runpod
efb093e198
add as VLLM_RUNPOD prefix and update readme
2026-02-24 17:37:57 -06:00
velaraptor-runpod
42443f735e
feat: allow engine args through VLLM_ and checks the engine args
2026-02-24 16:05:18 -06:00
chrisvela and GitHub
b7c6d4f9a2
feat: update dockerfile to 12.9.1 ( #267 )
...
Release / release (push) Waiting to run
* feat: update dockerfile to 12.9.1
* update readme on VLLM_NIGHTLY build arg
2026-02-19 10:13:14 +01:00