Commit Graph
112 Commits
Author SHA1 Message Date
velaraptor-runpod 32b29d4c6c fix: add deepgemm, update base image and hub for cuda 13.0 2026-05-20 17:43:42 -05:00
velaraptor-runpod dcea4fc4f9 fix dockerfile 2026-05-15 12:08:31 -04:00
velaraptor-runpod 9c139e8ceb update: update to 0.20.1, update dockerfile to cuda 13 2026-05-15 11:35:53 -04:00
velaraptor-runpod 678bb4be8f feat: update to 0.20.1 for patch fixes 2026-05-07 11:58:22 -05:00
velaraptor-runpodandClaude Sonnet 4.6 22356ee2b3 feat: upgrade vLLM to 0.20.0
- Bump vllm[flashinfer] to 0.20.0 in Dockerfile
- Remove io_processor param from OpenAIServingRender (dropped in 0.20.0)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 18:39:07 -05:00
velaraptor-runpodandClaude Sonnet 4.6 e6950bdebd feat: upgrade vLLM to 0.19.1
- Bump vllm[flashinfer] to 0.19.1 in Dockerfile
- Add OpenAIServingRender (new required dependency in 0.19.x serving layer)
- Pass openai_serving_render to all four serving class constructors
- Remove log_error_stack param (removed upstream in 0.19.x)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 17:41:11 -05:00
velaraptor-runpod 9de17d49b7 feat: update vllm to 0.18.1 2026-04-30 17:25:23 -05:00
velaraptor-runpod 296556a6f7 feat: update vllm to 0.17.1 2026-04-30 17:21:03 -05:00
velaraptor-runpod 6fbd480a26 fix: allow for transformers_version 2026-04-06 16:49:32 -05:00
velaraptor-runpod 3ef1fb8e7b requested changes 2026-04-03 15:53:40 -05:00
velaraptor-runpod d8ed3b5353 feat: update 0.16.0, add lmcache 2026-03-17 19:54:11 -05:00
velaraptor-runpod 2b5f07df63 feat: Update to 0.16.0, remove NUM_GPU_BLOCKS_OVERRIDE in hub default since 0 will break 2026-03-04 16:38:40 -06:00
velaraptor-runpod f103c142c1 feat: update vllm to 0.15.1 2026-02-24 17:44:49 -06:00
chrisvelaandGitHub b7c6d4f9a2 feat: update dockerfile to 12.9.1 (#267)
Release / release (push) Waiting to run
* feat: update dockerfile to 12.9.1

* update readme on VLLM_NIGHTLY build arg
2026-02-19 10:13:14 +01:00
velaraptor-runpod fefdbe21a9 update changes 2026-02-13 03:16:43 -06:00
velaraptor-runpod 2e8c251447 Merge branch 'main' into feat/update-vllm-v0.15.0 2026-02-13 03:01:05 -06:00
c45ac42acd vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes

* MAX_NUM_BATCHED_TOKENS fix and CUDA tester

* Sys kill worker instead of marking as failed

* upgrade to vllm 0.12.0

* Update to vllm 0.15.0 and lora fix

* Update for HUB and removal of deprected env variables

* reverted docker-bake changes

* removed leftovers

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/utils.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Clean up of docs and comments in code

* nit: lowercase p

* nit: lowercase p

---------

Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
2026-02-12 21:50:34 +01:00
velaraptor-runpod e1e9ef74ad add changes from pr 2026-02-06 18:10:09 -06:00
velaraptor-runpod 8eb55b90c1 add changes for v0.15.0 2026-02-05 17:24:16 -06:00
90c16b472d fix: update CUDA to 12.4.1 for Blackwell GPU support (#251)
Release / release (push) Waiting to run
* fix: update CUDA to 12.4.1 for Blackwell GPU support

- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json

This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.

Fixes: DR-1118

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* revert: remove NVIDIA B200 from default gpuIds

The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove FlashInfer to avoid JIT compilation errors

FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.

vLLM will use its built-in fallback sampling methods instead.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-13 22:01:36 +01:00
Tim Pietrusky fae16e7ee1 chore: update vllm to 0.11.0 2025-10-22 13:40:53 -07:00
Jhenner Tigreros 8b02a703b4 fix issues 2025-08-07 15:59:43 -05:00
Jhenner Tigreros d8863139d6 add model to test 2025-08-07 15:28:02 -05:00
Tim Pietrusky a4062fc488 feat: update to 0.9.1 2025-06-14 13:31:27 +02:00
pandyamarut 70cd1c8113 update vllm version 0.9.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-06-05 11:59:15 -07:00
Marut PandyaandGitHub 6db2c44d3b Update Dockerfile 2025-05-09 09:27:54 -07:00
pandyamarut f33e8d2bcd update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-04-21 13:14:32 -07:00
pandyamarut 3d067cd472 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-04-07 11:54:57 -07:00
pandyamarut acbdf63de8 update vllm 0.8.2
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-03-26 15:55:30 -07:00
pandyamarut 56dc4ad075 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-02-24 15:01:45 -08:00
pandyamarut c9791f1163 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-02-11 13:27:57 -08:00
pandyamarut fcbfe84f63 update vllm 0.7.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-01-28 14:18:10 -08:00
Hailong Yang a578c6df23 update docker file 2025-01-26 23:28:27 -05:00
pandyamarut 06c2bb1715 upgrade vllm version
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-12-30 15:20:11 -08:00
pandyamarut 4e10641d69 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-11-19 12:14:04 -08:00
pandyamarut c03ecc42fe update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-10-15 16:49:41 -07:00
pandyamarut 1420091588 update vll
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-27 19:23:50 -07:00
pandyamarut 2ac6a0108f update vllm v0.6.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-16 14:55:40 -07:00
pandyamarut 5e1c8c8128 update vllm version 0.5.5
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-08-28 22:46:45 -07:00
pandyamarut 9cb9336cf5 update vllm version 0.5.4
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-08-09 12:01:42 -07:00
pandyamarut bd96b5e0de update v0.5.3.post1
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-07-25 21:29:45 -07:00
alpayariyak 5bd6f3a75e 0.5.3, any vllm arg as env var, refactor and fixes, moving away from building separate image from vLLM fork 2024-07-25 12:41:48 -07:00
alpayariyak c8458fef2b preparation for 1.0.0 release 2024-06-12 12:28:03 -07:00
alpayariyak ec7ea0b760 Update default base image version to fix github actions build 2024-05-09 20:18:34 -04:00
Alpay Ariyakandalpayariyak 874379a0c5 1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) (#62) 2024-05-09 00:42:08 -04:00
alpayariyak db7167d57f 0.3.3 2024-03-05 19:14:35 +00:00
Alpay Ariyakandalpayariyak d91ccb866f 0.3.1: bug fixes 2024-02-29 02:55:44 -05:00
alpayariyak 6bcd9d7c67 Preparing for 0.3.0 2024-02-22 18:04:34 -05:00
alpayariyak 8de10468dd Working tokenizer and model download fix
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00