Commit Graph
102 Commits
Author SHA1 Message Date
Owen Qwen 08c6ff7490 Bump vLLM to 0.17.0
CI | Update runpod package version / Check python requirements file and update (push) Canceled after 0s
2026-03-26 12:13:42 -05:00
velaraptor-runpod 2b5f07df63 feat: Update to 0.16.0, remove NUM_GPU_BLOCKS_OVERRIDE in hub default since 0 will break 2026-03-04 16:38:40 -06:00
velaraptor-runpod f103c142c1 feat: update vllm to 0.15.1 2026-02-24 17:44:49 -06:00
chrisvelaandGitHub b7c6d4f9a2 feat: update dockerfile to 12.9.1 (#267)
Release / release (push) Waiting to run
* feat: update dockerfile to 12.9.1

* update readme on VLLM_NIGHTLY build arg
2026-02-19 10:13:14 +01:00
velaraptor-runpod fefdbe21a9 update changes 2026-02-13 03:16:43 -06:00
velaraptor-runpod 2e8c251447 Merge branch 'main' into feat/update-vllm-v0.15.0 2026-02-13 03:01:05 -06:00
c45ac42acd vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes

* MAX_NUM_BATCHED_TOKENS fix and CUDA tester

* Sys kill worker instead of marking as failed

* upgrade to vllm 0.12.0

* Update to vllm 0.15.0 and lora fix

* Update for HUB and removal of deprected env variables

* reverted docker-bake changes

* removed leftovers

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/utils.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Clean up of docs and comments in code

* nit: lowercase p

* nit: lowercase p

---------

Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
2026-02-12 21:50:34 +01:00
velaraptor-runpod e1e9ef74ad add changes from pr 2026-02-06 18:10:09 -06:00
velaraptor-runpod 8eb55b90c1 add changes for v0.15.0 2026-02-05 17:24:16 -06:00
90c16b472d fix: update CUDA to 12.4.1 for Blackwell GPU support (#251)
Release / release (push) Waiting to run
* fix: update CUDA to 12.4.1 for Blackwell GPU support

- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json

This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.

Fixes: DR-1118

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* revert: remove NVIDIA B200 from default gpuIds

The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* fix: remove FlashInfer to avoid JIT compilation errors

FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.

vLLM will use its built-in fallback sampling methods instead.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-13 22:01:36 +01:00
Tim Pietrusky fae16e7ee1 chore: update vllm to 0.11.0 2025-10-22 13:40:53 -07:00
Jhenner Tigreros 8b02a703b4 fix issues 2025-08-07 15:59:43 -05:00
Jhenner Tigreros d8863139d6 add model to test 2025-08-07 15:28:02 -05:00
Tim Pietrusky a4062fc488 feat: update to 0.9.1 2025-06-14 13:31:27 +02:00
pandyamarut 70cd1c8113 update vllm version 0.9.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-06-05 11:59:15 -07:00
Marut PandyaandGitHub 6db2c44d3b Update Dockerfile 2025-05-09 09:27:54 -07:00
pandyamarut f33e8d2bcd update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-04-21 13:14:32 -07:00
pandyamarut 3d067cd472 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-04-07 11:54:57 -07:00
pandyamarut acbdf63de8 update vllm 0.8.2
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-03-26 15:55:30 -07:00
pandyamarut 56dc4ad075 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-02-24 15:01:45 -08:00
pandyamarut c9791f1163 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-02-11 13:27:57 -08:00
pandyamarut fcbfe84f63 update vllm 0.7.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-01-28 14:18:10 -08:00
Hailong Yang a578c6df23 update docker file 2025-01-26 23:28:27 -05:00
pandyamarut 06c2bb1715 upgrade vllm version
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-12-30 15:20:11 -08:00
pandyamarut 4e10641d69 update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-11-19 12:14:04 -08:00
pandyamarut c03ecc42fe update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-10-15 16:49:41 -07:00
pandyamarut 1420091588 update vll
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-27 19:23:50 -07:00
pandyamarut 2ac6a0108f update vllm v0.6.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-16 14:55:40 -07:00
pandyamarut 5e1c8c8128 update vllm version 0.5.5
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-08-28 22:46:45 -07:00
pandyamarut 9cb9336cf5 update vllm version 0.5.4
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-08-09 12:01:42 -07:00
pandyamarut bd96b5e0de update v0.5.3.post1
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-07-25 21:29:45 -07:00
alpayariyak 5bd6f3a75e 0.5.3, any vllm arg as env var, refactor and fixes, moving away from building separate image from vLLM fork 2024-07-25 12:41:48 -07:00
alpayariyak c8458fef2b preparation for 1.0.0 release 2024-06-12 12:28:03 -07:00
alpayariyak ec7ea0b760 Update default base image version to fix github actions build 2024-05-09 20:18:34 -04:00
Alpay Ariyakandalpayariyak 874379a0c5 1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) (#62) 2024-05-09 00:42:08 -04:00
alpayariyak db7167d57f 0.3.3 2024-03-05 19:14:35 +00:00
Alpay Ariyakandalpayariyak d91ccb866f 0.3.1: bug fixes 2024-02-29 02:55:44 -05:00
alpayariyak 6bcd9d7c67 Preparing for 0.3.0 2024-02-22 18:04:34 -05:00
alpayariyak 8de10468dd Working tokenizer and model download fix
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00
alpayariyak e7b340d73b Update vLLM base image 2024-01-31 03:37:55 +00:00
alpayariyak 12d6f0778e Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer 2024-01-31 03:13:21 +00:00
Alpay AriyakandGitHub fa5556434c Bug fix 2024-01-29 11:26:22 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
Alpay Ariyakandalpayariyak 368c5f87fb Temporary Dockerfile Fix
Temporary Dockerfile fix
2024-01-25 20:48:26 -05:00
alpayariyak 584852f0f6 Docker-Protected HF Token, Refactor, Better Documentation 2024-01-18 18:20:51 -05:00
alpayariyak c59038e902 Fixed CUDA 11.8 workers 2023-12-20 07:50:14 +00:00
Alpay AriyakandGitHub c321d828ab Update Dockerfile 2023-12-19 00:58:45 -05:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00
alpayariyak 22de0b322d Complete refactor, improved overall functionality 2023-12-13 10:33:03 +00:00