Commit Graph
59 Commits
Author SHA1 Message Date
jhcipar d356c31675 feat: bump runpod-python version 2026-06-01 20:28:40 -04:00
velaraptor-runpod 32b29d4c6c fix: add deepgemm, update base image and hub for cuda 13.0 2026-05-20 17:43:42 -05:00
velaraptor-runpod a8b754b92a merge main 2026-05-01 10:12:43 -05:00
velaraptor-runpod 0cb8aeae77 chore: add specific runpod version 2026-05-01 09:47:23 -05:00
velaraptor-runpod 4f8a16df5d fix: update transformers to >=5 2026-04-30 19:41:09 -05:00
velaraptor-runpod 296556a6f7 feat: update vllm to 0.17.1 2026-04-30 17:21:03 -05:00
Tim Pietrusky 3403889528 fix: address review comments on responses/messages handlers and lmcache guard
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
  except blocks of _handle_responses_request and _handle_messages_request;
  emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
  blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
  actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
  and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
2026-04-23 11:28:50 +02:00
velaraptor-runpod 9035b0e07f fix: lmcache version 2026-04-09 14:32:19 -05:00
velaraptor-runpod 6fbd480a26 fix: allow for transformers_version 2026-04-06 16:49:32 -05:00
velaraptor-runpod 3ef1fb8e7b requested changes 2026-04-03 15:53:40 -05:00
velaraptor-runpod d9808815ee feat: update requirements.txt 2026-03-17 19:56:39 -05:00
c45ac42acd vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes

* MAX_NUM_BATCHED_TOKENS fix and CUDA tester

* Sys kill worker instead of marking as failed

* upgrade to vllm 0.12.0

* Update to vllm 0.15.0 and lora fix

* Update for HUB and removal of deprected env variables

* reverted docker-bake changes

* removed leftovers

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/utils.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Clean up of docs and comments in code

* nit: lowercase p

* nit: lowercase p

---------

Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
2026-02-12 21:50:34 +01:00
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
Witold WydmańskiandGitHub c896438f21 feat: bump transformers to allow Qwen3-VL (#225)
Release / release (push) Waiting to run
2025-11-14 17:23:34 +01:00
Jhenner TigrerosandGitHub f5a063956e Fix requirements.txt to support gpt-oss models 2025-08-07 15:13:37 -05:00
Marut PandyaandGitHub 6b64bb93bd Merge pull request #147 from mohamednaji7/BitsAndBytes
completing the "bitsandbytes" option - based on  https://docs.vllm.ai/en/stable/quantization/bnb.html
2025-04-21 11:39:53 -07:00
Patrick RachfordandGitHub b2eef50c3b Update requirements.txt
Use the latest runpod version from here:
https://github.com/runpod/runpod-python/releases/tag/1.7.7
2025-03-10 08:46:28 -07:00
Mohamed NagyandGitHub 9299f43b8b solving "typing_extensions" compatibility with "bitsandbytes"
```
2025-01-22 18:04:01 [INFO] > [stage-0 5/8] RUN --mount=type=cache,target=/root/.cache/pip python3 -m pip install --upgrade pip && python3 -m pip install --upgrade -r /requirements.txt:
2025-01-22 18:04:01 [INFO] #13 8.904
2025-01-22 18:04:01 [INFO] #13 8.904 The conflict is caused by:
2025-01-22 18:04:01 [INFO] #13 8.904 The user requested typing-extensions==4.7.1
2025-01-22 18:04:01 [INFO] #13 8.904 bitsandbytes 0.45.0 depends on typing_extensions>=4.8.0
```
2025-01-22 18:06:46 +02:00
mohamednaji7 7167985f23 including bitsandbytes "src:https://docs.vllm.ai/en/stable/quantization/bnb.html" 2025-01-21 13:49:36 +02:00
Dean QuiñanolaandGitHub de2876e659 Update requirements.txt 2024-10-08 10:48:10 -07:00
pandyamarut 1420091588 update vll
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-27 19:23:50 -07:00
pandyamarutandGitHub 6a15a9e750 Update package version 2024-08-07 22:38:37 +00:00
alpayariyak 4abe494635 Fix hf-transfer error 2024-05-09 23:51:07 +00:00
alpayariyak c8ee100d80 Small refactor 2024-03-06 17:08:57 +00:00
alpayariyak e191149259 OpenAI Compatibility, Dynamic Batching, Refactor 2024-02-21 04:36:04 +00:00
alpayariyak 370698442c Update RunPod SDK version and Docker Tag 2024-01-31 00:54:58 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak a5dc8b53b7 Fix Concurrency, Add Max Model Length 2024-01-16 17:58:46 -05:00
Justin Merrell 6e7a1953e4 fix: update runpod package 2024-01-12 09:26:08 -05:00
alpayariyak cdda5edab3 Bump runpod version 2023-12-29 02:18:07 +00:00
alpayariyak dad02e13d8 vLLM 0.2.6 Stable Worker 2023-12-20 23:06:57 +00:00
alpayariyak e4e3ce4a6e Small fix 2023-12-20 08:45:48 +00:00
alpayariyak c59038e902 Fixed CUDA 11.8 workers 2023-12-20 07:50:14 +00:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00
justinmerrellandGitHub 15d9de0e10 Update package version 2023-12-14 21:07:42 +00:00
alpayariyak 5e38a1a2fa Multi-GPU Fix 2023-12-09 10:36:26 +00:00
alpayariyak b705544bd4 model packing, cuda version selection on build 2023-12-06 18:52:21 +00:00
alpayariyak 48c520b240 Version Bump + small changes 2023-12-06 16:48:12 +00:00
Justin Merrell 98af834124 fix: added hf_transfer requirement 2023-11-21 20:50:12 -05:00
Justin Merrell ff3ed55c9f Update requirements.txt 2023-11-21 20:14:51 -05:00
Justin Merrell 5ca8e56bbe fix: clean up extra 2023-11-21 19:18:08 -05:00
Samuel Willandalpayariyak 4f792062aa vllm 0.2.1.post1 - speed boost, quantization, mistral support 2023-11-14 19:12:58 -05:00
Jorg Doku 5bccd40952 using vllm 0.1.7 2023-09-21 09:41:59 -05:00
Jorg Doku 62961ce773 update vllm version 2023-08-29 16:37:24 -05:00
Jorg Doku 72be7ff31d update 2023-08-07 11:28:36 -05:00
Jorg Doku 4e4d3a27fc fixed 2023-08-06 14:17:53 -05:00
Jorg Doku 401a4637fa latest version 2023-08-03 10:05:51 -05:00
Jorg Doku 728e4f4f2e updating the handler 2023-07-31 14:04:34 -05:00
Jorg Doku c862d043dc latest 2023-07-28 23:14:03 -05:00
Jorg Doku 9fe114c049 llama 2 2023-07-20 18:45:08 -05:00