Commit Graph
55 Commits
Author SHA1 Message Date
velaraptor-runpod d9808815ee feat: update requirements.txt 2026-03-17 19:56:39 -05:00
c45ac42acd vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 (#259)
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes

* MAX_NUM_BATCHED_TOKENS fix and CUDA tester

* Sys kill worker instead of marking as failed

* upgrade to vllm 0.12.0

* Update to vllm 0.15.0 and lora fix

* Update for HUB and removal of deprected env variables

* reverted docker-bake changes

* removed leftovers

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/utils.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Update src/handler.py

Co-authored-by: Dj Isaac <contact@dejaydev.com>

* Clean up of docs and comments in code

* nit: lowercase p

* nit: lowercase p

---------

Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
2026-02-12 21:50:34 +01:00
Tim PietruskyandGitHub 6f2381a9a1 chore(deps): update runpod to latest version (#242)
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
Witold WydmańskiandGitHub c896438f21 feat: bump transformers to allow Qwen3-VL (#225)
Release / release (push) Waiting to run
2025-11-14 17:23:34 +01:00
Jhenner TigrerosandGitHub f5a063956e Fix requirements.txt to support gpt-oss models 2025-08-07 15:13:37 -05:00
Marut PandyaandGitHub 6b64bb93bd Merge pull request #147 from mohamednaji7/BitsAndBytes
completing the "bitsandbytes" option - based on  https://docs.vllm.ai/en/stable/quantization/bnb.html
2025-04-21 11:39:53 -07:00
Patrick RachfordandGitHub b2eef50c3b Update requirements.txt
Use the latest runpod version from here:
https://github.com/runpod/runpod-python/releases/tag/1.7.7
2025-03-10 08:46:28 -07:00
Mohamed NagyandGitHub 9299f43b8b solving "typing_extensions" compatibility with "bitsandbytes"
```
2025-01-22 18:04:01 [INFO] > [stage-0 5/8] RUN --mount=type=cache,target=/root/.cache/pip python3 -m pip install --upgrade pip && python3 -m pip install --upgrade -r /requirements.txt:
2025-01-22 18:04:01 [INFO] #13 8.904
2025-01-22 18:04:01 [INFO] #13 8.904 The conflict is caused by:
2025-01-22 18:04:01 [INFO] #13 8.904 The user requested typing-extensions==4.7.1
2025-01-22 18:04:01 [INFO] #13 8.904 bitsandbytes 0.45.0 depends on typing_extensions>=4.8.0
```
2025-01-22 18:06:46 +02:00
mohamednaji7 7167985f23 including bitsandbytes "src:https://docs.vllm.ai/en/stable/quantization/bnb.html" 2025-01-21 13:49:36 +02:00
Dean QuiñanolaandGitHub de2876e659 Update requirements.txt 2024-10-08 10:48:10 -07:00
pandyamarut 1420091588 update vll
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-27 19:23:50 -07:00
pandyamarutandGitHub 6a15a9e750 Update package version 2024-08-07 22:38:37 +00:00
alpayariyak 4abe494635 Fix hf-transfer error 2024-05-09 23:51:07 +00:00
Alpay Ariyakandalpayariyak 874379a0c5 1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) (#62) 2024-05-09 00:42:08 -04:00
alpayariyak c8ee100d80 Small refactor 2024-03-06 17:08:57 +00:00
Alpay AriyakandGitHub 3549cf24d5 Merge branch 'main' into openai-sse-output 2024-02-21 19:56:55 -05:00
alpayariyak e191149259 OpenAI Compatibility, Dynamic Batching, Refactor 2024-02-21 04:36:04 +00:00
alpayariyak 7b3fd05542 Small refactor to tokenizer fix 2024-02-09 22:50:00 -05:00
Samuel Will b0e7b575f3 fix: default value for tokenizer 2024-02-09 10:24:58 +00:00
alpayariyak 8de10468dd Working tokenizer and model download fix
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak 370698442c Update RunPod SDK version and Docker Tag 2024-01-31 00:54:58 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak a5dc8b53b7 Fix Concurrency, Add Max Model Length 2024-01-16 17:58:46 -05:00
Justin Merrell 6e7a1953e4 fix: update runpod package 2024-01-12 09:26:08 -05:00
alpayariyak cdda5edab3 Bump runpod version 2023-12-29 02:18:07 +00:00
alpayariyak dad02e13d8 vLLM 0.2.6 Stable Worker 2023-12-20 23:06:57 +00:00
alpayariyak e4e3ce4a6e Small fix 2023-12-20 08:45:48 +00:00
alpayariyak c59038e902 Fixed CUDA 11.8 workers 2023-12-20 07:50:14 +00:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00
justinmerrellandGitHub 15d9de0e10 Update package version 2023-12-14 21:07:42 +00:00
alpayariyak 5e38a1a2fa Multi-GPU Fix 2023-12-09 10:36:26 +00:00
alpayariyak b705544bd4 model packing, cuda version selection on build 2023-12-06 18:52:21 +00:00
alpayariyak 48c520b240 Version Bump + small changes 2023-12-06 16:48:12 +00:00
Justin Merrell 98af834124 fix: added hf_transfer requirement 2023-11-21 20:50:12 -05:00
Justin Merrell ff3ed55c9f Update requirements.txt 2023-11-21 20:14:51 -05:00
Justin Merrell 5ca8e56bbe fix: clean up extra 2023-11-21 19:18:08 -05:00
alpayariyak e7c8de4678 Changes to worker 2023-11-15 20:53:04 -05:00
Samuel Willandalpayariyak 4f792062aa vllm 0.2.1.post1 - speed boost, quantization, mistral support 2023-11-14 19:12:58 -05:00
Jorg Doku 5bccd40952 using vllm 0.1.7 2023-09-21 09:41:59 -05:00
Jorg Doku 62961ce773 update vllm version 2023-08-29 16:37:24 -05:00
Jorg Doku 72be7ff31d update 2023-08-07 11:28:36 -05:00
Jorg Doku 4e4d3a27fc fixed 2023-08-06 14:17:53 -05:00
Jorg Doku 401a4637fa latest version 2023-08-03 10:05:51 -05:00
Jorg Doku 728e4f4f2e updating the handler 2023-07-31 14:04:34 -05:00
Jorg Doku c862d043dc latest 2023-07-28 23:14:03 -05:00
Jorg Doku 9fe114c049 llama 2 2023-07-20 18:45:08 -05:00
Jorg Doku ab3f0e1c58 update worker 2023-07-14 11:17:25 -05:00
Jorg Doku be65b7a99c update the worker 2023-07-14 00:46:14 -05:00
Jorg Doku 488a9a9451 temp 2023-07-05 22:34:26 -05:00
Jorg Doku c91fc100cd push 2023-07-05 22:32:32 -05:00