velaraptor-runpod
2e8c251447
Merge branch 'main' into feat/update-vllm-v0.15.0
2026-02-13 03:01:05 -06:00
c45ac42acd
vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 ( #259 )
...
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes
* MAX_NUM_BATCHED_TOKENS fix and CUDA tester
* Sys kill worker instead of marking as failed
* upgrade to vllm 0.12.0
* Update to vllm 0.15.0 and lora fix
* Update for HUB and removal of deprected env variables
* reverted docker-bake changes
* removed leftovers
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Update src/utils.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Clean up of docs and comments in code
* nit: lowercase p
* nit: lowercase p
---------
Co-authored-by: Dj Isaac <contact@dejaydev.com >
Co-authored-by: chrisvela <chris.vela@runpod.io >
2026-02-12 21:50:34 +01:00
velaraptor-runpod
340bc0b3c6
fix: served model name
2026-02-10 21:42:58 -06:00
velaraptor-runpod
e1e9ef74ad
add changes from pr
2026-02-06 18:10:09 -06:00
velaraptor-runpod
461f89cea6
add torch-c-dlpack-ext requirement
2026-02-06 17:03:39 -06:00
velaraptor-runpod
8eb55b90c1
add changes for v0.15.0
2026-02-05 17:24:16 -06:00
Tim Pietrusky and GitHub
6f2381a9a1
chore(deps): update runpod to latest version ( #242 )
Release / release (push) Waiting to run
2025-11-24 16:42:21 +01:00
Witold Wydmański and GitHub
c896438f21
feat: bump transformers to allow Qwen3-VL ( #225 )
Release / release (push) Waiting to run
2025-11-14 17:23:34 +01:00
Jhenner Tigreros and GitHub
f5a063956e
Fix requirements.txt to support gpt-oss models
2025-08-07 15:13:37 -05:00
Marut Pandya and GitHub
6b64bb93bd
Merge pull request #147 from mohamednaji7/BitsAndBytes
...
completing the "bitsandbytes" option - based on https://docs.vllm.ai/en/stable/quantization/bnb.html
2025-04-21 11:39:53 -07:00
Patrick Rachford and GitHub
b2eef50c3b
Update requirements.txt
...
Use the latest runpod version from here:
https://github.com/runpod/runpod-python/releases/tag/1.7.7
2025-03-10 08:46:28 -07:00
Mohamed Nagy and GitHub
9299f43b8b
solving "typing_extensions" compatibility with "bitsandbytes"
...
```
2025-01-22 18:04:01 [INFO] > [stage-0 5/8] RUN --mount=type=cache,target=/root/.cache/pip python3 -m pip install --upgrade pip && python3 -m pip install --upgrade -r /requirements.txt:
2025-01-22 18:04:01 [INFO] #13 8.904
2025-01-22 18:04:01 [INFO] #13 8.904 The conflict is caused by:
2025-01-22 18:04:01 [INFO] #13 8.904 The user requested typing-extensions==4.7.1
2025-01-22 18:04:01 [INFO] #13 8.904 bitsandbytes 0.45.0 depends on typing_extensions>=4.8.0
```
2025-01-22 18:06:46 +02:00
mohamednaji7
7167985f23
including bitsandbytes "src: https://docs.vllm.ai/en/stable/quantization/bnb.html "
2025-01-21 13:49:36 +02:00
Dean Quiñanola and GitHub
de2876e659
Update requirements.txt
2024-10-08 10:48:10 -07:00
pandyamarut
1420091588
update vll
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-09-27 19:23:50 -07:00
pandyamarut and GitHub
6a15a9e750
Update package version
2024-08-07 22:38:37 +00:00
alpayariyak
4abe494635
Fix hf-transfer error
2024-05-09 23:51:07 +00:00
Alpay Ariyak and alpayariyak
874379a0c5
1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) ( #62 )
2024-05-09 00:42:08 -04:00
alpayariyak
c8ee100d80
Small refactor
2024-03-06 17:08:57 +00:00
Alpay Ariyak and GitHub
3549cf24d5
Merge branch 'main' into openai-sse-output
2024-02-21 19:56:55 -05:00
alpayariyak
e191149259
OpenAI Compatibility, Dynamic Batching, Refactor
2024-02-21 04:36:04 +00:00
alpayariyak
7b3fd05542
Small refactor to tokenizer fix
2024-02-09 22:50:00 -05:00
Samuel Will
b0e7b575f3
fix: default value for tokenizer
2024-02-09 10:24:58 +00:00
alpayariyak
8de10468dd
Working tokenizer and model download fix
...
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak
370698442c
Update RunPod SDK version and Docker Tag
2024-01-31 00:54:58 -05:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak
a5dc8b53b7
Fix Concurrency, Add Max Model Length
2024-01-16 17:58:46 -05:00
Justin Merrell
6e7a1953e4
fix: update runpod package
2024-01-12 09:26:08 -05:00
alpayariyak
cdda5edab3
Bump runpod version
2023-12-29 02:18:07 +00:00
alpayariyak
dad02e13d8
vLLM 0.2.6 Stable Worker
2023-12-20 23:06:57 +00:00
alpayariyak
e4e3ce4a6e
Small fix
2023-12-20 08:45:48 +00:00
alpayariyak
c59038e902
Fixed CUDA 11.8 workers
2023-12-20 07:50:14 +00:00
alpayariyak
f324bef8a0
Bump to vLLM 0.2.6, fixes and improvements
2023-12-19 01:52:55 +00:00
justinmerrell and GitHub
15d9de0e10
Update package version
2023-12-14 21:07:42 +00:00
alpayariyak
5e38a1a2fa
Multi-GPU Fix
2023-12-09 10:36:26 +00:00
alpayariyak
b705544bd4
model packing, cuda version selection on build
2023-12-06 18:52:21 +00:00
alpayariyak
48c520b240
Version Bump + small changes
2023-12-06 16:48:12 +00:00
Justin Merrell
98af834124
fix: added hf_transfer requirement
2023-11-21 20:50:12 -05:00
Justin Merrell
ff3ed55c9f
Update requirements.txt
2023-11-21 20:14:51 -05:00
Justin Merrell
5ca8e56bbe
fix: clean up extra
2023-11-21 19:18:08 -05:00
alpayariyak
e7c8de4678
Changes to worker
2023-11-15 20:53:04 -05:00
Samuel Will and alpayariyak
4f792062aa
vllm 0.2.1.post1 - speed boost, quantization, mistral support
2023-11-14 19:12:58 -05:00
Jorg Doku
5bccd40952
using vllm 0.1.7
2023-09-21 09:41:59 -05:00
Jorg Doku
62961ce773
update vllm version
2023-08-29 16:37:24 -05:00
Jorg Doku
72be7ff31d
update
2023-08-07 11:28:36 -05:00
Jorg Doku
4e4d3a27fc
fixed
2023-08-06 14:17:53 -05:00
Jorg Doku
401a4637fa
latest version
2023-08-03 10:05:51 -05:00
Jorg Doku
728e4f4f2e
updating the handler
2023-07-31 14:04:34 -05:00
Jorg Doku
c862d043dc
latest
2023-07-28 23:14:03 -05:00
Jorg Doku
9fe114c049
llama 2
2023-07-20 18:45:08 -05:00