c45ac42acd
vLLM Worker v0.15.0 — Upgrade from v0.11.x to v0.15.0 ( #259 )
...
Release / release (push) Waiting to run
* VLLM upgrade to 0.12.0 and compatibility fixes
* MAX_NUM_BATCHED_TOKENS fix and CUDA tester
* Sys kill worker instead of marking as failed
* upgrade to vllm 0.12.0
* Update to vllm 0.15.0 and lora fix
* Update for HUB and removal of deprected env variables
* reverted docker-bake changes
* removed leftovers
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Update src/utils.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com >
* Clean up of docs and comments in code
* nit: lowercase p
* nit: lowercase p
---------
Co-authored-by: Dj Isaac <contact@dejaydev.com >
Co-authored-by: chrisvela <chris.vela@runpod.io >
2026-02-12 21:50:34 +01:00
alpayariyak
e191149259
OpenAI Compatibility, Dynamic Batching, Refactor
2024-02-21 04:36:04 +00:00
alpayariyak
fef8c81cb9
OpenAI Compatible worker, Refactor
2024-02-06 02:44:14 +00:00
alpayariyak
3bbcf0021b
Fix: Yield if tokens left in batch
...
Temp: default model for testing image
2024-02-01 00:52:39 +00:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak
ef3c303743
Added support for n parameter
2024-01-19 15:34:25 +00:00
alpayariyak
584852f0f6
Docker-Protected HF Token, Refactor, Better Documentation
2024-01-18 18:20:51 -05:00
alpayariyak
a5dc8b53b7
Fix Concurrency, Add Max Model Length
2024-01-16 17:58:46 -05:00
alpayariyak
fa358708bc
Return finish status
2023-12-29 09:08:51 +00:00
alpayariyak
e23ab549e4
Small fix
2023-12-29 08:44:54 +00:00
alpayariyak
a14c5cd388
New Worker Stable
2023-12-29 08:21:09 +00:00
alpayariyak
2475cd7a66
Local development, disable log requests
2023-12-29 02:21:24 +00:00
alpayariyak
ca8b02e392
Concurrency and Prompt Template fix
2023-12-29 00:42:40 +00:00
alpayariyak
a69ab1875d
vLLM job tracker, Refactor Concurrency Modifier, Serverless Config
2023-12-28 23:06:49 +00:00
alpayariyak
f7ac60f802
Fix Concurrency
2023-12-28 22:56:44 +00:00
alpayariyak
fe6e618692
Refactor code and Black formatting style
2023-12-22 04:08:55 +00:00
alpayariyak
3f413c1038
Documentation for Chat Templates, Messages format, Refactor Streaming, etc
2023-12-21 00:22:23 +00:00
alpayariyak
ba893ed148
Refactor prompt handling in handler.py
2023-12-20 18:40:43 -05:00
alpayariyak
582e21f97f
Chat Template
2023-12-20 18:37:54 -05:00
alpayariyak
d39844ac00
Small fixes
2023-12-13 13:50:36 +00:00
alpayariyak
bfa3703d54
Concurrency modifier
2023-12-13 12:26:01 +00:00
alpayariyak
22de0b322d
Complete refactor, improved overall functionality
2023-12-13 10:33:03 +00:00
alpayariyak
8156b1d8e0
Final changes
2023-12-08 10:29:24 +00:00
alpayariyak
5acbbf5528
Cleanup
2023-12-06 20:57:48 +00:00
alpayariyak
8dbd8f15dd
Token counting, cuda-related changes
2023-12-06 19:02:12 +00:00
alpayariyak
3a9393bafd
Batched Tokens, cleanup
2023-12-05 17:46:28 +00:00
alpayariyak
8d27261645
Logging change
2023-12-01 10:18:48 +00:00
alpayariyak
b242b0cedd
Refactor handler + add non-streaming, add utils.py
2023-12-01 10:12:56 +00:00
alpayariyak
daf2f4d07b
Make handler stream aggregate_text
2023-12-01 04:21:03 +00:00
Justin Merrell
8dbf197362
Update handler.py
2023-11-30 11:00:16 -05:00
alpayariyak
86916bf628
Update handler.py
2023-11-29 23:37:17 -05:00
Justin Merrell
f958014d10
Update handler.py
2023-11-29 22:41:59 -05:00
Justin Merrell
c455508931
Update handler.py
2023-11-29 22:01:17 -05:00
Justin Merrell
9b9c1cf503
cleanup
2023-11-29 21:32:40 -05:00
alpayariyak
820a21f32c
Download logic + other changes
2023-11-28 19:15:08 -05:00
alpayariyak
fdddeacf3c
Update handler.py
2023-11-28 18:49:15 -05:00
alpayariyak
7cadff29f5
edge case
2023-11-28 18:49:04 -05:00
alpayariyak
97480ff0ae
handler changes
2023-11-28 23:39:48 +00:00
Alpay Ariyak and Ubuntu
96612413ce
worker changes
2023-11-22 00:39:52 +00:00
alpayariyak
e7c8de4678
Changes to worker
2023-11-15 20:53:04 -05:00
Samuel Will and alpayariyak
4f792062aa
vllm 0.2.1.post1 - speed boost, quantization, mistral support
2023-11-14 19:12:58 -05:00
Jorg Doku
ba4e6e6df0
fix aggre
2023-09-05 23:51:35 -05:00
Jorg Doku
871196dafb
update
2023-09-05 22:06:46 -05:00
Jorg Doku
5e04e5ebb0
print out llm metrics, make metrics more verbose
2023-08-31 19:09:04 -05:00
Jorg Doku
54d229cb16
upate metrics
2023-08-31 18:26:37 -05:00
Jorg Doku
96c4d70d2e
metrics now included
2023-08-31 17:03:59 -05:00
Jorg Doku
c5662663ce
update response
2023-08-31 16:00:28 -05:00
Jorg Doku
9f1dd5a9b1
finished
2023-08-31 15:56:13 -05:00
Jorg Doku
9eb099ded2
fix bug
2023-08-31 14:42:18 -05:00
Jorg Doku
460f3dd348
update the handler
2023-08-31 13:48:57 -05:00