Samuel Will
b0e7b575f3
fix: default value for tokenizer
2024-02-09 10:24:58 +00:00
alpayariyak
8de10468dd
Working tokenizer and model download fix
...
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak
370698442c
Update RunPod SDK version and Docker Tag
2024-01-31 00:54:58 -05:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak
a5dc8b53b7
Fix Concurrency, Add Max Model Length
2024-01-16 17:58:46 -05:00
Justin Merrell
6e7a1953e4
fix: update runpod package
2024-01-12 09:26:08 -05:00
alpayariyak
cdda5edab3
Bump runpod version
2023-12-29 02:18:07 +00:00
alpayariyak
dad02e13d8
vLLM 0.2.6 Stable Worker
2023-12-20 23:06:57 +00:00
alpayariyak
e4e3ce4a6e
Small fix
2023-12-20 08:45:48 +00:00
alpayariyak
c59038e902
Fixed CUDA 11.8 workers
2023-12-20 07:50:14 +00:00
alpayariyak
f324bef8a0
Bump to vLLM 0.2.6, fixes and improvements
2023-12-19 01:52:55 +00:00
justinmerrell and GitHub
15d9de0e10
Update package version
2023-12-14 21:07:42 +00:00
alpayariyak
5e38a1a2fa
Multi-GPU Fix
2023-12-09 10:36:26 +00:00
alpayariyak
b705544bd4
model packing, cuda version selection on build
2023-12-06 18:52:21 +00:00
alpayariyak
48c520b240
Version Bump + small changes
2023-12-06 16:48:12 +00:00
Justin Merrell
98af834124
fix: added hf_transfer requirement
2023-11-21 20:50:12 -05:00
Justin Merrell
ff3ed55c9f
Update requirements.txt
2023-11-21 20:14:51 -05:00
Justin Merrell
5ca8e56bbe
fix: clean up extra
2023-11-21 19:18:08 -05:00
alpayariyak
e7c8de4678
Changes to worker
2023-11-15 20:53:04 -05:00
Samuel Will and alpayariyak
4f792062aa
vllm 0.2.1.post1 - speed boost, quantization, mistral support
2023-11-14 19:12:58 -05:00
Jorg Doku
5bccd40952
using vllm 0.1.7
2023-09-21 09:41:59 -05:00
Jorg Doku
62961ce773
update vllm version
2023-08-29 16:37:24 -05:00
Jorg Doku
72be7ff31d
update
2023-08-07 11:28:36 -05:00
Jorg Doku
4e4d3a27fc
fixed
2023-08-06 14:17:53 -05:00
Jorg Doku
401a4637fa
latest version
2023-08-03 10:05:51 -05:00
Jorg Doku
728e4f4f2e
updating the handler
2023-07-31 14:04:34 -05:00
Jorg Doku
c862d043dc
latest
2023-07-28 23:14:03 -05:00
Jorg Doku
9fe114c049
llama 2
2023-07-20 18:45:08 -05:00
Jorg Doku
ab3f0e1c58
update worker
2023-07-14 11:17:25 -05:00
Jorg Doku
be65b7a99c
update the worker
2023-07-14 00:46:14 -05:00
Jorg Doku
488a9a9451
temp
2023-07-05 22:34:26 -05:00
Jorg Doku
c91fc100cd
push
2023-07-05 22:32:32 -05:00
Jorg Doku
9c141ff30f
fix handler & req
2023-07-05 17:16:13 -05:00
Jorg Doku
c8ec26a782
update docker container
2023-07-05 15:47:45 -05:00
Jorg Doku
9fe3d4a09d
cleaning up handler for users
2023-07-05 13:20:08 -05:00
Jorg Doku
68184d424b
testing vllm
2023-07-05 03:51:21 -05:00
Jorg Doku and GitHub
38829c2720
Initial commit
2023-07-03 13:42:43 -05:00