alpayariyak
e191149259
OpenAI Compatibility, Dynamic Batching, Refactor
2024-02-21 04:36:04 +00:00
alpayariyak
a94ef66f71
Dynamic Batch Size [needs refactor]
2024-02-06 03:47:27 +00:00
alpayariyak
45081e4037
Add __init__.py
2024-02-06 02:44:35 +00:00
alpayariyak
fef8c81cb9
OpenAI Compatible worker, Refactor
2024-02-06 02:44:14 +00:00
alpayariyak
15b06bb687
Merge branch 'main' into openai-sse-output
2024-02-02 20:03:29 -05:00
alpayariyak
8de10468dd
Working tokenizer and model download fix
...
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak
b7051d37ca
Move test_openai_stream.py
2024-02-02 22:12:26 +00:00
alpayariyak
b1720a154d
Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision
2024-01-31 22:57:32 -05:00
alpayariyak
afa33a2875
Handle errors
2024-02-01 03:01:51 +00:00
alpayariyak
068303ce8f
OpenAI Proxy Server for EndPoints and more examples
2024-02-01 02:32:32 +00:00
alpayariyak
3bbcf0021b
Fix: Yield if tokens left in batch
...
Temp: default model for testing image
2024-02-01 00:52:39 +00:00
alpayariyak
dab8bad906
OpenAI Chat Completions Stream
2024-01-31 23:27:31 +00:00
alpayariyak
370698442c
Update RunPod SDK version and Docker Tag
2024-01-31 00:54:58 -05:00
alpayariyak
3e2cd080a2
Fix Tensor Parallel
2024-01-31 05:12:58 +00:00
alpayariyak
e7b340d73b
Update vLLM base image
2024-01-31 03:37:55 +00:00
alpayariyak
46eee12819
Simplify Tensor Parallel
2024-01-31 03:28:02 +00:00
alpayariyak
12d6f0778e
Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer
2024-01-31 03:13:21 +00:00
alpayariyak
97726372c0
Update release tag in README.md
2024-01-25 23:37:07 -05:00
alpayariyak
9fc8e1e54c
Non-streaming OpenAI Chat Completions
2024-01-25 23:24:18 -05:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak
15cc7cd36f
Updated Documentation
2024-01-19 10:48:12 -05:00
alpayariyak
ef3c303743
Added support for n parameter
2024-01-19 15:34:25 +00:00
alpayariyak
584852f0f6
Docker-Protected HF Token, Refactor, Better Documentation
2024-01-18 18:20:51 -05:00
alpayariyak
a5dc8b53b7
Fix Concurrency, Add Max Model Length
2024-01-16 17:58:46 -05:00
alpayariyak
914caaa374
Merge branch 'refactor'
2024-01-05 23:44:59 +07:00
alpayariyak
618dd8e0c4
Temporarily remove examples
2024-01-05 16:44:12 +00:00
alpayariyak
3e08a69291
Documentation
2024-01-05 16:40:52 +00:00
alpayariyak
fa358708bc
Return finish status
2023-12-29 09:08:51 +00:00
alpayariyak
e23ab549e4
Small fix
2023-12-29 08:44:54 +00:00
alpayariyak
a14c5cd388
New Worker Stable
2023-12-29 08:21:09 +00:00
alpayariyak
2475cd7a66
Local development, disable log requests
2023-12-29 02:21:24 +00:00
alpayariyak
cdda5edab3
Bump runpod version
2023-12-29 02:18:07 +00:00
alpayariyak
ca8b02e392
Concurrency and Prompt Template fix
2023-12-29 00:42:40 +00:00
alpayariyak
a69ab1875d
vLLM job tracker, Refactor Concurrency Modifier, Serverless Config
2023-12-28 23:06:49 +00:00
alpayariyak
f7ac60f802
Fix Concurrency
2023-12-28 22:56:44 +00:00
alpayariyak
fe6e618692
Refactor code and Black formatting style
2023-12-22 04:08:55 +00:00
alpayariyak
1fc64fef2b
Additional Examples and Documentation for Chat Template and Messages list
2023-12-21 00:38:14 +00:00
alpayariyak
3f413c1038
Documentation for Chat Templates, Messages format, Refactor Streaming, etc
2023-12-21 00:22:23 +00:00
alpayariyak
ba893ed148
Refactor prompt handling in handler.py
2023-12-20 18:40:43 -05:00
alpayariyak
582e21f97f
Chat Template
2023-12-20 18:37:54 -05:00
alpayariyak
dad02e13d8
vLLM 0.2.6 Stable Worker
2023-12-20 23:06:57 +00:00
alpayariyak
e4e3ce4a6e
Small fix
2023-12-20 08:45:48 +00:00
alpayariyak
c59038e902
Fixed CUDA 11.8 workers
2023-12-20 07:50:14 +00:00
alpayariyak
8a11a7e4ed
Add version
2023-12-19 02:57:06 +00:00
alpayariyak
8202d4cc63
Bug fix
2023-12-19 02:56:21 +00:00
alpayariyak
f324bef8a0
Bump to vLLM 0.2.6, fixes and improvements
2023-12-19 01:52:55 +00:00
alpayariyak
d39844ac00
Small fixes
2023-12-13 13:50:36 +00:00
alpayariyak
bfa3703d54
Concurrency modifier
2023-12-13 12:26:01 +00:00
alpayariyak
22de0b322d
Complete refactor, improved overall functionality
2023-12-13 10:33:03 +00:00
alpayariyak
ef89868570
Small changes
2023-12-12 20:42:27 +00:00
alpayariyak
91c47790a6
Merge branch 'merge-ready' of https://github.com/runpod-workers/worker-vllm into merge-ready
2023-12-12 20:40:34 +00:00
alpayariyak
e81007d79b
bug fix
2023-12-09 11:10:29 -08:00
alpayariyak
b51827388c
bug fix
2023-12-09 19:06:41 +00:00
alpayariyak
8578581c6a
Small changes
2023-12-09 18:50:00 +00:00
alpayariyak
f1afbcbe4f
CPU Limit
2023-12-09 18:47:45 +00:00
alpayariyak
5e38a1a2fa
Multi-GPU Fix
2023-12-09 10:36:26 +00:00
alpayariyak
4052a17af1
Merge branch 'merge-ready' of https://github.com/runpod-workers/worker-vllm into merge-ready
2023-12-08 10:29:36 +00:00
alpayariyak
8156b1d8e0
Final changes
2023-12-08 10:29:24 +00:00
alpayariyak
5acbbf5528
Cleanup
2023-12-06 20:57:48 +00:00
alpayariyak
8dbd8f15dd
Token counting, cuda-related changes
2023-12-06 19:02:12 +00:00
alpayariyak
b705544bd4
model packing, cuda version selection on build
2023-12-06 18:52:21 +00:00
alpayariyak
48c520b240
Version Bump + small changes
2023-12-06 16:48:12 +00:00
alpayariyak
66c23d45b0
Model Downloading, no GPU
2023-12-06 01:43:09 +00:00
alpayariyak
b4c61e61c4
Packing model into worker for internal use
2023-12-05 20:26:56 -05:00
alpayariyak
3a9393bafd
Batched Tokens, cleanup
2023-12-05 17:46:28 +00:00
alpayariyak
8d27261645
Logging change
2023-12-01 10:18:48 +00:00
alpayariyak
b242b0cedd
Refactor handler + add non-streaming, add utils.py
2023-12-01 10:12:56 +00:00
alpayariyak
daf2f4d07b
Make handler stream aggregate_text
2023-12-01 04:21:03 +00:00
alpayariyak
86916bf628
Update handler.py
2023-11-29 23:37:17 -05:00
alpayariyak
820a21f32c
Download logic + other changes
2023-11-28 19:15:08 -05:00
alpayariyak
fdddeacf3c
Update handler.py
2023-11-28 18:49:15 -05:00
alpayariyak
7cadff29f5
edge case
2023-11-28 18:49:04 -05:00
alpayariyak
97480ff0ae
handler changes
2023-11-28 23:39:48 +00:00
alpayariyak
7f73f41c9f
Merge remote-tracking branch 'origin/new-concurrency' into revamp2
2023-11-26 18:23:50 -05:00
Alpay Ariyak and Ubuntu
96612413ce
worker changes
2023-11-22 00:39:52 +00:00
alpayariyak
c34d79ea49
Entrypoint for testing
2023-11-20 18:56:00 -05:00
alpayariyak
8f17764f8a
Update Dockerfile
2023-11-15 20:55:17 -05:00
alpayariyak
e7c8de4678
Changes to worker
2023-11-15 20:53:04 -05:00