alpayariyak
db7167d57f
0.3.3
2024-03-05 19:14:35 +00:00
Alpay Ariyak and alpayariyak
d91ccb866f
0.3.1: bug fixes
0.3.1
2024-02-29 02:55:44 -05:00
Alpay Ariyak and GitHub
36e9b670ee
Add notice on what to do when HuggingFace is down
2024-02-28 17:10:39 -05:00
Alpay Ariyak and alpayariyak
91167b873a
v0.3.0: OpenAI Compatibility, Dynamic Stream Batching, Refactor, Error Catching
0.3.0
2024-02-23 22:34:00 -05:00
alpayariyak
819102cfd6
Merge branch 'openai-sse-output' of https://github.com/runpod-workers/worker-vllm into openai-sse-output
2024-02-24 03:18:37 +00:00
alpayariyak
985bbf1cb5
Fix Multi-GPU, tokenizer trust remote code
2024-02-24 03:18:25 +00:00
Alpay Ariyak and GitHub
6f5718f191
Update Docker image tag in README.md
2024-02-23 02:15:50 -05:00
Alpay Ariyak and GitHub
235d0d31d0
Merge pull request #50 from runpod-workers/tagged-releases
...
feat: auto build cuda version
2024-02-22 23:14:57 -05:00
alpayariyak
708f68d7f8
Final Bug fixes, configurable oai response role, served model name override, documentation
2024-02-23 04:14:03 +00:00
alpayariyak
b42d45ce0f
Bug fixes, refactors
2024-02-23 03:46:47 +00:00
Justin Merrell
7221caceff
feat: auto build cuda version
2024-02-22 20:48:54 -05:00
alpayariyak
a2d9535652
Update documentation further
2024-02-23 01:34:38 +00:00
alpayariyak
9129d0a252
New ENV Vars
2024-02-22 23:58:20 +00:00
alpayariyak
6bcd9d7c67
Preparing for 0.3.0
2024-02-22 18:04:34 -05:00
Alpay Ariyak and GitHub
b4204c612c
Merge pull request #48 from rachfop/patch-1
...
Fixes import statement in docs
2024-02-21 23:55:11 -05:00
Patrick Rachford and GitHub
7993818f5f
Update README.md
...
Remove formatting of tables
2024-02-21 19:28:03 -08:00
Patrick Rachford and GitHub
e97917cc14
Fixes import statement
...
Fixes import statements, formats tables, run black on code blocks
2024-02-21 19:17:23 -08:00
Alpay Ariyak and GitHub
3549cf24d5
Merge branch 'main' into openai-sse-output
2024-02-21 19:56:55 -05:00
alpayariyak
aed0408f19
Documentation for 0.3.0, small fixes and changes
2024-02-21 19:28:59 -05:00
alpayariyak
e191149259
OpenAI Compatibility, Dynamic Batching, Refactor
2024-02-21 04:36:04 +00:00
Alpay Ariyak and GitHub
2941db0fb8
Update worker-vllm version
0.2.3
2024-02-09 23:12:29 -05:00
Alpay Ariyak and GitHub
bfeb60c54e
Merge pull request #45 from willsamu/fix-tokenizer-input
...
fix: build error if no `TOKENIZER_NAME` provided
2024-02-09 22:51:18 -05:00
alpayariyak
7b3fd05542
Small refactor to tokenizer fix
2024-02-09 22:50:00 -05:00
Samuel Will
b0e7b575f3
fix: default value for tokenizer
2024-02-09 10:24:58 +00:00
alpayariyak
4f5e0d37c4
Fix tokenizer's trust_remote_code parameter
2024-02-08 23:53:13 +00:00
alpayariyak
a94ef66f71
Dynamic Batch Size [needs refactor]
2024-02-06 03:47:27 +00:00
alpayariyak
45081e4037
Add __init__.py
2024-02-06 02:44:35 +00:00
alpayariyak
fef8c81cb9
OpenAI Compatible worker, Refactor
2024-02-06 02:44:14 +00:00
alpayariyak
15b06bb687
Merge branch 'main' into openai-sse-output
2024-02-02 20:03:29 -05:00
Alpay Ariyak and GitHub
2b5b8dfb61
Fix Model and Tokenizer download for bake-in option, add revision configuration for both.
2024-02-02 19:58:50 -05:00
alpayariyak
8de10468dd
Working tokenizer and model download fix
...
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak
b7051d37ca
Move test_openai_stream.py
2024-02-02 22:12:26 +00:00
alpayariyak
b1720a154d
Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision
2024-01-31 22:57:32 -05:00
alpayariyak
afa33a2875
Handle errors
2024-02-01 03:01:51 +00:00
alpayariyak
068303ce8f
OpenAI Proxy Server for EndPoints and more examples
2024-02-01 02:32:32 +00:00
alpayariyak
3bbcf0021b
Fix: Yield if tokens left in batch
...
Temp: default model for testing image
2024-02-01 00:52:39 +00:00
alpayariyak
dab8bad906
OpenAI Chat Completions Stream
2024-01-31 23:27:31 +00:00
Casper
f4d7c75504
Snapshot download only tokenizer/config related things
2024-01-31 22:16:44 +01:00
Casper
3adc9e3336
Remove unused import
2024-01-31 18:27:49 +01:00
Casper
fd00a1ece3
Update to use snapshot_download
2024-01-31 18:24:57 +01:00
Casper
664dd35782
Download tokenizer upon build
2024-01-31 17:59:52 +01:00
alpayariyak
370698442c
Update RunPod SDK version and Docker Tag
2024-01-31 00:54:58 -05:00
alpayariyak
3e2cd080a2
Fix Tensor Parallel
0.2.2
2024-01-31 05:12:58 +00:00
alpayariyak
e7b340d73b
Update vLLM base image
2024-01-31 03:37:55 +00:00
alpayariyak
46eee12819
Simplify Tensor Parallel
2024-01-31 03:28:02 +00:00
alpayariyak
12d6f0778e
Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer
2024-01-31 03:13:21 +00:00
Alpay Ariyak and GitHub
fa5556434c
Bug fix
2024-01-29 11:26:22 -05:00
alpayariyak
97726372c0
Update release tag in README.md
2024-01-25 23:37:07 -05:00
alpayariyak
9fc8e1e54c
Non-streaming OpenAI Chat Completions
0.2.1
2024-01-25 23:24:18 -05:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
0.2.0
2024-01-25 20:49:15 -05:00