30cb56a3df
Fix MODEL_REVISION env var (Merge pull request #50 from joennlae/rename-revision and #67 from mikljohansson/main)
...
Fixed MODEL_REVISION environment variable
Co-Authored-By: Jannis Schönleber <1493766+joennlae@users.noreply.github.com >
Co-Authored-By: Mikael Johansson <mikl.johansson@gmail.com >
2024-05-09 20:45:07 -04:00
alpayariyak
9f2cb7b1d0
Update documentation to include rename of fp8_e5m2 to fp8
2024-05-09 23:59:11 +00:00
Alpay Ariyak and GitHub
4f61b04afe
Update README.md
2024-05-09 01:05:04 -04:00
Alpay Ariyak and alpayariyak
874379a0c5
1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) ( #62 )
2024-05-09 00:42:08 -04:00
alpayariyak
0a5b5bc095
Add badges
2024-03-15 21:23:52 -04:00
alpayariyak
2936e4d95d
Update automatic builds, documentation
2024-03-15 11:57:38 -04:00
Alpay Ariyak and GitHub
cee4e484d5
Update README.md for 0.3.2
2024-03-12 19:07:37 -04:00
alpayariyak
db7167d57f
0.3.3
2024-03-05 19:14:35 +00:00
Alpay Ariyak and alpayariyak
d91ccb866f
0.3.1: bug fixes
2024-02-29 02:55:44 -05:00
Alpay Ariyak and GitHub
36e9b670ee
Add notice on what to do when HuggingFace is down
2024-02-28 17:10:39 -05:00
Alpay Ariyak and alpayariyak
91167b873a
v0.3.0: OpenAI Compatibility, Dynamic Stream Batching, Refactor, Error Catching
2024-02-23 22:34:00 -05:00
Alpay Ariyak and GitHub
6f5718f191
Update Docker image tag in README.md
2024-02-23 02:15:50 -05:00
alpayariyak
708f68d7f8
Final Bug fixes, configurable oai response role, served model name override, documentation
2024-02-23 04:14:03 +00:00
alpayariyak
a2d9535652
Update documentation further
2024-02-23 01:34:38 +00:00
alpayariyak
9129d0a252
New ENV Vars
2024-02-22 23:58:20 +00:00
alpayariyak
6bcd9d7c67
Preparing for 0.3.0
2024-02-22 18:04:34 -05:00
Patrick Rachford and GitHub
7993818f5f
Update README.md
...
Remove formatting of tables
2024-02-21 19:28:03 -08:00
Patrick Rachford and GitHub
e97917cc14
Fixes import statement
...
Fixes import statements, formats tables, run black on code blocks
2024-02-21 19:17:23 -08:00
Alpay Ariyak and GitHub
3549cf24d5
Merge branch 'main' into openai-sse-output
2024-02-21 19:56:55 -05:00
alpayariyak
aed0408f19
Documentation for 0.3.0, small fixes and changes
2024-02-21 19:28:59 -05:00
Alpay Ariyak and GitHub
2941db0fb8
Update worker-vllm version
2024-02-09 23:12:29 -05:00
alpayariyak
8de10468dd
Working tokenizer and model download fix
...
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak
b1720a154d
Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision
2024-01-31 22:57:32 -05:00
alpayariyak
370698442c
Update RunPod SDK version and Docker Tag
2024-01-31 00:54:58 -05:00
alpayariyak
3e2cd080a2
Fix Tensor Parallel
2024-01-31 05:12:58 +00:00
alpayariyak
46eee12819
Simplify Tensor Parallel
2024-01-31 03:28:02 +00:00
alpayariyak
12d6f0778e
Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer
2024-01-31 03:13:21 +00:00
alpayariyak
97726372c0
Update release tag in README.md
2024-01-25 23:37:07 -05:00
alpayariyak
9fc8e1e54c
Non-streaming OpenAI Chat Completions
2024-01-25 23:24:18 -05:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak
15cc7cd36f
Updated Documentation
2024-01-19 10:48:12 -05:00
alpayariyak
584852f0f6
Docker-Protected HF Token, Refactor, Better Documentation
2024-01-18 18:20:51 -05:00
Alpay Ariyak and GitHub
65454c024a
Update README.md (temporary)
2024-01-17 00:16:56 -05:00
Justin Merrell
f023e2e097
Update README.md
2024-01-16 19:55:20 -05:00
alpayariyak
a5dc8b53b7
Fix Concurrency, Add Max Model Length
2024-01-16 17:58:46 -05:00
Alpay Ariyak and GitHub
23ff046586
Update README.md
2024-01-11 12:48:54 +07:00
alpayariyak
914caaa374
Merge branch 'refactor'
2024-01-05 23:44:59 +07:00
alpayariyak
618dd8e0c4
Temporarily remove examples
2024-01-05 16:44:12 +00:00
alpayariyak
3e08a69291
Documentation
2024-01-05 16:40:52 +00:00
alpayariyak
2475cd7a66
Local development, disable log requests
2023-12-29 02:21:24 +00:00
Alpay Ariyak and GitHub
a247a3afe1
Update README.md
2023-12-22 03:16:04 -05:00
alpayariyak
1fc64fef2b
Additional Examples and Documentation for Chat Template and Messages list
2023-12-21 00:38:14 +00:00
alpayariyak
3f413c1038
Documentation for Chat Templates, Messages format, Refactor Streaming, etc
2023-12-21 00:22:23 +00:00
alpayariyak
8a11a7e4ed
Add version
2023-12-19 02:57:06 +00:00
alpayariyak
f324bef8a0
Bump to vLLM 0.2.6, fixes and improvements
2023-12-19 01:52:55 +00:00
Alpay Ariyak and GitHub
d6c145bc1e
Update README.md
2023-12-18 20:08:35 -05:00
Alpay Ariyak and GitHub
840d5ff0d2
Update README.md
2023-12-16 01:30:49 -05:00
Justin Merrell
06660fb8b9
fix: update badge
2023-12-14 16:26:52 -05:00
alpayariyak
22de0b322d
Complete refactor, improved overall functionality
2023-12-13 10:33:03 +00:00
Alpay Ariyak and GitHub
2dbab5c254
Update README.md
2023-12-12 15:34:47 -05:00