kldzj
3e86d16892
fix: openai tool calling
2024-12-06 13:08:48 +01:00
Sven Knoblauch
677a01e8f3
update code for case of no lora adapter
2024-10-09 14:51:33 +02:00
Sven Knoblauch
5cd12ba331
add changes for lora adapter support and /v1/models endpoint
2024-10-09 11:01:12 +02:00
pandyamarut
1420091588
update vll
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-09-27 19:23:50 -07:00
pandyamarut
814f50af38
fix oai completion api error
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-09-06 11:03:47 -07:00
pandyamarut
9cb9336cf5
update vllm version 0.5.4
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-08-09 12:01:42 -07:00
pandyamarut
f3534a4ea7
fix openai compat
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-07-27 15:49:45 -07:00
pandyamarut
bd96b5e0de
update v0.5.3.post1
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-07-25 21:29:45 -07:00
alpayariyak
5bd6f3a75e
0.5.3, any vllm arg as env var, refactor and fixes, moving away from building separate image from vLLM fork
2024-07-25 12:41:48 -07:00
alpayariyak
a08d83f600
Allow any vLLM engine args as env vars, refactor
2024-07-02 19:44:01 +00:00
alpayariyak
bad5ddd892
Fix vLLM 0.4.1 bug for OpenAI /models route
2024-06-10 20:38:49 +00:00
Alpay Ariyak and alpayariyak
874379a0c5
1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) ( #62 )
2024-05-09 00:42:08 -04:00
alpayariyak
d25b6f9628
Fix sampling params
2024-03-12 22:15:57 +00:00
alpayariyak
c8ee100d80
Small refactor
2024-03-06 17:08:57 +00:00
alpayariyak
985bbf1cb5
Fix Multi-GPU, tokenizer trust remote code
2024-02-24 03:18:25 +00:00
alpayariyak
708f68d7f8
Final Bug fixes, configurable oai response role, served model name override, documentation
2024-02-23 04:14:03 +00:00
alpayariyak
b42d45ce0f
Bug fixes, refactors
2024-02-23 03:46:47 +00:00
alpayariyak
aed0408f19
Documentation for 0.3.0, small fixes and changes
2024-02-21 19:28:59 -05:00
alpayariyak
e191149259
OpenAI Compatibility, Dynamic Batching, Refactor
2024-02-21 04:36:04 +00:00
alpayariyak
a94ef66f71
Dynamic Batch Size [needs refactor]
2024-02-06 03:47:27 +00:00
alpayariyak
fef8c81cb9
OpenAI Compatible worker, Refactor
2024-02-06 02:44:14 +00:00
alpayariyak
15b06bb687
Merge branch 'main' into openai-sse-output
2024-02-02 20:03:29 -05:00
alpayariyak
8de10468dd
Working tokenizer and model download fix
...
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak
b1720a154d
Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision
2024-01-31 22:57:32 -05:00
alpayariyak
afa33a2875
Handle errors
2024-02-01 03:01:51 +00:00
alpayariyak
3bbcf0021b
Fix: Yield if tokens left in batch
...
Temp: default model for testing image
2024-02-01 00:52:39 +00:00
alpayariyak
dab8bad906
OpenAI Chat Completions Stream
2024-01-31 23:27:31 +00:00
alpayariyak
3e2cd080a2
Fix Tensor Parallel
2024-01-31 05:12:58 +00:00
alpayariyak
46eee12819
Simplify Tensor Parallel
2024-01-31 03:28:02 +00:00
alpayariyak
12d6f0778e
Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer
2024-01-31 03:13:21 +00:00
alpayariyak
9fc8e1e54c
Non-streaming OpenAI Chat Completions
2024-01-25 23:24:18 -05:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak
584852f0f6
Docker-Protected HF Token, Refactor, Better Documentation
2024-01-18 18:20:51 -05:00
Alpay Ariyak and GitHub
ed48093792
bug fix
2024-01-16 18:38:26 -05:00
alpayariyak
a5dc8b53b7
Fix Concurrency, Add Max Model Length
2024-01-16 17:58:46 -05:00
alpayariyak
2475cd7a66
Local development, disable log requests
2023-12-29 02:21:24 +00:00
alpayariyak
ca8b02e392
Concurrency and Prompt Template fix
2023-12-29 00:42:40 +00:00
alpayariyak
a69ab1875d
vLLM job tracker, Refactor Concurrency Modifier, Serverless Config
2023-12-28 23:06:49 +00:00
alpayariyak
f7ac60f802
Fix Concurrency
2023-12-28 22:56:44 +00:00
alpayariyak
fe6e618692
Refactor code and Black formatting style
2023-12-22 04:08:55 +00:00