WorstDev01
a58a76783d
Add multimodal limit parameter support
...
- Updated convert_limit_mm_per_prompt to handle multiple types
- Added limit_mm_per_prompt parameter for image and video limits
Note: Consider adjusting default values - perhaps image limit > 1 or video=1
2025-06-10 23:36:47 +02:00
pandyamarut
99b952e55e
set default max_token size
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-02-24 17:22:44 -08:00
pandyamarut
56dc4ad075
update vllm
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-02-24 15:01:45 -08:00
sihamouda
0e3359a70e
add engine argument for multimodels input
2025-01-19 02:44:39 +01:00
pandyamarut
1420091588
update vll
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-09-27 19:23:50 -07:00
pandyamarut
814f50af38
fix oai completion api error
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-09-06 11:03:47 -07:00
alpayariyak
5bd6f3a75e
0.5.3, any vllm arg as env var, refactor and fixes, moving away from building separate image from vLLM fork
2024-07-25 12:41:48 -07:00
alpayariyak
f19ce12ab0
Fix Building Docker with model built-in #71
2024-06-07 15:29:46 -07:00
Alpay Ariyak and alpayariyak
874379a0c5
1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) ( #62 )
2024-05-09 00:42:08 -04:00
alpayariyak
d25b6f9628
Fix sampling params
2024-03-12 22:15:57 +00:00
alpayariyak
c8ee100d80
Small refactor
2024-03-06 17:08:57 +00:00
alpayariyak
aed0408f19
Documentation for 0.3.0, small fixes and changes
2024-02-21 19:28:59 -05:00
alpayariyak
e191149259
OpenAI Compatibility, Dynamic Batching, Refactor
2024-02-21 04:36:04 +00:00
alpayariyak
a94ef66f71
Dynamic Batch Size [needs refactor]
2024-02-06 03:47:27 +00:00
alpayariyak
fef8c81cb9
OpenAI Compatible worker, Refactor
2024-02-06 02:44:14 +00:00
alpayariyak
9fc8e1e54c
Non-streaming OpenAI Chat Completions
2024-01-25 23:24:18 -05:00
alpayariyak
4cebe66b36
0.2.0 Release
...
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak
584852f0f6
Docker-Protected HF Token, Refactor, Better Documentation
2024-01-18 18:20:51 -05:00
alpayariyak
fe6e618692
Refactor code and Black formatting style
2023-12-22 04:08:55 +00:00
alpayariyak
582e21f97f
Chat Template
2023-12-20 18:37:54 -05:00
alpayariyak
dad02e13d8
vLLM 0.2.6 Stable Worker
2023-12-20 23:06:57 +00:00
alpayariyak
8202d4cc63
Bug fix
2023-12-19 02:56:21 +00:00
alpayariyak
f324bef8a0
Bump to vLLM 0.2.6, fixes and improvements
2023-12-19 01:52:55 +00:00
alpayariyak
bfa3703d54
Concurrency modifier
2023-12-13 12:26:01 +00:00
alpayariyak
22de0b322d
Complete refactor, improved overall functionality
2023-12-13 10:33:03 +00:00
alpayariyak
5acbbf5528
Cleanup
2023-12-06 20:57:48 +00:00
alpayariyak
3a9393bafd
Batched Tokens, cleanup
2023-12-05 17:46:28 +00:00
alpayariyak
8d27261645
Logging change
2023-12-01 10:18:48 +00:00
alpayariyak
b242b0cedd
Refactor handler + add non-streaming, add utils.py
2023-12-01 10:12:56 +00:00