Commit Graph
154 Commits
Author SHA1 Message Date
pandyamarut dc8c88027a update serving classes
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-01-28 20:58:07 -08:00
pandyamarut c703254f71 dynamic lora loading
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-01-28 19:00:57 -08:00
sihamouda 0e3359a70e add engine argument for multimodels input 2025-01-19 02:44:39 +01:00
pandyamarut 66ea8b1110 update engine
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-12-30 17:13:03 -08:00
kldzj b7787d8ca8 fix: remove unused import 2024-12-06 16:29:06 +01:00
kldzj b1fca5d257 fix: move tool flags away from engine args 2024-12-06 16:07:23 +01:00
kldzj 3e86d16892 fix: openai tool calling 2024-12-06 13:08:48 +01:00
Nikolai Kolodziej 8df7f41f1d fix: set empty tool_call_parser to None 2024-11-24 06:55:43 +01:00
Nikolai Kolodziej 4d7b8c03c0 feat: tool calling flags 2024-11-24 06:51:41 +01:00
Marut PandyaandGitHub 6e8696c12a Merge pull request #121 from sven-knoblauch/lora-modules
add changes for lora adapter support and /v1/models endpoint
2024-10-31 15:07:02 -04:00
pandyamarut c03ecc42fe update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-10-15 16:49:41 -07:00
Sven Knoblauch 677a01e8f3 update code for case of no lora adapter 2024-10-09 14:51:33 +02:00
Sven Knoblauch 5cd12ba331 add changes for lora adapter support and /v1/models endpoint 2024-10-09 11:01:12 +02:00
pandyamarut 1420091588 update vll
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-27 19:23:50 -07:00
pandyamarut 814f50af38 fix oai completion api error
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-09-06 11:03:47 -07:00
pandyamarut 967eaba573 change to float
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-08-09 14:41:07 -07:00
pandyamarut 9cb9336cf5 update vllm version 0.5.4
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-08-09 12:01:42 -07:00
pandyamarut f3534a4ea7 fix openai compat
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-07-27 15:49:45 -07:00
pandyamarut 0f8657e58d update env default args
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-07-26 16:52:51 -07:00
pandyamarut bd96b5e0de update v0.5.3.post1
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-07-25 21:29:45 -07:00
alpayariyak 5bd6f3a75e 0.5.3, any vllm arg as env var, refactor and fixes, moving away from building separate image from vLLM fork 2024-07-25 12:41:48 -07:00
alpayariyak a08d83f600 Allow any vLLM engine args as env vars, refactor 2024-07-02 19:44:01 +00:00
alpayariyak 0e1e38326a Fix deprecated max_context_len_to_capture engine argument 2024-06-13 17:48:05 +00:00
alpayariyak bad5ddd892 Fix vLLM 0.4.1 bug for OpenAI /models route 2024-06-10 20:38:49 +00:00
alpayariyak f19ce12ab0 Fix Building Docker with model built-in #71 2024-06-07 15:29:46 -07:00
alpayariyak 1bb6f84541 Deprecated kv cache dtype warning 2024-05-10 16:29:20 +00:00
Alpay Ariyakandalpayariyak 874379a0c5 1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) (#62) 2024-05-09 00:42:08 -04:00
alpayariyak d25b6f9628 Fix sampling params 2024-03-12 22:15:57 +00:00
alpayariyak c8ee100d80 Small refactor 2024-03-06 17:08:57 +00:00
alpayariyak db7167d57f 0.3.3 2024-03-05 19:14:35 +00:00
Alpay Ariyakandalpayariyak d91ccb866f 0.3.1: bug fixes 2024-02-29 02:55:44 -05:00
alpayariyak 985bbf1cb5 Fix Multi-GPU, tokenizer trust remote code 2024-02-24 03:18:25 +00:00
alpayariyak 708f68d7f8 Final Bug fixes, configurable oai response role, served model name override, documentation 2024-02-23 04:14:03 +00:00
alpayariyak b42d45ce0f Bug fixes, refactors 2024-02-23 03:46:47 +00:00
alpayariyak 9129d0a252 New ENV Vars 2024-02-22 23:58:20 +00:00
alpayariyak 6bcd9d7c67 Preparing for 0.3.0 2024-02-22 18:04:34 -05:00
alpayariyak aed0408f19 Documentation for 0.3.0, small fixes and changes 2024-02-21 19:28:59 -05:00
alpayariyak e191149259 OpenAI Compatibility, Dynamic Batching, Refactor 2024-02-21 04:36:04 +00:00
alpayariyak a94ef66f71 Dynamic Batch Size [needs refactor] 2024-02-06 03:47:27 +00:00
alpayariyak 45081e4037 Add __init__.py 2024-02-06 02:44:35 +00:00
alpayariyak fef8c81cb9 OpenAI Compatible worker, Refactor 2024-02-06 02:44:14 +00:00
alpayariyak 15b06bb687 Merge branch 'main' into openai-sse-output 2024-02-02 20:03:29 -05:00
alpayariyak 8de10468dd Working tokenizer and model download fix
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak b7051d37ca Move test_openai_stream.py 2024-02-02 22:12:26 +00:00
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00
alpayariyak afa33a2875 Handle errors 2024-02-01 03:01:51 +00:00
alpayariyak 068303ce8f OpenAI Proxy Server for EndPoints and more examples 2024-02-01 02:32:32 +00:00
alpayariyak 3bbcf0021b Fix: Yield if tokens left in batch
Temp: default model for testing image
2024-02-01 00:52:39 +00:00
alpayariyak dab8bad906 OpenAI Chat Completions Stream 2024-01-31 23:27:31 +00:00
Casper f4d7c75504 Snapshot download only tokenizer/config related things 2024-01-31 22:16:44 +01:00