velaraptor-runpod
7ec10b98cd
Update utils.py
2026-02-12 15:28:31 -06:00
velaraptor-runpod
340bc0b3c6
fix: served model name
2026-02-10 21:42:58 -06:00
velaraptor-runpod
e1e9ef74ad
add changes from pr
2026-02-06 18:10:09 -06:00
velaraptor-runpod
8eb55b90c1
add changes for v0.15.0
2026-02-05 17:24:16 -06:00
chrisvela and GitHub
3851d53f93
add ENABLE_EXPERT_PARALLEL engine arg for MoE models ( #239 )
...
Release / release (push) Waiting to run
* enable expert parallel arg for moe models
* add ENABLE_EXPERT_PARALLEL to hub config
2025-11-17 19:25:19 +01:00
Tim Pietrusky and GitHub
66e1b1605b
Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
...
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
2025-10-22 22:54:21 +02:00
Tim Pietrusky
121a3dd44b
fix: parse value for RAW_OPENAI_OUTPUT correctly
2025-08-13 12:20:44 +02:00
Tim Pietrusky
e2e111b942
fix: allow "None" as string for setting env variables (like quantization)
2025-08-13 10:27:00 +02:00
Jhenner Tigreros
fb0c030797
fix initialization on openaiservingmodels
2025-08-07 16:40:06 -05:00
Tim Pietrusky
1e9a731380
Fix Mistral tokenizer initialization: let vLLM handle tokenizer for mistral models
2025-06-14 15:31:56 +02:00
Tim Pietrusky
c47e649a24
Add CONFIG_FORMAT environment variable support
2025-06-14 15:12:46 +02:00
pandyamarut
1a93932ab2
add multi modal env var
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-06-10 17:02:42 -07:00
WorstDev01
a58a76783d
Add multimodal limit parameter support
...
- Updated convert_limit_mm_per_prompt to handle multiple types
- Added limit_mm_per_prompt parameter for image and video limits
Note: Consider adjusting default values - perhaps image limit > 1 or video=1
2025-06-10 23:36:47 +02:00
pandyamarut
70cd1c8113
update vllm version 0.9.0
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-06-05 11:59:15 -07:00
Soren Dreano
11f96a09d7
remove requirements for MODEL_NAME in local_model_args.json
...
We want to use the same local_model_args.json for multiple models
which have different names. It would be very convenient to only
have a single local_args file and not have to create it every time
A warning should be enough for users
2025-05-23 17:59:19 +02:00
Marut Pandya and GitHub
23e8ecf85b
Merge pull request #169 from RedHitMark/main
...
fix lora and multi-lora
2025-05-14 16:25:26 -07:00
Marut Pandya and GitHub
a9786a2481
Revert "fix: added back limit_mm_per_prompt to engine args"
2025-05-07 12:29:33 -07:00
Marut Pandya and GitHub
4e474c41c8
Merge pull request #177 from aleksandar-babic/main
...
fix: added back limit_mm_per_prompt to engine args
2025-05-03 21:01:58 -07:00
Marut Pandya and GitHub
6b64bb93bd
Merge pull request #147 from mohamednaji7/BitsAndBytes
...
completing the "bitsandbytes" option - based on https://docs.vllm.ai/en/stable/quantization/bnb.html
2025-04-21 11:39:53 -07:00
Aleksandar Babic
cfc258674b
chore: added trailing comma to the final arg
2025-04-21 10:06:08 -04:00
Aleksandar Babic
8beafed06b
fix: added back limit_mm_per_prompt to engine args
2025-04-21 10:02:46 -04:00
RedHitMark
084d000324
fix lora and multi-lora
2025-03-14 21:45:09 +01:00
pandyamarut
99b952e55e
set default max_token size
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-02-24 17:22:44 -08:00
pandyamarut
56dc4ad075
update vllm
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-02-24 15:01:45 -08:00
Marut Pandya and GitHub
2b1d618287
Revert "Enabling model caching."
2025-02-18 10:45:47 -08:00
Marut Pandya and GitHub
6fc770415d
Merge pull request #157 from runpod-workers/m-c
...
Enabling model caching.
2025-02-06 11:52:19 -08:00
pandyamarut
30dd7c1eb5
update engine.py
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-01-28 21:40:00 -08:00
pandyamarut
dc8c88027a
update serving classes
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-01-28 20:58:07 -08:00
pandyamarut
c703254f71
dynamic lora loading
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-01-28 19:00:57 -08:00
mohamednaji7
131c17569f
correct access to "args" dictionary
2025-01-21 22:33:28 +02:00
mohamednaji7
a27f72a33a
inforce args.quantization for bnb load_froamt
2025-01-21 14:01:32 +02:00
sihamouda
0e3359a70e
add engine argument for multimodels input
2025-01-19 02:44:39 +01:00
Marut Pandya and GitHub
3f0a20d28e
Merge pull request #141 from runpod-workers/main
...
Rebase
2025-01-02 20:49:48 -08:00
pandyamarut
66ea8b1110
update engine
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-12-30 17:13:03 -08:00
kldzj
b7787d8ca8
fix: remove unused import
2024-12-06 16:29:06 +01:00
kldzj
b1fca5d257
fix: move tool flags away from engine args
2024-12-06 16:07:23 +01:00
kldzj
3e86d16892
fix: openai tool calling
2024-12-06 13:08:48 +01:00
Nikolai Kolodziej
8df7f41f1d
fix: set empty tool_call_parser to None
2024-11-24 06:55:43 +01:00
Nikolai Kolodziej
4d7b8c03c0
feat: tool calling flags
2024-11-24 06:51:41 +01:00
pandyamarut
6c6bf50379
update env
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-11-22 15:14:57 -08:00
pandyamarut
27a2ee5754
add model cache
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-11-22 13:11:12 -08:00
Marut Pandya and GitHub
6e8696c12a
Merge pull request #121 from sven-knoblauch/lora-modules
...
add changes for lora adapter support and /v1/models endpoint
2024-10-31 15:07:02 -04:00
pandyamarut
c03ecc42fe
update vllm
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-10-15 16:49:41 -07:00
Sven Knoblauch
677a01e8f3
update code for case of no lora adapter
2024-10-09 14:51:33 +02:00
Sven Knoblauch
5cd12ba331
add changes for lora adapter support and /v1/models endpoint
2024-10-09 11:01:12 +02:00
pandyamarut
1420091588
update vll
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-09-27 19:23:50 -07:00
pandyamarut
814f50af38
fix oai completion api error
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-09-06 11:03:47 -07:00
pandyamarut
967eaba573
change to float
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-08-09 14:41:07 -07:00
pandyamarut
9cb9336cf5
update vllm version 0.5.4
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-08-09 12:01:42 -07:00
pandyamarut
f3534a4ea7
fix openai compat
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2024-07-27 15:49:45 -07:00