Tim Pietrusky and GitHub
66e1b1605b
Merge pull request #226 from runpod-workers/fix/cse-839-max-concurrency
...
Release / release (push) Waiting to run
fix: max concurrency = 30 instead of 300
v2.9.5
2025-10-22 22:54:21 +02:00
Tim Pietrusky and GitHub
60c8f257a8
Merge pull request #227 from runpod-workers/chore/vllm-0.11.0
...
chore: update vllm to 0.11.0
2025-10-22 22:53:51 +02:00
Tim Pietrusky
fae16e7ee1
chore: update vllm to 0.11.0
2025-10-22 13:40:53 -07:00
max4c
2becd35345
Revert "fix: added back the HF_TOKEN ( #219 )"
...
Release / release (push) Waiting to run
This reverts commit 33d88df6c0 .
v2.9.4
2025-09-23 12:38:24 -07:00
33d88df6c0
fix: added back the HF_TOKEN ( #219 )
...
Release / release (push) Waiting to run
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io >
v2.9.3
2025-09-23 19:25:12 +02:00
Tim Pietrusky
ecd562e112
fix: remove "access token" as this is handled by the platform
Release / release (push) Waiting to run
v2.9.2
2025-09-19 21:00:47 +02:00
5cffaab8e8
docs: how to use the reasoning parser ( #218 )
...
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io >
2025-09-17 07:49:37 +02:00
a0fe1dfdad
feat: better hub support & concise README for the main repo ( #215 )
...
Release / release (push) Waiting to run
* feat: moved config into docs; added banner; auto detect "messages" in input
* docs: moved config into docs
* chore: added .DS_Store
* chore: get the original stuff working again
* chore: remove all changes
* docs: reduced toc and added small config table
---------
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io >
v2.9.1
2025-09-01 16:47:48 +02:00
Tim Pietrusky
0e0d6df859
docs: updated example-tag for dev and release
2025-08-28 14:28:19 +02:00
Tim Pietrusky
d1718aec00
ci: removed "github release" step as that is not needed
2025-08-28 14:27:50 +02:00
5f0fc69d75
feat: prepare worker-vllm for the hub ( #214 )
...
Release / release (push) Waiting to run
* docs: remove outdated video; remove old info; added missing config for tools
* ci: use proper release for dev (pr only) and production (release only)
* ci(hub): added openai example; use smollm2 as base model
* docs: added conventions to be able to work with ai ide's
* chore: remove outdated stuff
* chore: update copyright to 2025
* ci: added github permissions
* feat: added gpuIds, gputCount and allowedCudaVersions; removed default value for LOAD_FORMAT to check which influence this has on the ui
---------
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io >
v2.9.0
2025-08-28 10:09:27 +02:00
Marut Pandya and GitHub
aef1187a30
Merge pull request #211 from runpod-workers/fix/allow-none-as-string
...
fix: allow "None" as value & parse the value of RAW_OPENAI_OUTPUT correctly
2025-08-21 11:35:54 -07:00
Tim Pietrusky
f7514dea4b
refactor: moved MODEL_NAME & HF_TOKEN out of advanced into the top section
2025-08-18 16:05:43 +02:00
Tim Pietrusky
121a3dd44b
fix: parse value for RAW_OPENAI_OUTPUT correctly
2025-08-13 12:20:44 +02:00
Tim Pietrusky
e2e111b942
fix: allow "None" as string for setting env variables (like quantization)
2025-08-13 10:27:00 +02:00
Marut Pandya and GitHub
7aa17463d3
Merge pull request #200 from runpod-workers/release/0.10.0
...
chore(release): 0.10.0
v2.8.0
2025-08-11 15:33:51 -07:00
Marut Pandya and GitHub
72a643dd0c
Merge pull request #207 from JhennerTigreros/main
...
Update requirements and engine creation to support new 0.10.0 vLLM version
2025-08-09 09:08:19 -07:00
Marut Pandya and GitHub
15f569f970
Merge pull request #208 from runpod-workers/revert-202-feat/proper-deployment
...
[Revert]"feat: added dev & release workflows; added conventions to support AI IDE"
2025-08-09 09:06:53 -07:00
Marut Pandya and GitHub
2f2bd4c749
Revert "feat: added dev & release workflows; added conventions to support AI IDE"
2025-08-09 09:01:52 -07:00
Jhenner Tigreros
fb0c030797
fix initialization on openaiservingmodels
2025-08-07 16:40:06 -05:00
Jhenner Tigreros
8b02a703b4
fix issues
2025-08-07 15:59:43 -05:00
Jhenner Tigreros
d8863139d6
add model to test
2025-08-07 15:28:02 -05:00
Jhenner Tigreros and GitHub
f5a063956e
Fix requirements.txt to support gpt-oss models
2025-08-07 15:13:37 -05:00
Marut Pandya and GitHub
18748fd73e
Merge pull request #202 from runpod-workers/feat/proper-deployment
...
feat: added dev & release workflows; added conventions to support AI IDE
2025-08-04 17:13:47 -07:00
Tim Pietrusky
0133c23be8
ci: added manual workflow trigger for releases
2025-08-04 09:58:23 +02:00
Tim Pietrusky
a129cff47d
docs: use "version" instead of actual version, so that people can check the releases
2025-07-31 12:03:17 +02:00
Tim Pietrusky
30f2c4630e
refactor: use correct version
2025-07-31 12:02:46 +02:00
Tim Pietrusky
b98636e432
feat: added "dev" and "release" workflows; removed "vllm-base-image" as it's not needed
2025-07-28 16:40:20 +02:00
pandyamarut
185205c750
chore(release): 0.10.0
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-07-25 10:18:06 -07:00
pandyamarut
b948e530a1
chore(release): 0.10.0
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-07-25 10:17:50 -07:00
Marut Pandya and GitHub
4f22d6f107
Merge pull request #193 from runpod-workers/release/0.9.1
...
chore(release): v0.9.1
2025-06-26 13:23:33 -07:00
pandyamarut
8839689132
chore(release): v0.9.1
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-06-26 11:06:23 -07:00
Marut Pandya and GitHub
5ddc2326cd
Merge pull request #191 from runpod-workers/feat/0.9.1
...
feat: update to 0.9.1 & added CONFIG_FORMAT to run magistral
2025-06-26 10:17:54 -07:00
Tim Pietrusky
46a3300cbb
chore: reverted changeds to only focus on vllm update
2025-06-20 14:56:39 +02:00
Tim Pietrusky
1e9a731380
Fix Mistral tokenizer initialization: let vLLM handle tokenizer for mistral models
2025-06-14 15:31:56 +02:00
Tim Pietrusky
c47e649a24
Add CONFIG_FORMAT environment variable support
2025-06-14 15:12:46 +02:00
Tim Pietrusky
7192bcaeef
Trigger build automatically on feat/0.9.1 branch
2025-06-14 13:49:22 +02:00
Tim Pietrusky
8665ffb78d
Revert workflow back to original configuration
2025-06-14 13:47:16 +02:00
Tim Pietrusky
57431b30ad
Fix workflow: use standard GitHub runners and actions
2025-06-14 13:42:46 +02:00
Tim Pietrusky
437a84c77a
ci: added workflow to build the image
2025-06-14 13:31:36 +02:00
Tim Pietrusky
a4062fc488
feat: update to 0.9.1
2025-06-14 13:31:27 +02:00
Marut Pandya and GitHub
9631407c1d
Merge pull request #189 from runpod-workers/hf-mm
...
add multi modal env var
2025-06-10 17:03:18 -07:00
pandyamarut
1a93932ab2
add multi modal env var
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-06-10 17:02:42 -07:00
Marut Pandya and GitHub
4c6c88c3b9
Merge pull request #188 from WorstDev01/add-mm-limit-parameter
...
Add multimodal limit parameter support
2025-06-10 14:59:21 -07:00
WorstDev01
a58a76783d
Add multimodal limit parameter support
...
- Updated convert_limit_mm_per_prompt to handle multiple types
- Added limit_mm_per_prompt parameter for image and video limits
Note: Consider adjusting default values - perhaps image limit > 1 or video=1
2025-06-10 23:36:47 +02:00
Marut Pandya and GitHub
26919c8849
Merge pull request #185 from runpod-workers/up-0.9.0
...
Version upgrade
v2.6.0
2025-06-05 12:00:05 -07:00
pandyamarut
70cd1c8113
update vllm version 0.9.0
...
Signed-off-by: pandyamarut <pandyamarut@gmail.com >
2025-06-05 11:59:15 -07:00
Marut Pandya and GitHub
deeff579f8
Merge pull request #183 from SorenDreano/fix/model_name_in_local_args
...
remove requirements for MODEL_NAME in local_model_args.json
2025-06-04 12:45:00 -07:00
Soren Dreano
11f96a09d7
remove requirements for MODEL_NAME in local_model_args.json
...
We want to use the same local_model_args.json for multiple models
which have different names. It would be very convenient to only
have a single local_args file and not have to create it every time
A warning should be enough for users
2025-05-23 17:59:19 +02:00
Marut Pandya and GitHub
23e8ecf85b
Merge pull request #169 from RedHitMark/main
...
fix lora and multi-lora
2025-05-14 16:25:26 -07:00