Commit Graph
445 Commits
Author SHA1 Message Date
Tim Pietrusky 0e0d6df859 docs: updated example-tag for dev and release 2025-08-28 14:28:19 +02:00
Tim Pietrusky d1718aec00 ci: removed "github release" step as that is not needed 2025-08-28 14:27:50 +02:00
5f0fc69d75 feat: prepare worker-vllm for the hub (#214)
Release / release (push) Waiting to run
* docs: remove outdated video; remove old info; added missing config for tools

* ci: use proper release for dev (pr only) and production (release only)

* ci(hub): added openai example; use smollm2 as base model

* docs: added conventions to be able to work with ai ide's

* chore: remove outdated stuff

* chore: update copyright to 2025

* ci: added github permissions

* feat: added gpuIds, gputCount and allowedCudaVersions; removed default value for LOAD_FORMAT to check which influence this has on the ui

---------

Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
v2.9.0
2025-08-28 10:09:27 +02:00
Marut PandyaandGitHub aef1187a30 Merge pull request #211 from runpod-workers/fix/allow-none-as-string
fix: allow "None" as value & parse the value of RAW_OPENAI_OUTPUT correctly
2025-08-21 11:35:54 -07:00
Tim Pietrusky f7514dea4b refactor: moved MODEL_NAME & HF_TOKEN out of advanced into the top section 2025-08-18 16:05:43 +02:00
Tim Pietrusky 121a3dd44b fix: parse value for RAW_OPENAI_OUTPUT correctly 2025-08-13 12:20:44 +02:00
Tim Pietrusky e2e111b942 fix: allow "None" as string for setting env variables (like quantization) 2025-08-13 10:27:00 +02:00
Marut PandyaandGitHub 7aa17463d3 Merge pull request #200 from runpod-workers/release/0.10.0
chore(release): 0.10.0
v2.8.0
2025-08-11 15:33:51 -07:00
Marut PandyaandGitHub 72a643dd0c Merge pull request #207 from JhennerTigreros/main
Update requirements and engine creation to support new 0.10.0 vLLM version
2025-08-09 09:08:19 -07:00
Marut PandyaandGitHub 15f569f970 Merge pull request #208 from runpod-workers/revert-202-feat/proper-deployment
[Revert]"feat: added dev & release workflows; added conventions to support AI IDE"
2025-08-09 09:06:53 -07:00
Marut PandyaandGitHub 2f2bd4c749 Revert "feat: added dev & release workflows; added conventions to support AI IDE" 2025-08-09 09:01:52 -07:00
Jhenner Tigreros fb0c030797 fix initialization on openaiservingmodels 2025-08-07 16:40:06 -05:00
Jhenner Tigreros 8b02a703b4 fix issues 2025-08-07 15:59:43 -05:00
Jhenner Tigreros d8863139d6 add model to test 2025-08-07 15:28:02 -05:00
Jhenner TigrerosandGitHub f5a063956e Fix requirements.txt to support gpt-oss models 2025-08-07 15:13:37 -05:00
Marut PandyaandGitHub 18748fd73e Merge pull request #202 from runpod-workers/feat/proper-deployment
feat: added dev & release workflows; added conventions to support AI IDE
2025-08-04 17:13:47 -07:00
Tim Pietrusky 0133c23be8 ci: added manual workflow trigger for releases 2025-08-04 09:58:23 +02:00
Tim Pietrusky a129cff47d docs: use "version" instead of actual version, so that people can check the releases 2025-07-31 12:03:17 +02:00
Tim Pietrusky 30f2c4630e refactor: use correct version 2025-07-31 12:02:46 +02:00
Tim Pietrusky b98636e432 feat: added "dev" and "release" workflows; removed "vllm-base-image" as it's not needed 2025-07-28 16:40:20 +02:00
pandyamarut 185205c750 chore(release): 0.10.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-07-25 10:18:06 -07:00
pandyamarut b948e530a1 chore(release): 0.10.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-07-25 10:17:50 -07:00
Marut PandyaandGitHub 4f22d6f107 Merge pull request #193 from runpod-workers/release/0.9.1
chore(release): v0.9.1
2025-06-26 13:23:33 -07:00
pandyamarut 8839689132 chore(release): v0.9.1
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-06-26 11:06:23 -07:00
Marut PandyaandGitHub 5ddc2326cd Merge pull request #191 from runpod-workers/feat/0.9.1
feat: update to 0.9.1 & added CONFIG_FORMAT to run magistral
2025-06-26 10:17:54 -07:00
Tim Pietrusky 46a3300cbb chore: reverted changeds to only focus on vllm update 2025-06-20 14:56:39 +02:00
Tim Pietrusky 1e9a731380 Fix Mistral tokenizer initialization: let vLLM handle tokenizer for mistral models 2025-06-14 15:31:56 +02:00
Tim Pietrusky c47e649a24 Add CONFIG_FORMAT environment variable support 2025-06-14 15:12:46 +02:00
Tim Pietrusky 7192bcaeef Trigger build automatically on feat/0.9.1 branch 2025-06-14 13:49:22 +02:00
Tim Pietrusky 8665ffb78d Revert workflow back to original configuration 2025-06-14 13:47:16 +02:00
Tim Pietrusky 57431b30ad Fix workflow: use standard GitHub runners and actions 2025-06-14 13:42:46 +02:00
Tim Pietrusky 437a84c77a ci: added workflow to build the image 2025-06-14 13:31:36 +02:00
Tim Pietrusky a4062fc488 feat: update to 0.9.1 2025-06-14 13:31:27 +02:00
Marut PandyaandGitHub 9631407c1d Merge pull request #189 from runpod-workers/hf-mm
add multi modal env var
2025-06-10 17:03:18 -07:00
pandyamarut 1a93932ab2 add multi modal env var
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-06-10 17:02:42 -07:00
Marut PandyaandGitHub 4c6c88c3b9 Merge pull request #188 from WorstDev01/add-mm-limit-parameter
Add multimodal limit parameter support
2025-06-10 14:59:21 -07:00
WorstDev01 a58a76783d Add multimodal limit parameter support
- Updated convert_limit_mm_per_prompt to handle multiple types
- Added limit_mm_per_prompt parameter for image and video limits

Note: Consider adjusting default values - perhaps image limit > 1 or video=1
2025-06-10 23:36:47 +02:00
Marut PandyaandGitHub 26919c8849 Merge pull request #185 from runpod-workers/up-0.9.0
Version upgrade
v2.6.0
2025-06-05 12:00:05 -07:00
pandyamarut 70cd1c8113 update vllm version 0.9.0
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-06-05 11:59:15 -07:00
Marut PandyaandGitHub deeff579f8 Merge pull request #183 from SorenDreano/fix/model_name_in_local_args
remove requirements for MODEL_NAME in local_model_args.json
2025-06-04 12:45:00 -07:00
Soren Dreano 11f96a09d7 remove requirements for MODEL_NAME in local_model_args.json
We want to use the same local_model_args.json for multiple models
which have different names. It would be very convenient to only
have a single local_args file and not have to create it every time

A warning should be enough for users
2025-05-23 17:59:19 +02:00
Marut PandyaandGitHub 23e8ecf85b Merge pull request #169 from RedHitMark/main
fix lora and multi-lora
2025-05-14 16:25:26 -07:00
Marut PandyaandGitHub 6db2c44d3b Update Dockerfile 2025-05-09 09:27:54 -07:00
Marut PandyaandGitHub 9b7ca4d0b0 Update tests.json v2.5.1 2025-05-08 17:12:59 -07:00
Marut PandyaandGitHub 2a4eaf0356 Merge pull request #182 from runpod-workers/up-0.8.5
update vllm
v2.5.0
2025-05-08 11:40:23 -07:00
pandyamarut ba19cc97bf update vllm
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2025-05-08 11:34:58 -07:00
Marut PandyaandGitHub 6075f2c590 Merge pull request #181 from runpod-workers/revert-177-main
Revert "fix: added back limit_mm_per_prompt to engine args"
2025-05-07 12:31:15 -07:00
Marut PandyaandGitHub a9786a2481 Revert "fix: added back limit_mm_per_prompt to engine args" 2025-05-07 12:29:33 -07:00
Marut PandyaandGitHub 0b6bc7a2be Update tests.json 2025-05-07 11:14:50 -07:00
Marut PandyaandGitHub d0ab58ee17 Merge pull request #180 from muhsinking/patch-1
Update README.md table to fix table of contents links
2025-05-03 21:02:37 -07:00