Commit Graph
67 Commits
Author SHA1 Message Date
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00
alpayariyak 370698442c Update RunPod SDK version and Docker Tag 2024-01-31 00:54:58 -05:00
alpayariyak 3e2cd080a2 Fix Tensor Parallel 2024-01-31 05:12:58 +00:00
alpayariyak e7b340d73b Update vLLM base image 2024-01-31 03:37:55 +00:00
alpayariyak 46eee12819 Simplify Tensor Parallel 2024-01-31 03:28:02 +00:00
alpayariyak 12d6f0778e Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer 2024-01-31 03:13:21 +00:00
alpayariyak 97726372c0 Update release tag in README.md 2024-01-25 23:37:07 -05:00
alpayariyak 9fc8e1e54c Non-streaming OpenAI Chat Completions 2024-01-25 23:24:18 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak 15cc7cd36f Updated Documentation 2024-01-19 10:48:12 -05:00
alpayariyak ef3c303743 Added support for n parameter 2024-01-19 15:34:25 +00:00
alpayariyak 584852f0f6 Docker-Protected HF Token, Refactor, Better Documentation 2024-01-18 18:20:51 -05:00
alpayariyak a5dc8b53b7 Fix Concurrency, Add Max Model Length 2024-01-16 17:58:46 -05:00
alpayariyak 914caaa374 Merge branch 'refactor' 2024-01-05 23:44:59 +07:00
alpayariyak 618dd8e0c4 Temporarily remove examples 2024-01-05 16:44:12 +00:00
alpayariyak 3e08a69291 Documentation 2024-01-05 16:40:52 +00:00
alpayariyak fa358708bc Return finish status 2023-12-29 09:08:51 +00:00
alpayariyak e23ab549e4 Small fix 2023-12-29 08:44:54 +00:00
alpayariyak a14c5cd388 New Worker Stable 2023-12-29 08:21:09 +00:00
alpayariyak 2475cd7a66 Local development, disable log requests 2023-12-29 02:21:24 +00:00
alpayariyak cdda5edab3 Bump runpod version 2023-12-29 02:18:07 +00:00
alpayariyak ca8b02e392 Concurrency and Prompt Template fix 2023-12-29 00:42:40 +00:00
alpayariyak a69ab1875d vLLM job tracker, Refactor Concurrency Modifier, Serverless Config 2023-12-28 23:06:49 +00:00
alpayariyak f7ac60f802 Fix Concurrency 2023-12-28 22:56:44 +00:00
alpayariyak fe6e618692 Refactor code and Black formatting style 2023-12-22 04:08:55 +00:00
alpayariyak 1fc64fef2b Additional Examples and Documentation for Chat Template and Messages list 2023-12-21 00:38:14 +00:00
alpayariyak 3f413c1038 Documentation for Chat Templates, Messages format, Refactor Streaming, etc 2023-12-21 00:22:23 +00:00
alpayariyak ba893ed148 Refactor prompt handling in handler.py 2023-12-20 18:40:43 -05:00
alpayariyak 582e21f97f Chat Template 2023-12-20 18:37:54 -05:00
alpayariyak dad02e13d8 vLLM 0.2.6 Stable Worker 2023-12-20 23:06:57 +00:00
alpayariyak e4e3ce4a6e Small fix 2023-12-20 08:45:48 +00:00
alpayariyak c59038e902 Fixed CUDA 11.8 workers 2023-12-20 07:50:14 +00:00
alpayariyak 8a11a7e4ed Add version 2023-12-19 02:57:06 +00:00
alpayariyak 8202d4cc63 Bug fix 2023-12-19 02:56:21 +00:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00
alpayariyak d39844ac00 Small fixes 2023-12-13 13:50:36 +00:00
alpayariyak bfa3703d54 Concurrency modifier 2023-12-13 12:26:01 +00:00
alpayariyak 22de0b322d Complete refactor, improved overall functionality 2023-12-13 10:33:03 +00:00
alpayariyak ef89868570 Small changes 2023-12-12 20:42:27 +00:00
alpayariyak 91c47790a6 Merge branch 'merge-ready' of https://github.com/runpod-workers/worker-vllm into merge-ready 2023-12-12 20:40:34 +00:00
alpayariyak e81007d79b bug fix 2023-12-09 11:10:29 -08:00
alpayariyak b51827388c bug fix 2023-12-09 19:06:41 +00:00
alpayariyak 8578581c6a Small changes 2023-12-09 18:50:00 +00:00
alpayariyak f1afbcbe4f CPU Limit 2023-12-09 18:47:45 +00:00
alpayariyak 5e38a1a2fa Multi-GPU Fix 2023-12-09 10:36:26 +00:00
alpayariyak 4052a17af1 Merge branch 'merge-ready' of https://github.com/runpod-workers/worker-vllm into merge-ready 2023-12-08 10:29:36 +00:00
alpayariyak 8156b1d8e0 Final changes 2023-12-08 10:29:24 +00:00
alpayariyak 5acbbf5528 Cleanup 2023-12-06 20:57:48 +00:00
alpayariyak 8dbd8f15dd Token counting, cuda-related changes 2023-12-06 19:02:12 +00:00
alpayariyak b705544bd4 model packing, cuda version selection on build 2023-12-06 18:52:21 +00:00
alpayariyak 48c520b240 Version Bump + small changes 2023-12-06 16:48:12 +00:00
alpayariyak 66c23d45b0 Model Downloading, no GPU 2023-12-06 01:43:09 +00:00
alpayariyak b4c61e61c4 Packing model into worker for internal use 2023-12-05 20:26:56 -05:00
alpayariyak 3a9393bafd Batched Tokens, cleanup 2023-12-05 17:46:28 +00:00
alpayariyak 8d27261645 Logging change 2023-12-01 10:18:48 +00:00
alpayariyak b242b0cedd Refactor handler + add non-streaming, add utils.py 2023-12-01 10:12:56 +00:00
alpayariyak daf2f4d07b Make handler stream aggregate_text 2023-12-01 04:21:03 +00:00
alpayariyak 86916bf628 Update handler.py 2023-11-29 23:37:17 -05:00
alpayariyak 820a21f32c Download logic + other changes 2023-11-28 19:15:08 -05:00
alpayariyak fdddeacf3c Update handler.py 2023-11-28 18:49:15 -05:00
alpayariyak 7cadff29f5 edge case 2023-11-28 18:49:04 -05:00
alpayariyak 97480ff0ae handler changes 2023-11-28 23:39:48 +00:00
alpayariyak 7f73f41c9f Merge remote-tracking branch 'origin/new-concurrency' into revamp2 2023-11-26 18:23:50 -05:00
Alpay AriyakandUbuntu 96612413ce worker changes 2023-11-22 00:39:52 +00:00
alpayariyak c34d79ea49 Entrypoint for testing 2023-11-20 18:56:00 -05:00
alpayariyak 8f17764f8a Update Dockerfile 2023-11-15 20:55:17 -05:00
alpayariyak e7c8de4678 Changes to worker 2023-11-15 20:53:04 -05:00