194 Commits
Author SHA1 Message Date
Alpay AriyakandGitHub 2941db0fb8 Update worker-vllm version 0.2.3 2024-02-09 23:12:29 -05:00
Alpay AriyakandGitHub bfeb60c54e Merge pull request #45 from willsamu/fix-tokenizer-input
fix: build error if no `TOKENIZER_NAME` provided
2024-02-09 22:51:18 -05:00
alpayariyak 7b3fd05542 Small refactor to tokenizer fix 2024-02-09 22:50:00 -05:00
Samuel Will b0e7b575f3 fix: default value for tokenizer 2024-02-09 10:24:58 +00:00
alpayariyak 4f5e0d37c4 Fix tokenizer's trust_remote_code parameter 2024-02-08 23:53:13 +00:00
Alpay AriyakandGitHub 2b5b8dfb61 Fix Model and Tokenizer download for bake-in option, add revision configuration for both. 2024-02-02 19:58:50 -05:00
alpayariyak 8de10468dd Working tokenizer and model download fix
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00
Casper f4d7c75504 Snapshot download only tokenizer/config related things 2024-01-31 22:16:44 +01:00
Casper 3adc9e3336 Remove unused import 2024-01-31 18:27:49 +01:00
Casper fd00a1ece3 Update to use snapshot_download 2024-01-31 18:24:57 +01:00
Casper 664dd35782 Download tokenizer upon build 2024-01-31 17:59:52 +01:00
alpayariyak 370698442c Update RunPod SDK version and Docker Tag 2024-01-31 00:54:58 -05:00
alpayariyak 3e2cd080a2 Fix Tensor Parallel 0.2.2 2024-01-31 05:12:58 +00:00
alpayariyak e7b340d73b Update vLLM base image 2024-01-31 03:37:55 +00:00
alpayariyak 46eee12819 Simplify Tensor Parallel 2024-01-31 03:28:02 +00:00
alpayariyak 12d6f0778e Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer 2024-01-31 03:13:21 +00:00
Alpay AriyakandGitHub fa5556434c Bug fix 2024-01-29 11:26:22 -05:00
alpayariyak 97726372c0 Update release tag in README.md 2024-01-25 23:37:07 -05:00
alpayariyak 9fc8e1e54c Non-streaming OpenAI Chat Completions 0.2.1 2024-01-25 23:24:18 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
0.2.0
2024-01-25 20:49:15 -05:00
Alpay Ariyakandalpayariyak 368c5f87fb Temporary Dockerfile Fix
Temporary Dockerfile fix
2024-01-25 20:48:26 -05:00
alpayariyak 15cc7cd36f Updated Documentation 2024-01-19 10:48:12 -05:00
alpayariyak ef3c303743 Added support for n parameter 2024-01-19 15:34:25 +00:00
alpayariyak 584852f0f6 Docker-Protected HF Token, Refactor, Better Documentation 2024-01-18 18:20:51 -05:00
Alpay AriyakandGitHub 65454c024a Update README.md (temporary) 2024-01-17 00:16:56 -05:00
Justin Merrell f023e2e097 Update README.md 2024-01-16 19:55:20 -05:00
Alpay AriyakandGitHub ed48093792 bug fix 0.1.0 2024-01-16 18:38:26 -05:00
alpayariyak a5dc8b53b7 Fix Concurrency, Add Max Model Length 2024-01-16 17:58:46 -05:00
Justin Merrell 6e7a1953e4 fix: update runpod package 2024-01-12 09:26:08 -05:00
Alpay AriyakandGitHub 23ff046586 Update README.md 2024-01-11 12:48:54 +07:00
alpayariyak 914caaa374 Merge branch 'refactor' 2024-01-05 23:44:59 +07:00
alpayariyak 618dd8e0c4 Temporarily remove examples 2024-01-05 16:44:12 +00:00
alpayariyak 3e08a69291 Documentation 2024-01-05 16:40:52 +00:00
alpayariyak fa358708bc Return finish status 2023-12-29 09:08:51 +00:00
alpayariyak e23ab549e4 Small fix 2023-12-29 08:44:54 +00:00
alpayariyak a14c5cd388 New Worker Stable 2023-12-29 08:21:09 +00:00
alpayariyak 2475cd7a66 Local development, disable log requests 2023-12-29 02:21:24 +00:00
alpayariyak cdda5edab3 Bump runpod version 2023-12-29 02:18:07 +00:00
alpayariyak ca8b02e392 Concurrency and Prompt Template fix 2023-12-29 00:42:40 +00:00
Alpay AriyakandGitHub 4a6bf9f0ee Concurrency hotfix 2023-12-28 19:03:46 -05:00
alpayariyak a69ab1875d vLLM job tracker, Refactor Concurrency Modifier, Serverless Config 2023-12-28 23:06:49 +00:00
alpayariyak f7ac60f802 Fix Concurrency 2023-12-28 22:56:44 +00:00
Alpay AriyakandGitHub a247a3afe1 Update README.md 2023-12-22 03:16:04 -05:00
alpayariyak fe6e618692 Refactor code and Black formatting style 2023-12-22 04:08:55 +00:00
alpayariyak 1fc64fef2b Additional Examples and Documentation for Chat Template and Messages list 2023-12-21 00:38:14 +00:00
Alpay AriyakandGitHub 38c48c2423 Chat Template Feature, Message List, Small Refactor
Chat Template Feature, Message List, Small Refactor
2023-12-20 19:24:29 -05:00
alpayariyak 3f413c1038 Documentation for Chat Templates, Messages format, Refactor Streaming, etc 2023-12-21 00:22:23 +00:00
alpayariyak ba893ed148 Refactor prompt handling in handler.py 2023-12-20 18:40:43 -05:00
alpayariyak 582e21f97f Chat Template 2023-12-20 18:37:54 -05:00