Commit Graph
73 Commits
Author SHA1 Message Date
pandyamarut 9cb9336cf5 update vllm version 0.5.4
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-08-09 12:01:42 -07:00
pandyamarut bd96b5e0de update v0.5.3.post1
Signed-off-by: pandyamarut <pandyamarut@gmail.com>
2024-07-25 21:29:45 -07:00
alpayariyak 5bd6f3a75e 0.5.3, any vllm arg as env var, refactor and fixes, moving away from building separate image from vLLM fork 2024-07-25 12:41:48 -07:00
alpayariyak c8458fef2b preparation for 1.0.0 release 2024-06-12 12:28:03 -07:00
alpayariyak ec7ea0b760 Update default base image version to fix github actions build 2024-05-09 20:18:34 -04:00
Alpay Ariyakandalpayariyak 874379a0c5 1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) (#62) 2024-05-09 00:42:08 -04:00
alpayariyak db7167d57f 0.3.3 2024-03-05 19:14:35 +00:00
Alpay Ariyakandalpayariyak d91ccb866f 0.3.1: bug fixes 2024-02-29 02:55:44 -05:00
alpayariyak 6bcd9d7c67 Preparing for 0.3.0 2024-02-22 18:04:34 -05:00
alpayariyak 8de10468dd Working tokenizer and model download fix
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00
alpayariyak e7b340d73b Update vLLM base image 2024-01-31 03:37:55 +00:00
alpayariyak 12d6f0778e Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer 2024-01-31 03:13:21 +00:00
Alpay AriyakandGitHub fa5556434c Bug fix 2024-01-29 11:26:22 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
Alpay Ariyakandalpayariyak 368c5f87fb Temporary Dockerfile Fix
Temporary Dockerfile fix
2024-01-25 20:48:26 -05:00
alpayariyak 584852f0f6 Docker-Protected HF Token, Refactor, Better Documentation 2024-01-18 18:20:51 -05:00
alpayariyak c59038e902 Fixed CUDA 11.8 workers 2023-12-20 07:50:14 +00:00
Alpay AriyakandGitHub c321d828ab Update Dockerfile 2023-12-19 00:58:45 -05:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00
alpayariyak 22de0b322d Complete refactor, improved overall functionality 2023-12-13 10:33:03 +00:00
alpayariyak ef89868570 Small changes 2023-12-12 20:42:27 +00:00
alpayariyak b51827388c bug fix 2023-12-09 19:06:41 +00:00
alpayariyak 8578581c6a Small changes 2023-12-09 18:50:00 +00:00
alpayariyak f1afbcbe4f CPU Limit 2023-12-09 18:47:45 +00:00
alpayariyak 5e38a1a2fa Multi-GPU Fix 2023-12-09 10:36:26 +00:00
Justin Merrell bcf37e7bd2 Update Dockerfile 2023-12-08 20:42:15 -05:00
Justin Merrell deb6f67a25 Update Dockerfile 2023-12-08 20:42:05 -05:00
Justin Merrell 11e6495304 Update Dockerfile 2023-12-08 20:28:44 -05:00
Justin Merrell 35a74475f2 Update Dockerfile 2023-12-08 20:23:23 -05:00
alpayariyak 8dbd8f15dd Token counting, cuda-related changes 2023-12-06 19:02:12 +00:00
alpayariyak b705544bd4 model packing, cuda version selection on build 2023-12-06 18:52:21 +00:00
alpayariyak 48c520b240 Version Bump + small changes 2023-12-06 16:48:12 +00:00
alpayariyak 66c23d45b0 Model Downloading, no GPU 2023-12-06 01:43:09 +00:00
alpayariyak b4c61e61c4 Packing model into worker for internal use 2023-12-05 20:26:56 -05:00
Justin Merrell bd3ede60f3 Update Dockerfile 2023-11-29 22:32:30 -05:00
Justin Merrell 4131b7a16d Update Dockerfile 2023-11-29 22:32:13 -05:00
Justin Merrell 9b9c1cf503 cleanup 2023-11-29 21:32:40 -05:00
alpayariyak 820a21f32c Download logic + other changes 2023-11-28 19:15:08 -05:00
alpayariyak 97480ff0ae handler changes 2023-11-28 23:39:48 +00:00
Justin Merrell 98af834124 fix: added hf_transfer requirement 2023-11-21 20:50:12 -05:00
Justin Merrell d06fb9ea01 Update Dockerfile 2023-11-21 20:20:15 -05:00
Justin Merrell 5ca8e56bbe fix: clean up extra 2023-11-21 19:18:08 -05:00
alpayariyak c34d79ea49 Entrypoint for testing 2023-11-20 18:56:00 -05:00
alpayariyak 8f17764f8a Update Dockerfile 2023-11-15 20:55:17 -05:00
alpayariyak e7c8de4678 Changes to worker 2023-11-15 20:53:04 -05:00
Samuel Willandalpayariyak 4f792062aa vllm 0.2.1.post1 - speed boost, quantization, mistral support 2023-11-14 19:12:58 -05:00
Jorg Doku 06183813cf take advantage of layering 2023-08-31 18:30:59 -05:00
Jorg Doku 54d229cb16 upate metrics 2023-08-31 18:26:37 -05:00
Jorg Doku 84d6047cbf update 2023-08-31 17:41:39 -05:00