Commit Graph
63 Commits
Author SHA1 Message Date
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00
alpayariyak e7b340d73b Update vLLM base image 2024-01-31 03:37:55 +00:00
alpayariyak 12d6f0778e Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer 2024-01-31 03:13:21 +00:00
Alpay AriyakandGitHub fa5556434c Bug fix 2024-01-29 11:26:22 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
Alpay Ariyakandalpayariyak 368c5f87fb Temporary Dockerfile Fix
Temporary Dockerfile fix
2024-01-25 20:48:26 -05:00
alpayariyak 584852f0f6 Docker-Protected HF Token, Refactor, Better Documentation 2024-01-18 18:20:51 -05:00
alpayariyak c59038e902 Fixed CUDA 11.8 workers 2023-12-20 07:50:14 +00:00
Alpay AriyakandGitHub c321d828ab Update Dockerfile 2023-12-19 00:58:45 -05:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00
alpayariyak 22de0b322d Complete refactor, improved overall functionality 2023-12-13 10:33:03 +00:00
alpayariyak ef89868570 Small changes 2023-12-12 20:42:27 +00:00
alpayariyak b51827388c bug fix 2023-12-09 19:06:41 +00:00
alpayariyak 8578581c6a Small changes 2023-12-09 18:50:00 +00:00
alpayariyak f1afbcbe4f CPU Limit 2023-12-09 18:47:45 +00:00
alpayariyak 5e38a1a2fa Multi-GPU Fix 2023-12-09 10:36:26 +00:00
Justin Merrell bcf37e7bd2 Update Dockerfile 2023-12-08 20:42:15 -05:00
Justin Merrell deb6f67a25 Update Dockerfile 2023-12-08 20:42:05 -05:00
Justin Merrell 11e6495304 Update Dockerfile 2023-12-08 20:28:44 -05:00
Justin Merrell 35a74475f2 Update Dockerfile 2023-12-08 20:23:23 -05:00
alpayariyak 8dbd8f15dd Token counting, cuda-related changes 2023-12-06 19:02:12 +00:00
alpayariyak b705544bd4 model packing, cuda version selection on build 2023-12-06 18:52:21 +00:00
alpayariyak 48c520b240 Version Bump + small changes 2023-12-06 16:48:12 +00:00
alpayariyak 66c23d45b0 Model Downloading, no GPU 2023-12-06 01:43:09 +00:00
alpayariyak b4c61e61c4 Packing model into worker for internal use 2023-12-05 20:26:56 -05:00
Justin Merrell bd3ede60f3 Update Dockerfile 2023-11-29 22:32:30 -05:00
Justin Merrell 4131b7a16d Update Dockerfile 2023-11-29 22:32:13 -05:00
Justin Merrell 9b9c1cf503 cleanup 2023-11-29 21:32:40 -05:00
alpayariyak 820a21f32c Download logic + other changes 2023-11-28 19:15:08 -05:00
alpayariyak 97480ff0ae handler changes 2023-11-28 23:39:48 +00:00
Justin Merrell 98af834124 fix: added hf_transfer requirement 2023-11-21 20:50:12 -05:00
Justin Merrell d06fb9ea01 Update Dockerfile 2023-11-21 20:20:15 -05:00
Justin Merrell 5ca8e56bbe fix: clean up extra 2023-11-21 19:18:08 -05:00
alpayariyak c34d79ea49 Entrypoint for testing 2023-11-20 18:56:00 -05:00
alpayariyak 8f17764f8a Update Dockerfile 2023-11-15 20:55:17 -05:00
alpayariyak e7c8de4678 Changes to worker 2023-11-15 20:53:04 -05:00
Samuel Willandalpayariyak 4f792062aa vllm 0.2.1.post1 - speed boost, quantization, mistral support 2023-11-14 19:12:58 -05:00
Jorg Doku 06183813cf take advantage of layering 2023-08-31 18:30:59 -05:00
Jorg Doku 54d229cb16 upate metrics 2023-08-31 18:26:37 -05:00
Jorg Doku 84d6047cbf update 2023-08-31 17:41:39 -05:00
Jorg Doku 96c4d70d2e metrics now included 2023-08-31 17:03:59 -05:00
Jorg Doku 51d231c37a done 2023-08-25 18:26:28 -05:00
Jorg Doku e3cdd46a99 update 2023-08-25 12:03:43 -05:00
Jorg Doku d6cce5d5b6 update 2023-08-24 15:46:30 -05:00
Jorg Doku 33ab5600bf update 2023-08-24 13:58:38 -05:00
Jorg Doku 9bf9039151 updating the llm worker 2023-08-18 09:40:02 -05:00
Jorg Doku 72be7ff31d update 2023-08-07 11:28:36 -05:00
Jorg Doku 4e4d3a27fc fixed 2023-08-06 14:17:53 -05:00
Jorg Doku 401a4637fa latest version 2023-08-03 10:05:51 -05:00
Jorg Doku b78a3d4429 update 2023-07-31 14:01:38 -05:00