Commit Graph
127 Commits
Author SHA1 Message Date
alpayariyak d25b6f9628 Fix sampling params 2024-03-12 22:15:57 +00:00
alpayariyak c8ee100d80 Small refactor 2024-03-06 17:08:57 +00:00
alpayariyak db7167d57f 0.3.3 2024-03-05 19:14:35 +00:00
Alpay Ariyakandalpayariyak d91ccb866f 0.3.1: bug fixes 2024-02-29 02:55:44 -05:00
alpayariyak 985bbf1cb5 Fix Multi-GPU, tokenizer trust remote code 2024-02-24 03:18:25 +00:00
alpayariyak 708f68d7f8 Final Bug fixes, configurable oai response role, served model name override, documentation 2024-02-23 04:14:03 +00:00
alpayariyak b42d45ce0f Bug fixes, refactors 2024-02-23 03:46:47 +00:00
alpayariyak 9129d0a252 New ENV Vars 2024-02-22 23:58:20 +00:00
alpayariyak 6bcd9d7c67 Preparing for 0.3.0 2024-02-22 18:04:34 -05:00
alpayariyak aed0408f19 Documentation for 0.3.0, small fixes and changes 2024-02-21 19:28:59 -05:00
alpayariyak e191149259 OpenAI Compatibility, Dynamic Batching, Refactor 2024-02-21 04:36:04 +00:00
alpayariyak a94ef66f71 Dynamic Batch Size [needs refactor] 2024-02-06 03:47:27 +00:00
alpayariyak 45081e4037 Add __init__.py 2024-02-06 02:44:35 +00:00
alpayariyak fef8c81cb9 OpenAI Compatible worker, Refactor 2024-02-06 02:44:14 +00:00
alpayariyak 15b06bb687 Merge branch 'main' into openai-sse-output 2024-02-02 20:03:29 -05:00
alpayariyak 8de10468dd Working tokenizer and model download fix
Fix handler startup
2024-02-02 19:55:35 -05:00
alpayariyak b7051d37ca Move test_openai_stream.py 2024-02-02 22:12:26 +00:00
alpayariyak b1720a154d Added download of model extras into weights folder, separate download of tokenizer, making engine.py utilize downloaded tokenizer, model and tokenizer revision 2024-01-31 22:57:32 -05:00
alpayariyak afa33a2875 Handle errors 2024-02-01 03:01:51 +00:00
alpayariyak 068303ce8f OpenAI Proxy Server for EndPoints and more examples 2024-02-01 02:32:32 +00:00
alpayariyak 3bbcf0021b Fix: Yield if tokens left in batch
Temp: default model for testing image
2024-02-01 00:52:39 +00:00
alpayariyak dab8bad906 OpenAI Chat Completions Stream 2024-01-31 23:27:31 +00:00
Casper f4d7c75504 Snapshot download only tokenizer/config related things 2024-01-31 22:16:44 +01:00
Casper 3adc9e3336 Remove unused import 2024-01-31 18:27:49 +01:00
Casper fd00a1ece3 Update to use snapshot_download 2024-01-31 18:24:57 +01:00
Casper 664dd35782 Download tokenizer upon build 2024-01-31 17:59:52 +01:00
alpayariyak 3e2cd080a2 Fix Tensor Parallel 2024-01-31 05:12:58 +00:00
alpayariyak 46eee12819 Simplify Tensor Parallel 2024-01-31 03:28:02 +00:00
alpayariyak 12d6f0778e Fixed Model bake-in, added Custom Chat Templates, Custom Tokenizer 2024-01-31 03:13:21 +00:00
alpayariyak 9fc8e1e54c Non-streaming OpenAI Chat Completions 2024-01-25 23:24:18 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak ef3c303743 Added support for n parameter 2024-01-19 15:34:25 +00:00
alpayariyak 584852f0f6 Docker-Protected HF Token, Refactor, Better Documentation 2024-01-18 18:20:51 -05:00
Alpay AriyakandGitHub ed48093792 bug fix 2024-01-16 18:38:26 -05:00
alpayariyak a5dc8b53b7 Fix Concurrency, Add Max Model Length 2024-01-16 17:58:46 -05:00
alpayariyak fa358708bc Return finish status 2023-12-29 09:08:51 +00:00
alpayariyak e23ab549e4 Small fix 2023-12-29 08:44:54 +00:00
alpayariyak a14c5cd388 New Worker Stable 2023-12-29 08:21:09 +00:00
alpayariyak 2475cd7a66 Local development, disable log requests 2023-12-29 02:21:24 +00:00
alpayariyak ca8b02e392 Concurrency and Prompt Template fix 2023-12-29 00:42:40 +00:00
alpayariyak a69ab1875d vLLM job tracker, Refactor Concurrency Modifier, Serverless Config 2023-12-28 23:06:49 +00:00
alpayariyak f7ac60f802 Fix Concurrency 2023-12-28 22:56:44 +00:00
alpayariyak fe6e618692 Refactor code and Black formatting style 2023-12-22 04:08:55 +00:00
alpayariyak 3f413c1038 Documentation for Chat Templates, Messages format, Refactor Streaming, etc 2023-12-21 00:22:23 +00:00
alpayariyak ba893ed148 Refactor prompt handling in handler.py 2023-12-20 18:40:43 -05:00
alpayariyak 582e21f97f Chat Template 2023-12-20 18:37:54 -05:00
alpayariyak dad02e13d8 vLLM 0.2.6 Stable Worker 2023-12-20 23:06:57 +00:00
alpayariyak c59038e902 Fixed CUDA 11.8 workers 2023-12-20 07:50:14 +00:00
alpayariyak 8202d4cc63 Bug fix 2023-12-19 02:56:21 +00:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00