Commit Graph
23 Commits
Author SHA1 Message Date
alpayariyak 5bd6f3a75e 0.5.3, any vllm arg as env var, refactor and fixes, moving away from building separate image from vLLM fork 2024-07-25 12:41:48 -07:00
alpayariyak f19ce12ab0 Fix Building Docker with model built-in #71 2024-06-07 15:29:46 -07:00
Alpay Ariyakandalpayariyak 874379a0c5 1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) (#62) 2024-05-09 00:42:08 -04:00
alpayariyak d25b6f9628 Fix sampling params 2024-03-12 22:15:57 +00:00
alpayariyak c8ee100d80 Small refactor 2024-03-06 17:08:57 +00:00
alpayariyak aed0408f19 Documentation for 0.3.0, small fixes and changes 2024-02-21 19:28:59 -05:00
alpayariyak e191149259 OpenAI Compatibility, Dynamic Batching, Refactor 2024-02-21 04:36:04 +00:00
alpayariyak a94ef66f71 Dynamic Batch Size [needs refactor] 2024-02-06 03:47:27 +00:00
alpayariyak fef8c81cb9 OpenAI Compatible worker, Refactor 2024-02-06 02:44:14 +00:00
alpayariyak 9fc8e1e54c Non-streaming OpenAI Chat Completions 2024-01-25 23:24:18 -05:00
alpayariyak 4cebe66b36 0.2.0 Release
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
2024-01-25 20:49:15 -05:00
alpayariyak 584852f0f6 Docker-Protected HF Token, Refactor, Better Documentation 2024-01-18 18:20:51 -05:00
alpayariyak fe6e618692 Refactor code and Black formatting style 2023-12-22 04:08:55 +00:00
alpayariyak 582e21f97f Chat Template 2023-12-20 18:37:54 -05:00
alpayariyak dad02e13d8 vLLM 0.2.6 Stable Worker 2023-12-20 23:06:57 +00:00
alpayariyak 8202d4cc63 Bug fix 2023-12-19 02:56:21 +00:00
alpayariyak f324bef8a0 Bump to vLLM 0.2.6, fixes and improvements 2023-12-19 01:52:55 +00:00
alpayariyak bfa3703d54 Concurrency modifier 2023-12-13 12:26:01 +00:00
alpayariyak 22de0b322d Complete refactor, improved overall functionality 2023-12-13 10:33:03 +00:00
alpayariyak 5acbbf5528 Cleanup 2023-12-06 20:57:48 +00:00
alpayariyak 3a9393bafd Batched Tokens, cleanup 2023-12-05 17:46:28 +00:00
alpayariyak 8d27261645 Logging change 2023-12-01 10:18:48 +00:00
alpayariyak b242b0cedd Refactor handler + add non-streaming, add utils.py 2023-12-01 10:12:56 +00:00