0.2.0 Release

- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
- Over 3x lighter Docker image size.
- OpenAI Chat Completion output format (optional to use).
- Extremely fast image build time.
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
- New environment variables for various configuration.
- vLLM Version: 0.2.7
This commit is contained in:
alpayariyak
2024-01-25 20:49:15 -05:00
parent 368c5f87fb
commit 4cebe66b36
10 changed files with 365 additions and 139 deletions
+4 -1
View File
@@ -1,6 +1,9 @@
hf_transfer
ray
pandas
pyarrow
runpod==1.5.2
huggingface-hub
packaging
typing-extensions==4.7.1
pydantic
pydantic