- You no longer need a linux-based machine or NVIDIA GPUs to build the worker. - Over 3x lighter Docker image size. - OpenAI Chat Completion output format (optional to use). - Extremely fast image build time. - Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token. - Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt. - New environment variables for various configuration. - vLLM Version: 0.2.7
12 lines
282 B
Bash
12 lines
282 B
Bash
#!/bin/bash
|
|
|
|
git clone https://github.com/runpod/vllm-fork-for-sls-worker.git
|
|
|
|
cp -r vllm-fork-for-sls-worker vllm-12.1.0
|
|
cp -r vllm-fork-for-sls-worker vllm-11.8.0
|
|
rm -rf vllm-fork-for-sls-worker
|
|
|
|
cd vllm-11.8.0
|
|
git checkout cuda11.8
|
|
|
|
echo "vLLM Base Image Builder Setup Complete." |