1.0.0preview update for Llama 3 support and more (vLLM 0.3.3 -> 0.4.2) (#62)
This commit is contained in:
committed by
alpayariyak
parent
0a5b5bc095
commit
874379a0c5
+8
-6
@@ -1,5 +1,6 @@
|
||||
ARG WORKER_CUDA_VERSION=11.8.0
|
||||
FROM runpod/worker-vllm:base-0.3.2-cuda${WORKER_CUDA_VERSION} AS vllm-base
|
||||
ARG BASE_IMAGE_VERSION=1.0.0
|
||||
FROM runpod/worker-vllm:base-${BASE_IMAGE_VERSION}-cuda${WORKER_CUDA_VERSION} AS vllm-base
|
||||
|
||||
RUN apt-get update -y \
|
||||
&& apt-get install -y python3-pip
|
||||
@@ -19,7 +20,7 @@ ARG MODEL_REVISION=""
|
||||
ARG TOKENIZER_REVISION=""
|
||||
|
||||
ENV MODEL_NAME=$MODEL_NAME \
|
||||
MODEL_REVISION=$REVISION \
|
||||
MODEL_REVISION=$MODEL_REVISION \
|
||||
TOKENIZER_NAME=$TOKENIZER_NAME \
|
||||
TOKENIZER_REVISION=$TOKENIZER_REVISION \
|
||||
BASE_PATH=$BASE_PATH \
|
||||
@@ -27,11 +28,11 @@ ENV MODEL_NAME=$MODEL_NAME \
|
||||
HF_DATASETS_CACHE="${BASE_PATH}/huggingface-cache/datasets" \
|
||||
HUGGINGFACE_HUB_CACHE="${BASE_PATH}/huggingface-cache/hub" \
|
||||
HF_HOME="${BASE_PATH}/huggingface-cache/hub" \
|
||||
HF_TRANSFER=1
|
||||
HF_HUB_ENABLE_HF_TRANSFER=1
|
||||
|
||||
ENV PYTHONPATH="/:/vllm-installation"
|
||||
ENV PYTHONPATH="/:/vllm-workspace"
|
||||
|
||||
COPY builder/download_model.py /download_model.py
|
||||
COPY src/download_model.py /download_model.py
|
||||
RUN --mount=type=secret,id=HF_TOKEN,required=false \
|
||||
if [ -f /run/secrets/HF_TOKEN ]; then \
|
||||
export HF_TOKEN=$(cat /run/secrets/HF_TOKEN); \
|
||||
@@ -42,7 +43,8 @@ RUN --mount=type=secret,id=HF_TOKEN,required=false \
|
||||
|
||||
# Add source files
|
||||
COPY src /src
|
||||
|
||||
# Remove download_model.py
|
||||
RUN rm /download_model.py
|
||||
|
||||
# Start the handler
|
||||
CMD ["python3", "/src/handler.py"]
|
||||
@@ -1,29 +1,28 @@
|
||||
<div align="center">
|
||||
|
||||
# vLLM Serverless Endpoint Worker
|
||||
# OpenAI-Compatible vLLM Serverless Endpoint Worker
|
||||
Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https://github.com/vllm-project/vllm) Inference Engine on RunPod Serverless with just a few clicks.
|
||||
|
||||

|
||||

|
||||

|
||||

|
||||
\
|
||||

|
||||
|
||||

|
||||
|
||||
> [!NOTE]
|
||||
> Update 1.0.0preview is now available, use the image tag `runpod/worker-vllm:dev-cuda12.1.0` or `runpod/worker-vllm:dev-cuda11.8.0`.
|
||||
>
|
||||
> 1. vLLM was updated from version `0.3.3` to `0.4.2` in our latest release, adding compatibility for Llama 3 and other models, as well as increasing performance.
|
||||
>
|
||||
> 2. Worker vLLM is now cached on all RunPod machines, speeding up deployment.
|
||||
>
|
||||
> We will soon be adding more features from the updates, such as multi-LoRA, multi-modality, and more.
|
||||
</div>
|
||||
|
||||
### Worker vLLM 0.3.0: What's New since 0.2.0:
|
||||
- **🚀 Full OpenAI Compatibility 🚀**
|
||||
|
||||
You may now use your deployment with any OpenAI Codebase by changing **only 3 lines** in total. The supported routes are <ins>Chat Completions</ins>, <ins>Completions</ins>, and <ins>Models</ins> - with both streaming and non-streaming.
|
||||
- **Dynamic Batch Size** - time-to-first token(TTFT) as fast no batching, while maintaining the performance of batched token streaming throughout the request.
|
||||
- vLLM 0.2.7 -> 0.3.2
|
||||
- Gemma, DeepSeek MoE and OLMo support.
|
||||
- FP8 KV Cache support
|
||||
- New supported parameters
|
||||
- We're working on adding support for Multi-LoRA ⚙️
|
||||
- Support for a wide range of new settings for your endpoint, such as Custom chat templates.
|
||||
- Fixed Tensor Parallelism, baking model into images, and more bugs.
|
||||
- Refactors and general improvements.
|
||||
## NEW: UI for Deploying vLLM Worker on RunPod console:
|
||||

|
||||
|
||||
|
||||
## Table of Contents
|
||||
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
|
||||
@@ -56,11 +55,7 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https:
|
||||
|
||||
# Setting up the Serverless Worker
|
||||
|
||||
### Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended]
|
||||
> [!TIP]
|
||||
> This is the quickest and easiest way to tes your model, as it does not require you to build a Docker image, upload heavy models to DockerHub and wait for workers to download them. You can use this option to deploy your model in a few clicks. For even more convenience, attach a network storage volume to your Endpoint, which will download the model once and share it across all workers.
|
||||
>
|
||||
> However, for actual deployment, it is recommended that you build an image with the model baked in, which is described in Option 2 - this will ensure the fastest load speeds.
|
||||
### Option 1: Deploy Any Model Using Pre-Built Docker Image from RunPod Web Console, you can also use the new UI. [Recommended]
|
||||
|
||||
We now offer a pre-built Docker Image for the vLLM Worker that you can configure entirely with Environment Variables when creating the RunPod Serverless Endpoint:
|
||||
|
||||
@@ -82,7 +77,7 @@ Below is a summary of the available RunPod Worker images, categorized by image s
|
||||
#### Prerequisites
|
||||
- RunPod Account
|
||||
|
||||
#### Environment Variables
|
||||
#### Environment Variables/Settings
|
||||
> Note: `0` is equivalent to `False` and `1` is equivalent to `True` for boolean values.
|
||||
|
||||
| Name | Default | Type/Choices | Description |
|
||||
@@ -176,6 +171,8 @@ Below are all supported model architectures (and examples of each) that you can
|
||||
- Baichuan & Baichuan2 (`baichuan-inc/Baichuan2-13B-Chat`, `baichuan-inc/Baichuan-7B`, etc.)
|
||||
- BLOOM (`bigscience/bloom`, `bigscience/bloomz`, etc.)
|
||||
- ChatGLM (`THUDM/chatglm2-6b`, `THUDM/chatglm3-6b`, etc.)
|
||||
- Command-R (`CohereForAI/c4ai-command-r-v01`, etc.)
|
||||
- DBRX (`databricks/dbrx-base`, `databricks/dbrx-instruct` etc.)
|
||||
- DeciLM (`Deci/DeciLM-7B`, `Deci/DeciLM-7B-instruct`, etc.)
|
||||
- Falcon (`tiiuae/falcon-7b`, `tiiuae/falcon-40b`, `tiiuae/falcon-rw-7b`, etc.)
|
||||
- Gemma (`google/gemma-2b`, `google/gemma-7b`, etc.)
|
||||
@@ -185,16 +182,23 @@ Below are all supported model architectures (and examples of each) that you can
|
||||
- GPT-NeoX (`EleutherAI/gpt-neox-20b`, `databricks/dolly-v2-12b`, `stabilityai/stablelm-tuned-alpha-7b`, etc.)
|
||||
- InternLM (`internlm/internlm-7b`, `internlm/internlm-chat-7b`, etc.)
|
||||
- InternLM2 (`internlm/internlm2-7b`, `internlm/internlm2-chat-7b`, etc.)
|
||||
- LLaMA & LLaMA-2 (`meta-llama/Llama-2-70b-hf`, `lmsys/vicuna-13b-v1.3`, `young-geng/koala`, `openlm-research/open_llama_13b`, etc.)
|
||||
- Jais (`core42/jais-13b`, `core42/jais-13b-chat`, `core42/jais-30b-v3`, `core42/jais-30b-chat-v3`, etc.)
|
||||
- LLaMA, Llama 2, and Meta Llama 3 (`meta-llama/Meta-Llama-3-8B-Instruct`, `meta-llama/Meta-Llama-3-70B-Instruct`, `meta-llama/Llama-2-70b-hf`, `lmsys/vicuna-13b-v1.3`, `young-geng/koala`, `openlm-research/open_llama_13b`, etc.)
|
||||
- MiniCPM (`openbmb/MiniCPM-2B-sft-bf16`, `openbmb/MiniCPM-2B-dpo-bf16`, etc.)
|
||||
- Mistral (`mistralai/Mistral-7B-v0.1`, `mistralai/Mistral-7B-Instruct-v0.1`, etc.)
|
||||
- Mixtral (`mistralai/Mixtral-8x7B-v0.1`, `mistralai/Mixtral-8x7B-Instruct-v0.1`, etc.)
|
||||
- Mixtral (`mistralai/Mixtral-8x7B-v0.1`, `mistralai/Mixtral-8x7B-Instruct-v0.1`, `mistral-community/Mixtral-8x22B-v0.1`, etc.)
|
||||
- MPT (`mosaicml/mpt-7b`, `mosaicml/mpt-30b`, etc.)
|
||||
- OLMo (`allenai/OLMo-1B`, `allenai/OLMo-7B`, etc.)
|
||||
- OLMo (`allenai/OLMo-1B-hf`, `allenai/OLMo-7B-hf`, etc.)
|
||||
- OPT (`facebook/opt-66b`, `facebook/opt-iml-max-30b`, etc.)
|
||||
- Orion (`OrionStarAI/Orion-14B-Base`, `OrionStarAI/Orion-14B-Chat`, etc.)
|
||||
- Phi (`microsoft/phi-1_5`, `microsoft/phi-2`, etc.)
|
||||
- Phi-3 (`microsoft/Phi-3-mini-4k-instruct`, `microsoft/Phi-3-mini-128k-instruct`, etc.)
|
||||
- Qwen (`Qwen/Qwen-7B`, `Qwen/Qwen-7B-Chat`, etc.)
|
||||
- Qwen2 (`Qwen/Qwen2-7B-beta`, `Qwen/Qwen-7B-Chat-beta`, etc.)
|
||||
- Qwen2 (`Qwen/Qwen1.5-7B`, `Qwen/Qwen1.5-7B-Chat`, etc.)
|
||||
- Qwen2MoE (`Qwen/Qwen1.5-MoE-A2.7B`, `Qwen/Qwen1.5-MoE-A2.7B-Chat`, etc.)
|
||||
- StableLM(`stabilityai/stablelm-3b-4e1t`, `stabilityai/stablelm-base-alpha-7b-v2`, etc.)
|
||||
- Starcoder2(`bigcode/starcoder2-3b`, `bigcode/starcoder2-7b`, `bigcode/starcoder2-15b`, etc.)
|
||||
- Xverse (`xverse/XVERSE-7B-Chat`, `xverse/XVERSE-13B-Chat`, `xverse/XVERSE-65B-Chat`, etc.)
|
||||
- Yi (`01-ai/Yi-6B`, `01-ai/Yi-34B`, etc.)
|
||||
|
||||
# Usage: OpenAI Compatibility
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
variable "PUSH" {
|
||||
default = "true"
|
||||
}
|
||||
|
||||
variable "REPOSITORY" {
|
||||
default = "runpod"
|
||||
}
|
||||
|
||||
variable "BASE_IMAGE_VERSION" {
|
||||
default = "1.0.0preview"
|
||||
}
|
||||
|
||||
group "all" {
|
||||
targets = ["base", "main"]
|
||||
}
|
||||
|
||||
group "base" {
|
||||
targets = ["base-1180", "base-1210"]
|
||||
}
|
||||
|
||||
group "main" {
|
||||
targets = ["worker-1180", "worker-1210"]
|
||||
}
|
||||
|
||||
target "base-1180" {
|
||||
tags = ["${REPOSITORY}/worker-vllm:base-${BASE_IMAGE_VERSION}-cuda11.8.0"]
|
||||
context = "vllm-base-image"
|
||||
dockerfile = "Dockerfile"
|
||||
args = {
|
||||
WORKER_CUDA_VERSION = "11.8.0"
|
||||
}
|
||||
output = ["type=docker,push=${PUSH}"]
|
||||
}
|
||||
|
||||
target "base-1210" {
|
||||
tags = ["${REPOSITORY}/worker-vllm:base-${BASE_IMAGE_VERSION}-cuda12.1.0"]
|
||||
context = "vllm-base-image"
|
||||
dockerfile = "Dockerfile"
|
||||
args = {
|
||||
WORKER_CUDA_VERSION = "12.1.0"
|
||||
}
|
||||
output = ["type=docker,push=${PUSH}"]
|
||||
}
|
||||
|
||||
target "worker-1180" {
|
||||
tags = ["${REPOSITORY}/worker-vllm:worker-${BASE_IMAGE_VERSION}-cuda11.8.0"]
|
||||
context = "."
|
||||
dockerfile = "Dockerfile"
|
||||
args = {
|
||||
BASE_IMAGE_VERSION = "${BASE_IMAGE_VERSION}"
|
||||
WORKER_CUDA_VERSION = "11.8.0"
|
||||
}
|
||||
output = ["type=docker,push=${PUSH}"]
|
||||
}
|
||||
|
||||
target "worker-1210" {
|
||||
tags = ["${REPOSITORY}/worker-vllm:worker-${BASE_IMAGE_VERSION}-cuda12.1.0"]
|
||||
context = "."
|
||||
dockerfile = "Dockerfile"
|
||||
args = {
|
||||
BASE_IMAGE_VERSION = "${BASE_IMAGE_VERSION}"
|
||||
WORKER_CUDA_VERSION = "12.1.0"
|
||||
}
|
||||
output = ["type=docker,push=${PUSH}"]
|
||||
}
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 27 MiB |
@@ -1,5 +1,6 @@
|
||||
import os
|
||||
import shutil
|
||||
from tensorize import serialize_model
|
||||
from huggingface_hub import snapshot_download
|
||||
from vllm.model_executor.weight_utils import prepare_hf_model_weights, Disabledtqdm
|
||||
|
||||
@@ -41,6 +42,10 @@ if __name__ == "__main__":
|
||||
model_folder, hf_weights_files, use_safetensors = prepare_hf_model_weights(model_name_or_path=model, revision=revisions["model"], cache_dir=download_dir)
|
||||
model_extras_folder = download_extras_or_tokenizer(model, download_dir, revisions["model"], extras=True)
|
||||
move_files(model_extras_folder, model_folder)
|
||||
|
||||
if os.environ.get("TENSORIZE_MODEL"):
|
||||
|
||||
|
||||
|
||||
with open("/local_model_path.txt", "w") as f:
|
||||
f.write(model_folder)
|
||||
+6
-1
@@ -5,6 +5,7 @@ import json
|
||||
from dotenv import load_dotenv
|
||||
from torch.cuda import device_count
|
||||
from typing import AsyncGenerator
|
||||
import time
|
||||
|
||||
from vllm import AsyncLLMEngine, AsyncEngineArgs
|
||||
from vllm.entrypoints.openai.serving_chat import OpenAIServingChat
|
||||
@@ -100,7 +101,11 @@ class vLLMEngine:
|
||||
|
||||
def _initialize_llm(self):
|
||||
try:
|
||||
return AsyncLLMEngine.from_engine_args(AsyncEngineArgs(**self.config))
|
||||
start = time.time()
|
||||
engine = AsyncLLMEngine.from_engine_args(AsyncEngineArgs(**self.config))
|
||||
end = time.time()
|
||||
logging.info(f"Initialized vLLM engine in {end - start:.2f}s")
|
||||
return engine
|
||||
except Exception as e:
|
||||
logging.error("Error initializing vLLM engine: %s", e)
|
||||
raise e
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
import logging
|
||||
from http import HTTPStatus
|
||||
from typing import Any, Dict
|
||||
from vllm.utils import random_uuid
|
||||
from vllm.entrypoints.openai.protocol import ErrorResponse
|
||||
from vllm import SamplingParams
|
||||
|
||||
+82
-21
@@ -20,15 +20,21 @@ RUN apt-get update -y \
|
||||
# Set working directory
|
||||
WORKDIR /vllm-installation
|
||||
|
||||
RUN ldconfig /usr/local/cuda-$(echo "$WORKER_CUDA_VERSION" | sed 's/\.0$//')/compat/
|
||||
|
||||
# Install build and runtime dependencies
|
||||
COPY vllm/requirements-${WORKER_CUDA_VERSION}.txt requirements.txt
|
||||
COPY vllm/requirements-common.txt requirements-common.txt
|
||||
COPY vllm/requirements-cuda${WORKER_CUDA_VERSION}.txt requirements-cuda.txt
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
pip install -r requirements.txt
|
||||
pip install -r requirements-cuda.txt
|
||||
|
||||
# Install development dependencies
|
||||
COPY vllm/requirements-dev.txt requirements-dev.txt
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
pip install -r requirements-dev.txt
|
||||
|
||||
ARG torch_cuda_arch_list='7.0 7.5 8.0 8.6 8.9 9.0+PTX'
|
||||
ENV TORCH_CUDA_ARCH_LIST=${torch_cuda_arch_list}
|
||||
|
||||
FROM dev AS build
|
||||
|
||||
@@ -40,26 +46,69 @@ COPY vllm/requirements-build.txt requirements-build.txt
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
pip install -r requirements-build.txt
|
||||
|
||||
# install compiler cache to speed up compilation leveraging local or remote caching
|
||||
RUN apt-get update -y && apt-get install -y ccache
|
||||
|
||||
# Copy necessary files
|
||||
COPY vllm/csrc csrc
|
||||
COPY vllm/setup.py setup.py
|
||||
COPY vllm/cmake cmake
|
||||
COPY vllm/CMakeLists.txt CMakeLists.txt
|
||||
COPY vllm/requirements-common.txt requirements-common.txt
|
||||
COPY vllm/requirements-cuda${WORKER_CUDA_VERSION}.txt requirements-cuda.txt
|
||||
COPY vllm/pyproject.toml pyproject.toml
|
||||
COPY vllm/vllm/__init__.py vllm/__init__.py
|
||||
COPY vllm/vllm vllm
|
||||
|
||||
# Set environment variables for building extensions
|
||||
ARG torch_cuda_arch_list='7.0 7.5 8.0 8.6 8.9 9.0+PTX'
|
||||
ENV TORCH_CUDA_ARCH_LIST=${torch_cuda_arch_list}
|
||||
ARG max_jobs=48
|
||||
ENV MAX_JOBS=${max_jobs}
|
||||
ARG nvcc_threads=1024
|
||||
ENV NVCC_THREADS=${nvcc_threads}
|
||||
ENV WORKER_CUDA_VERSION=${WORKER_CUDA_VERSION}
|
||||
ENV VLLM_INSTALL_PUNICA_KERNELS=0
|
||||
# Build extensions
|
||||
RUN ldconfig /usr/local/cuda-$(echo "$WORKER_CUDA_VERSION" | sed 's/\.0$//')/compat/
|
||||
RUN python3 setup.py build_ext --inplace
|
||||
ENV CCACHE_DIR=/root/.cache/ccache
|
||||
RUN --mount=type=cache,target=/root/.cache/ccache \
|
||||
--mount=type=cache,target=/root/.cache/pip \
|
||||
python3 setup.py bdist_wheel --dist-dir=dist
|
||||
|
||||
FROM nvidia/cuda:${WORKER_CUDA_VERSION}-runtime-ubuntu22.04 AS vllm-base
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
pip cache remove vllm_nccl*
|
||||
|
||||
FROM dev as flash-attn-builder
|
||||
# max jobs used for build
|
||||
# flash attention version
|
||||
ARG flash_attn_version=v2.5.8
|
||||
ENV FLASH_ATTN_VERSION=${flash_attn_version}
|
||||
|
||||
WORKDIR /usr/src/flash-attention-v2
|
||||
|
||||
# Download the wheel or build it if a pre-compiled release doesn't exist
|
||||
RUN pip --verbose wheel flash-attn==${FLASH_ATTN_VERSION} \
|
||||
--no-build-isolation --no-deps --no-cache-dir
|
||||
|
||||
FROM dev as NCCL-installer
|
||||
|
||||
# Re-declare ARG after FROM
|
||||
ARG WORKER_CUDA_VERSION
|
||||
|
||||
# Update and install necessary libraries
|
||||
RUN apt-get update -y \
|
||||
&& apt-get install -y wget
|
||||
|
||||
# Install NCCL library
|
||||
RUN if [ "$WORKER_CUDA_VERSION" = "11.8.0" ]; then \
|
||||
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.0-1_all.deb \
|
||||
&& dpkg -i cuda-keyring_1.0-1_all.deb \
|
||||
&& apt-get update \
|
||||
&& apt install -y libnccl2=2.15.5-1+cuda11.8 libnccl-dev=2.15.5-1+cuda11.8; \
|
||||
elif [ "$WORKER_CUDA_VERSION" = "12.1.0" ]; then \
|
||||
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.0-1_all.deb \
|
||||
&& dpkg -i cuda-keyring_1.0-1_all.deb \
|
||||
&& apt-get update \
|
||||
&& apt install -y libnccl2=2.17.1-1+cuda12.1 libnccl-dev=2.17.1-1+cuda12.1; \
|
||||
else \
|
||||
echo "Unsupported CUDA version: $WORKER_CUDA_VERSION"; \
|
||||
exit 1; \
|
||||
fi
|
||||
|
||||
FROM nvidia/cuda:${WORKER_CUDA_VERSION}-base-ubuntu22.04 AS vllm-base
|
||||
|
||||
# Re-declare ARG after FROM
|
||||
ARG WORKER_CUDA_VERSION
|
||||
@@ -69,20 +118,32 @@ RUN apt-get update -y \
|
||||
&& apt-get install -y python3-pip
|
||||
|
||||
# Set working directory
|
||||
WORKDIR /vllm-installation
|
||||
WORKDIR /vllm-workspace
|
||||
|
||||
RUN ldconfig /usr/local/cuda-$(echo "$WORKER_CUDA_VERSION" | sed 's/\.0$//')/compat/
|
||||
|
||||
# Install runtime dependencies
|
||||
COPY vllm/requirements-${WORKER_CUDA_VERSION}.txt requirements.txt
|
||||
RUN --mount=type=bind,from=build,src=/vllm-installation/dist,target=/vllm-workspace/dist \
|
||||
--mount=type=cache,target=/root/.cache/pip \
|
||||
pip install dist/*.whl --verbose
|
||||
|
||||
RUN --mount=type=bind,from=flash-attn-builder,src=/usr/src/flash-attention-v2,target=/usr/src/flash-attention-v2 \
|
||||
--mount=type=cache,target=/root/.cache/pip \
|
||||
pip install /usr/src/flash-attention-v2/*.whl --no-cache-dir
|
||||
|
||||
FROM vllm-base AS runtime
|
||||
|
||||
# install additional dependencies for openai api server
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
pip install -r requirements.txt
|
||||
|
||||
# Copy built files from the build stage
|
||||
COPY --from=build /vllm-installation/vllm/*.so /vllm-installation/vllm/
|
||||
COPY vllm/vllm vllm
|
||||
pip install accelerate hf_transfer modelscope tensorizer
|
||||
|
||||
# Set PYTHONPATH environment variable
|
||||
ENV PYTHONPATH="/"
|
||||
|
||||
# Copy NCCL library
|
||||
COPY --from=NCCL-installer /usr/lib/x86_64-linux-gnu/libnccl.so.2 /usr/lib/x86_64-linux-gnu/libnccl.so.2
|
||||
# Set the VLLM_NCCL_SO_PATH environment variable
|
||||
ENV VLLM_NCCL_SO_PATH="/usr/lib/x86_64-linux-gnu/libnccl.so.2"
|
||||
|
||||
|
||||
# Validate the installation
|
||||
RUN python3 -c "import sys; print(sys.path); import vllm; print(vllm.__file__)"
|
||||
RUN python3 -c "import vllm; print(vllm.__file__)"
|
||||
+1
-1
Submodule vllm-base-image/vllm updated: c46d230a62...ba8f5e79e1
@@ -1 +1,3 @@
|
||||
version: '0.3.3'
|
||||
version: '0.3.3'
|
||||
dev_version: '0.4.2'
|
||||
worker_dev_version: '1.0.0preview'
|
||||
Reference in New Issue
Block a user