Multi-GPU Fix
This commit is contained in:
+10
-3
@@ -27,16 +27,23 @@ RUN if [ "$CUDA_VERSION" == "12.1.0" ]; then \
|
|||||||
ADD src .
|
ADD src .
|
||||||
|
|
||||||
ARG MODEL_NAME=""
|
ARG MODEL_NAME=""
|
||||||
ARG MODEL_BASE_PATH=""
|
ARG MODEL_BASE_PATH="/runpod-volume/"
|
||||||
ARG HF_TOKEN=""
|
ARG HF_TOKEN=""
|
||||||
ENV HF_TOKEN=$HF_TOKEN
|
ARG QUANTIZATION=""
|
||||||
|
|
||||||
# Conditionally run download_model.py
|
# Conditionally run download_model.py
|
||||||
RUN if [ -n "$MODEL_NAME" ] && [ -n "$MODEL_BASE_PATH" ]; then \
|
RUN if [ -n "$MODEL_NAME" ]; then \
|
||||||
|
export HF_TOKEN=$HF_TOKEN; \
|
||||||
python3.11 /download_model.py --model $MODEL_NAME --download_dir $MODEL_BASE_PATH; \
|
python3.11 /download_model.py --model $MODEL_NAME --download_dir $MODEL_BASE_PATH; \
|
||||||
export MODEL_NAME=$MODEL_NAME; \
|
export MODEL_NAME=$MODEL_NAME; \
|
||||||
export MODEL_BASE_PATH=$MODEL_BASE_PATH; \
|
export MODEL_BASE_PATH=$MODEL_BASE_PATH; \
|
||||||
fi
|
fi
|
||||||
|
|
||||||
|
RUN if [ -n "$QUANTIZATION" ]; then \
|
||||||
|
export QUANTIZATION=$QUANTIZATION; \
|
||||||
|
fi
|
||||||
|
|
||||||
|
RUN python3.11 -m pip install pydantic==1.10.13
|
||||||
|
|
||||||
# Start the handler
|
# Start the handler
|
||||||
CMD ["python3.11", "/handler.py"]
|
CMD ["python3.11", "/handler.py"]
|
||||||
|
|||||||
@@ -10,17 +10,18 @@
|
|||||||
</div>
|
</div>
|
||||||
|
|
||||||
## Setting up the Serverless Worker
|
## Setting up the Serverless Worker
|
||||||
### Docker Arguments
|
|
||||||
#### Required:
|
### Option 1: Pre-Built Image
|
||||||
- `MODEL_NAME`: The Hugging Face model to use.
|
|
||||||
- `STREAMING`: Whether to use HTTP Streaming or not.
|
|
||||||
More information on receiving streaming responses from Serverless Endpoints can be found at [Endpoint URLs](https://docs.runpod.io/docs/serverless-endpoint-urls#streamjob_id), and a detailed example at [Llama2 7B Chat | Streaming Token Outputs](https://docs.runpod.io/reference/llama2-7b-chat#streaming-token-outputs).
|
### Option 2: Build Image with Model Inside
|
||||||
#### Optional:
|
#### Docker Arguments:
|
||||||
- `TOKENIZER`: The specified tokenizer to use. If you want to use the default tokenizer for the model, do not provide this docker argument at all.
|
- `MODEL_NAME`: the Hugging Face model to use.
|
||||||
|
- `MODEL_BASE_PATH`: directory to store the model in
|
||||||
|
- `HUGGING_FACE_HUB_TOKEN`: Your Hugging Face token to access private or gated models. You can get your token [here](https://huggingface.co/settings/token).
|
||||||
- `QUANTIZATION`: `awq` to use AWQ Quantization (Base model must be in AWQ format). `squeezellm` for SqueezeLLM quantization - preliminary support.
|
- `QUANTIZATION`: `awq` to use AWQ Quantization (Base model must be in AWQ format). `squeezellm` for SqueezeLLM quantization - preliminary support.
|
||||||
### Environment Variables:
|
### Environment Variables:
|
||||||
#### Optional:
|
|
||||||
- `HUGGING_FACE_HUB_TOKEN`: Your Hugging Face token to access private or gated models. You can get your token [here](https://huggingface.co/settings/token).
|
|
||||||
### Compatible Models
|
### Compatible Models
|
||||||
- LLaMA & LLaMA-2
|
- LLaMA & LLaMA-2
|
||||||
- Mistral
|
- Mistral
|
||||||
|
|||||||
@@ -1,3 +1,3 @@
|
|||||||
hf_transfer
|
hf_transfer
|
||||||
runpod==1.4.0
|
runpod==1.4.0
|
||||||
huggingface-hub==0.19.4
|
huggingface-hub
|
||||||
|
|||||||
Reference in New Issue
Block a user