Multi-GPU Fix

This commit is contained in:
alpayariyak
2023-12-09 10:36:26 +00:00
parent bcf37e7bd2
commit 5e38a1a2fa
3 changed files with 21 additions and 13 deletions
+10 -3
View File
@@ -27,16 +27,23 @@ RUN if [ "$CUDA_VERSION" == "12.1.0" ]; then \
ADD src . ADD src .
ARG MODEL_NAME="" ARG MODEL_NAME=""
ARG MODEL_BASE_PATH="" ARG MODEL_BASE_PATH="/runpod-volume/"
ARG HF_TOKEN="" ARG HF_TOKEN=""
ENV HF_TOKEN=$HF_TOKEN ARG QUANTIZATION=""
# Conditionally run download_model.py # Conditionally run download_model.py
RUN if [ -n "$MODEL_NAME" ] && [ -n "$MODEL_BASE_PATH" ]; then \ RUN if [ -n "$MODEL_NAME" ]; then \
export HF_TOKEN=$HF_TOKEN; \
python3.11 /download_model.py --model $MODEL_NAME --download_dir $MODEL_BASE_PATH; \ python3.11 /download_model.py --model $MODEL_NAME --download_dir $MODEL_BASE_PATH; \
export MODEL_NAME=$MODEL_NAME; \ export MODEL_NAME=$MODEL_NAME; \
export MODEL_BASE_PATH=$MODEL_BASE_PATH; \ export MODEL_BASE_PATH=$MODEL_BASE_PATH; \
fi fi
RUN if [ -n "$QUANTIZATION" ]; then \
export QUANTIZATION=$QUANTIZATION; \
fi
RUN python3.11 -m pip install pydantic==1.10.13
# Start the handler # Start the handler
CMD ["python3.11", "/handler.py"] CMD ["python3.11", "/handler.py"]
+10 -9
View File
@@ -10,17 +10,18 @@
</div> </div>
## Setting up the Serverless Worker ## Setting up the Serverless Worker
### Docker Arguments
#### Required: ### Option 1: Pre-Built Image
- `MODEL_NAME`: The Hugging Face model to use.
- `STREAMING`: Whether to use HTTP Streaming or not.
More information on receiving streaming responses from Serverless Endpoints can be found at [Endpoint URLs](https://docs.runpod.io/docs/serverless-endpoint-urls#streamjob_id), and a detailed example at [Llama2 7B Chat | Streaming Token Outputs](https://docs.runpod.io/reference/llama2-7b-chat#streaming-token-outputs). ### Option 2: Build Image with Model Inside
#### Optional: #### Docker Arguments:
- `TOKENIZER`: The specified tokenizer to use. If you want to use the default tokenizer for the model, do not provide this docker argument at all. - `MODEL_NAME`: the Hugging Face model to use.
- `MODEL_BASE_PATH`: directory to store the model in
- `HUGGING_FACE_HUB_TOKEN`: Your Hugging Face token to access private or gated models. You can get your token [here](https://huggingface.co/settings/token).
- `QUANTIZATION`: `awq` to use AWQ Quantization (Base model must be in AWQ format). `squeezellm` for SqueezeLLM quantization - preliminary support. - `QUANTIZATION`: `awq` to use AWQ Quantization (Base model must be in AWQ format). `squeezellm` for SqueezeLLM quantization - preliminary support.
### Environment Variables: ### Environment Variables:
#### Optional:
- `HUGGING_FACE_HUB_TOKEN`: Your Hugging Face token to access private or gated models. You can get your token [here](https://huggingface.co/settings/token).
### Compatible Models ### Compatible Models
- LLaMA & LLaMA-2 - LLaMA & LLaMA-2
- Mistral - Mistral
+1 -1
View File
@@ -1,3 +1,3 @@
hf_transfer hf_transfer
runpod==1.4.0 runpod==1.4.0
huggingface-hub==0.19.4 huggingface-hub