From 00477f213197e3f084c3bd5921418fa9fd81b8dc Mon Sep 17 00:00:00 2001 From: Jorg Doku Date: Thu, 3 Aug 2023 10:29:14 -0500 Subject: [PATCH 1/7] Update README.md --- README.md | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 76089a5..6a71af3 100644 --- a/README.md +++ b/README.md @@ -1,14 +1,30 @@
-

LLM Endpoint | Worker

+

vLLM Endpoint | Serverless Worker

[![CI | Test Worker](https://github.com/runpod-workers/worker-template/actions/workflows/CI-test_worker.yml/badge.svg)](https://github.com/runpod-workers/worker-template/actions/workflows/CI-test_worker.yml)   [![Docker Image](https://github.com/runpod-workers/worker-template/actions/workflows/CD-docker_dev.yml/badge.svg)](https://github.com/runpod-workers/worker-template/actions/workflows/CD-docker_dev.yml) -🚀 | A simple worker that can be used as a starting point to build your own custom RunPod Endpoint API worker. +🚀 | This serverless worker utilizes vLLM (very Large Language Model) behind the scenes and is integrated into RunPod's serverless environment. It supports dynamic auto-scaling using the built-in RunPod autoscaling feature.
+#### Docker Arguments: +1. `HUGGING_FACE_HUB_TOKEN`: Your private Hugging Face token. This token is required for downloading models that necessitate agreement to an End User License Agreement (EULA), such as the llama2 family of models. +2. `MODEL_NAME`: The Hugging Face model to use. Please ensure that the chosen model is supported by vLLM. Refer to the list of supported models for compatibility. +3. `TOKENIZER`: (Optional) The specified tokenizer to use. If you want to use the default tokenizer for the model, do not provide this docker argument at all. +4. `STREAMING`: Whether to use HTTP Streaming or not. Specify True if you want to enable HTTP Streaming; otherwise, omit this argument. + +#### llama2 7B Chat: +`docker build . --platform linux/amd64 --build-arg HUGGING_FACE_HUB_TOKEN=your_hugging_face_token_here --build-arg MODEL_NAME=meta-llama/Llama-2-7b-chat-hf --build-arg TOKENIZER=hf-internal-testing/llama-tokenizer --build-arg STREAMING=True` + +#### llama2 13B Chat: +`docker build . --platform linux/amd64 --build-arg HUGGING_FACE_HUB_TOKEN=your_hugging_face_token_here --build-arg MODEL_NAME=meta-llama/Llama-2-13b-chat-hf --build-arg TOKENIZER=hf-internal-testing/llama-tokenizer --build-arg STREAMING=True` + +Please make sure to replace your_hugging_face_token_here with your actual Hugging Face token to enable model downloads that require it. + +Ensure that you have Docker installed and properly set up before running the docker build commands. Once built, you can deploy this serverless worker in your desired environment with confidence that it will automatically scale based on demand. For further inquiries or assistance, feel free to contact our support team. + ## 📖 | Getting Started 1. Clone this repository. From 7b1e0c29d71e485291b7ef7bb74f4630fd11585b Mon Sep 17 00:00:00 2001 From: Jorg Doku Date: Thu, 3 Aug 2023 10:30:27 -0500 Subject: [PATCH 2/7] Update README.md --- README.md | 35 ----------------------------------- 1 file changed, 35 deletions(-) diff --git a/README.md b/README.md index 6a71af3..68c2349 100644 --- a/README.md +++ b/README.md @@ -24,38 +24,3 @@ Please make sure to replace your_hugging_face_token_here with your actual Hugging Face token to enable model downloads that require it. Ensure that you have Docker installed and properly set up before running the docker build commands. Once built, you can deploy this serverless worker in your desired environment with confidence that it will automatically scale based on demand. For further inquiries or assistance, feel free to contact our support team. - -## 📖 | Getting Started - -1. Clone this repository. -2. (Optional) Add DockerHub credentials to GitHub Secrets. -3. Add your code to the `src` directory. -4. Update the `handler.py` file to load models and process requests. -5. Add any dependencies to the `requirements.txt` file. -6. Add any other build time scripts to the`builder` directory, for example, downloading models. -7. Update the `Dockerfile` to include any additional dependencies. - -### CI/CD - -This repository is setup to automatically build and push a docker image to the GitHub Container Registry. You will need to add the following to the GitHub Secrets for this repository to enable this functionality: - -- `DOCKERHUB_USERNAME` | Your DockerHub username for logging in. -- `DOCKERHUB_TOKEN` | Your DockerHub token for logging in. -- `DOCKERHUB_REPO` | The name of the repository you want to push to. -- `DOCKERHUB_IMG` | The name of the image you want to push to. - -The `CD-docker_dev.yml` file will build the image and push it to the `dev` tag, while the `CD-docker_release.yml` file will build the image on releases and tag it with the release version. - -The `CI-test_worker.yml` file will test the worker using the input provided by the `--test_input` argument when calling the file containing your handler. Be sure to update this workflow to install any dependencies you need to run your tests. - -## 💡 | Best Practices - -System dempendency installation, model caching, and other shell tasks should be added to the `builder/setup.sh` this will allow you to easily setup your Dockerfile as well as run CI/CD tasks. - -Models should be part of your docker image, this can be accomplished by either copying them into the image or downloading them during the build process. - -If using the input validation utility from the runpod python package, create a `schemas` python file where you can define the schemas, then import that file into your `handler.py` file. - -## 🔗 | Links - -🐳 [Docker Container](https://hub.docker.com/r/runpod/serverless-hello-world) From da057023728a5eda29623c77de485f8c7d63ef25 Mon Sep 17 00:00:00 2001 From: Jorg Doku Date: Thu, 3 Aug 2023 10:31:56 -0500 Subject: [PATCH 3/7] Update README.md --- README.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/README.md b/README.md index 68c2349..3a4271f 100644 --- a/README.md +++ b/README.md @@ -24,3 +24,11 @@ Please make sure to replace your_hugging_face_token_here with your actual Hugging Face token to enable model downloads that require it. Ensure that you have Docker installed and properly set up before running the docker build commands. Once built, you can deploy this serverless worker in your desired environment with confidence that it will automatically scale based on demand. For further inquiries or assistance, feel free to contact our support team. + +## Test Inputs +The following inputs can be used for testing the model: +`{ + "input": { + "prompt": "Who is the president of the United States?" + } +}` From 569b3059f0e8b3de2e2c50d8369f830c1efb6dd2 Mon Sep 17 00:00:00 2001 From: Jorg Doku Date: Thu, 3 Aug 2023 10:33:52 -0500 Subject: [PATCH 4/7] Update README.md --- README.md | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index 3a4271f..be92b32 100644 --- a/README.md +++ b/README.md @@ -27,8 +27,13 @@ Ensure that you have Docker installed and properly set up before running the doc ## Test Inputs The following inputs can be used for testing the model: -`{ +```json +{ "input": { - "prompt": "Who is the president of the United States?" + "prompt": "Who is the president of the United States?", + "sampling_params": { + "max_tokens": 100 + } } -}` +} +``` From 53f9967ea4f43ac500733ce1289b321e2df1eb0f Mon Sep 17 00:00:00 2001 From: Jorg Doku Date: Thu, 3 Aug 2023 10:45:32 -0500 Subject: [PATCH 5/7] Update README.md --- README.md | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) diff --git a/README.md b/README.md index be92b32..a367970 100644 --- a/README.md +++ b/README.md @@ -25,6 +25,25 @@ Please make sure to replace your_hugging_face_token_here with your actual Huggin Ensure that you have Docker installed and properly set up before running the docker build commands. Once built, you can deploy this serverless worker in your desired environment with confidence that it will automatically scale based on demand. For further inquiries or assistance, feel free to contact our support team. + +## Model Inputs +```json +| Argument | Type | Description | +|--------------------|-----------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| n | int | Number of output sequences to return for the given prompt. | +| best_of | Optional[int] | Number of output sequences that are generated from the prompt. From these `best_of` sequences, the top `n` sequences are returned. `best_of` must be greater than or equal to `n`. This is treated as the beam width when `use_beam_search` is True. By default, `best_of` is set to `n`. | +| presence_penalty | float | Float that penalizes new tokens based on whether they appear in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. | +| frequency_penalty | float | Float that penalizes new tokens based on their frequency in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. | +| temperature | float | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. | +| top_p | float | Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. | +| top_k | int | Integer that controls the number of top tokens to consider. Set to -1 to consider all tokens. | +| use_beam_search | bool | Whether to use beam search instead of sampling. | +| stop | Union[None, str, List[str]] | List of strings that stop the generation when they are generated. The returned output will not contain the stop strings. | +| ignore_eos | bool | Whether to ignore the EOS token and continue generating tokens after the EOS token is generated. | +| max_tokens | int | Maximum number of tokens to generate per output sequence. | +| logprobs | Optional[int] | Number of log probabilities to return per output token. | +``` + ## Test Inputs The following inputs can be used for testing the model: ```json From 04baa8903daee9224e4dd9dbe8b69f69a24b3cb1 Mon Sep 17 00:00:00 2001 From: Jorg Doku Date: Thu, 3 Aug 2023 10:45:57 -0500 Subject: [PATCH 6/7] Update README.md --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index a367970..04fb569 100644 --- a/README.md +++ b/README.md @@ -27,7 +27,7 @@ Ensure that you have Docker installed and properly set up before running the doc ## Model Inputs -```json +``` | Argument | Type | Description | |--------------------|-----------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------| | n | int | Number of output sequences to return for the given prompt. | From 9e836cc6d1124d304489161f644257b1b4f6c113 Mon Sep 17 00:00:00 2001 From: Jorg Doku Date: Thu, 3 Aug 2023 10:48:40 -0500 Subject: [PATCH 7/7] Update README.md --- README.md | 28 ++++++++++++++-------------- 1 file changed, 14 insertions(+), 14 deletions(-) diff --git a/README.md b/README.md index 04fb569..af5bd02 100644 --- a/README.md +++ b/README.md @@ -28,20 +28,20 @@ Ensure that you have Docker installed and properly set up before running the doc ## Model Inputs ``` -| Argument | Type | Description | -|--------------------|-----------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| n | int | Number of output sequences to return for the given prompt. | -| best_of | Optional[int] | Number of output sequences that are generated from the prompt. From these `best_of` sequences, the top `n` sequences are returned. `best_of` must be greater than or equal to `n`. This is treated as the beam width when `use_beam_search` is True. By default, `best_of` is set to `n`. | -| presence_penalty | float | Float that penalizes new tokens based on whether they appear in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. | -| frequency_penalty | float | Float that penalizes new tokens based on their frequency in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. | -| temperature | float | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. | -| top_p | float | Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. | -| top_k | int | Integer that controls the number of top tokens to consider. Set to -1 to consider all tokens. | -| use_beam_search | bool | Whether to use beam search instead of sampling. | -| stop | Union[None, str, List[str]] | List of strings that stop the generation when they are generated. The returned output will not contain the stop strings. | -| ignore_eos | bool | Whether to ignore the EOS token and continue generating tokens after the EOS token is generated. | -| max_tokens | int | Maximum number of tokens to generate per output sequence. | -| logprobs | Optional[int] | Number of log probabilities to return per output token. | +| Argument | Type | Default | Description | +|--------------------|-----------------|-----------|------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| n | int | 1 | Number of output sequences to return for the given prompt. | +| best_of | Optional[int] | None | Number of output sequences that are generated from the prompt. From these `best_of` sequences, the top `n` sequences are returned. `best_of` must be greater than or equal to `n`. This is treated as the beam width when `use_beam_search` is True. By default, `best_of` is set to `n`. | +| presence_penalty | float | 0.0 | Float that penalizes new tokens based on whether they appear in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. | +| frequency_penalty | float | 0.0 | Float that penalizes new tokens based on their frequency in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. | +| temperature | float | 1.0 | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. | +| top_p | float | 1.0 | Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. | +| top_k | int | -1 | Integer that controls the number of top tokens to consider. Set to -1 to consider all tokens. | +| use_beam_search | bool | False | Whether to use beam search instead of sampling. | +| stop | Union[None, str, List[str]] | None | List of strings that stop the generation when they are generated. The returned output will not contain the stop strings. | +| ignore_eos | bool | False | Whether to ignore the EOS token and continue generating tokens after the EOS token is generated. | +| max_tokens | int | 256 | Maximum number of tokens to generate per output sequence. | +| logprobs | Optional[int] | None | Number of log probabilities to return per output token. | ``` ## Test Inputs