Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
91167b873a | ||
|
|
819102cfd6 | ||
|
|
985bbf1cb5 | ||
|
|
6f5718f191 | ||
|
|
235d0d31d0 | ||
|
|
708f68d7f8 | ||
|
|
b42d45ce0f | ||
|
|
7221caceff | ||
|
|
a2d9535652 | ||
|
|
9129d0a252 | ||
|
|
6bcd9d7c67 | ||
|
|
b4204c612c | ||
|
|
7993818f5f | ||
|
|
e97917cc14 | ||
|
|
3549cf24d5 | ||
|
|
aed0408f19 | ||
|
|
e191149259 | ||
|
|
a94ef66f71 | ||
|
|
45081e4037 | ||
|
|
fef8c81cb9 | ||
|
|
15b06bb687 | ||
|
|
b7051d37ca | ||
|
|
afa33a2875 | ||
|
|
068303ce8f | ||
|
|
3bbcf0021b | ||
|
|
dab8bad906 |
@@ -19,6 +19,10 @@ jobs:
|
||||
# DO is a custom runner deployed on DigitalOcean, only available for workflows under the runpod-workers organization.
|
||||
# If you would like to use this workflow, you can replace DO with ubuntu-latest or any other runner.
|
||||
|
||||
strategy:
|
||||
matrix:
|
||||
cuda_version: [11.8.0, 12.1.0]
|
||||
|
||||
steps:
|
||||
- name: Set up QEMU
|
||||
uses: docker/setup-qemu-action@v2
|
||||
@@ -37,4 +41,5 @@ jobs:
|
||||
uses: docker/build-push-action@v4
|
||||
with:
|
||||
push: true
|
||||
tags: ${{ vars.DOCKERHUB_REPO }}/${{ vars.DOCKERHUB_IMG }}:${{ (github.event_name == 'release' && github.event.release.tag_name) || (github.event_name == 'workflow_dispatch' && github.event.inputs.image_tag) || 'dev' }}
|
||||
tags: ${{ vars.DOCKERHUB_REPO }}/${{ vars.DOCKERHUB_IMG }}:${{ (github.event_name == 'release' && github.event.release.tag_name) || (github.event_name == 'workflow_dispatch' && github.event.inputs.image_tag) || 'dev' }}-cuda${{ matrix.cuda_version }}
|
||||
build-args: WORKER_CUDA_VERSION=${{ matrix.cuda_version }}
|
||||
|
||||
+2
-1
@@ -2,4 +2,5 @@
|
||||
runpod.toml
|
||||
*.pyc
|
||||
.env
|
||||
test/*
|
||||
test/*
|
||||
vllm-base/vllm-*
|
||||
|
||||
+1
-1
@@ -1,5 +1,5 @@
|
||||
ARG WORKER_CUDA_VERSION=11.8.0
|
||||
FROM runpod/worker-vllm:base-0.2.2-cuda${WORKER_CUDA_VERSION} AS vllm-base
|
||||
FROM runpod/worker-vllm:base-0.3.0-cuda${WORKER_CUDA_VERSION} AS vllm-base
|
||||
|
||||
RUN apt-get update -y \
|
||||
&& apt-get install -y python3-pip
|
||||
|
||||
@@ -7,86 +7,121 @@
|
||||
Deploy Blazing-fast LLMs powered by [vLLM](https://github.com/vllm-project/vllm) on RunPod Serverless in a few clicks.
|
||||
</div>
|
||||
|
||||
### Worker vLLM 0.2.0 - What's New
|
||||
- You no longer need a linux-based machine or NVIDIA GPUs to build the worker.
|
||||
- Over 3x lighter Docker image size.
|
||||
- OpenAI Chat Completion output format (optional to use).
|
||||
- Extremely fast image build time.
|
||||
- Docker Secrets-protected Hugging Face token support for building the image with a model baked in without exposing your token.
|
||||
- Support for `n` and `best_of` sampling parameters, which allow you to generate multiple responses from a single prompt.
|
||||
- New environment variables for various configuration.
|
||||
- vLLM Version: 0.2.7
|
||||
### Worker vLLM 0.3.0: What's New since 0.2.0:
|
||||
- **🚀 Full OpenAI Compatibility 🚀**
|
||||
|
||||
You may now use your deployment with any OpenAI Codebase by changing **only 3 lines** in total. The supported routes are <ins>Chat Completions</ins>, <ins>Completions</ins>, and <ins>Models</ins> - with both streaming and non-streaming.
|
||||
- **Dynamic Batch Size** - time-to-first token as fast no batching, while maintaining the performance of batched token streaming throughout the request.
|
||||
- vLLM 0.2.7 -> 0.3.2
|
||||
- Gemma, DeepSeek MoE and OLMo support.
|
||||
- FP8 KV Cache support
|
||||
- New supported parameters
|
||||
- We're working on adding support for Multi-LoRA ⚙️
|
||||
- Support for a wide range of new settings for your endpoint, such as Custom chat templates.
|
||||
- Fixed Tensor Parallelism, baking model into images, and more bugs.
|
||||
- Refactors and general improvements.
|
||||
|
||||
## Table of Contents
|
||||
- [Setting up the Serverless Worker](#setting-up-the-serverless-worker)
|
||||
- [Option 1: Deploy Any Model Using Pre-Built Docker Image [**RECOMMENDED**]](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
||||
- [Option 1: Deploy Any Model Using Pre-Built Docker Image **[RECOMMENDED]**](#option-1-deploy-any-model-using-pre-built-docker-image-recommended)
|
||||
- [Prerequisites](#prerequisites)
|
||||
- [Environment Variables](#environment-variables)
|
||||
- [LLM Settings](#llm-settings)
|
||||
- [Tokenizer Settings](#tokenizer-settings)
|
||||
- [Tensor Parallelism (Multi-GPU) Settings](#tensor-parallelism-multi-gpu-settings)
|
||||
- [System Settings](#system-settings)
|
||||
- [Streaming Batch Size](#streaming-batch-size)
|
||||
- [OpenAI Settings](#openai-settings)
|
||||
- [Serverless Settings](#serverless-settings)
|
||||
- [Option 2: Build Docker Image with Model Inside](#option-2-build-docker-image-with-model-inside)
|
||||
- [Prerequisites](#prerequisites-1)
|
||||
- [Arguments](#arguments)
|
||||
- [Example: Building an image with OpenChat-3.5](#example-building-an-image-with-openchat-35)
|
||||
- [(Optional) Including Huggingface Token](#optional-including-huggingface-token)
|
||||
- [Compatible Models](#compatible-models)
|
||||
- [Usage](#usage)
|
||||
- [Endpoint Model Inputs](#endpoint-model-inputs)
|
||||
- [Compatible Model Architectures](#compatible-model-architectures)
|
||||
- [Usage: OpenAI Compatibility](#usage-openai-compatibility)
|
||||
- [Modifying your OpenAI Codebase to use your deployed vLLM Worker](#modifying-your-openai-codebase-to-use-your-deployed-vllm-worker)
|
||||
- [OpenAI Request Input Parameters](#openai-request-input-parameters)
|
||||
- [Chat Completions](#chat-completions)
|
||||
- [Completions](#completions)
|
||||
- [Examples: Using your RunPod endpoint with OpenAI](#examples-using-your-runpod-endpoint-with-openai)
|
||||
- [Usage: standard](#non-openai-usage)
|
||||
- [Input Request Parameters](#input-request-parameters)
|
||||
- [Text Input Formats](#text-input-formats)
|
||||
- [1. `prompt`](#1-prompt)
|
||||
- [2. `messages`](#2-messages)
|
||||
- [Sampling Parameters](#sampling-parameters)
|
||||
|
||||
## Setting up the Serverless Worker
|
||||
# Setting up the Serverless Worker
|
||||
|
||||
### Option 1: Deploy Any Model Using Pre-Built Docker Image [Recommended]
|
||||
> [!TIP]
|
||||
> This is the recommended way to deploy your model, as it does not require you to build a Docker image, upload heavy models to DockerHub and wait for workers to download them. Instead, use this option to deploy your model in a few clicks. For even more convenience, attach a network storage volume to your Endpoint, which will download the model once and share it across all workers.
|
||||
|
||||
We now offer a pre-built Docker Image for the vLLM Worker that you can configure entirely with Environment Variables when creating the RunPod Serverless Endpoint:
|
||||
|
||||
<div align="center">
|
||||
---
|
||||
|
||||
Stable Image: ```runpod/worker-vllm:0.2.3```
|
||||
## RunPod Worker Images
|
||||
|
||||
Development Image: ```runpod/worker-vllm:dev```
|
||||
Below is a summary of the available RunPod Worker images, categorized by image stability and CUDA version compatibility.
|
||||
|
||||
</div>
|
||||
| CUDA Version | Stable Image Tag | Development Image Tag | Note |
|
||||
|--------------|-----------------------------------|-----------------------------------|----------------------------------------------------------------------|
|
||||
| 11.8.0 | `runpod/worker-vllm:0.3.0-cuda11.8.0` | `runpod/worker-vllm:dev-cuda11.8.0` | Available on all RunPod Workers without additional selection needed. |
|
||||
| 12.1.0 | `runpod/worker-vllm:0.3.0-cuda12.1.0` | `runpod/worker-vllm:dev-cuda12.1.0` | When creating an Endpoint, select CUDA Version 12.2 and 12.1 in the filter. |
|
||||
|
||||
This table provides a quick reference to the image tags you should use based on the desired CUDA version and image stability (Stable or Development). Ensure to follow the selection note for CUDA 12.1.0 compatibility.
|
||||
|
||||
---
|
||||
|
||||
#### Prerequisites
|
||||
- RunPod Account
|
||||
|
||||
#### Environment Variables
|
||||
> Note: `0` is equivalent to `False` and `1` is equivalent to `True` for boolean values.
|
||||
|
||||
**Required**:
|
||||
- `MODEL_NAME`: Hugging Face Model Repository (e.g., `openchat/openchat-3.5-1210`).
|
||||
| Name | Default | Type/Choices | Description |
|
||||
|-------------------------------------|----------------------|-------------------------------------------|-------------|
|
||||
**LLM Settings**
|
||||
| `MODEL_NAME`**\*** | - | `str` | Hugging Face Model Repository (e.g., `openchat/openchat-3.5-1210`). |
|
||||
| `MODEL_REVISION` | `None` | `str` |Model revision(branch) to load. |
|
||||
| `MAX_MODEL_LENGTH` | Model's maximum | `int` |Maximum number of tokens for the engine to handle per request. |
|
||||
| `BASE_PATH` | `/runpod-volume` | `str` |Storage directory for Huggingface cache and model. Utilizes network storage if attached when pointed at `/runpod-volume`, which will have only one worker download the model once, which all workers will be able to load. If no network volume is present, creates a local directory within each worker. |
|
||||
| `LOAD_FORMAT` | `auto` | `str` |Format to load model in. |
|
||||
| `HF_TOKEN` | - | `str` |Hugging Face token for private and gated models. |
|
||||
| `QUANTIZATION` | `None` | `awq`, `squeezellm`, `gptq` |Quantization of given model. The model must already be quantized. |
|
||||
| `TRUST_REMOTE_CODE` | `0` | boolean as `int` |Trust remote code for Hugging Face models. Can help with Mixtral 8x7B, Quantized models, and unusual models/architectures.
|
||||
| `SEED` | `0` | `int` |Sets random seed for operations. |
|
||||
| `KV_CACHE_DTYPE` | `auto` | boolean as `int` |Data type for kv cache storage. Uses `DTYPE` if set to `auto`. |
|
||||
| `DTYPE` | `auto` | `auto`, `half`, `float16`, `bfloat16`, `float`, `float32` |Sets datatype/precision for model weights and activations. |
|
||||
**Tokenizer Settings**
|
||||
| `TOKENIZER_NAME` | `None` | `str` |Tokenizer repository to use a different tokenizer than the model's default. |
|
||||
| `TOKENIZER_REVISION` | `None` | `str` |Tokenizer revision to load. |
|
||||
| `CUSTOM_CHAT_TEMPLATE` | `None` | `str` of single-line jinja template |Custom chat jinja template. [More Info](https://huggingface.co/docs/transformers/chat_templating) |
|
||||
**System, GPU, and Tensor Parallelism(Multi-GPU) Settings**
|
||||
| `GPU_MEMORY_UTILIZATION` | `0.95` | `float` |Sets GPU VRAM utilization. |
|
||||
| `MAX_PARALLEL_LOADING_WORKERS` | `None` | `int` |Load model sequentially in multiple batches, to avoid RAM OOM when using tensor parallel and large models. |
|
||||
| `BLOCK_SIZE` | `16` | `8`, `16`, `32` |Token block size for contiguous chunks of tokens. |
|
||||
| `SWAP_SPACE` | `4` | `int` |CPU swap space size (GiB) per GPU. |
|
||||
| `ENFORCE_EAGER` | `0` | boolean as `int` |Always use eager-mode PyTorch. If False(`0`), will use eager mode and CUDA graph in hybrid for maximal performance and flexibility. |
|
||||
| `MAX_CONTEXT_LEN_TO_CAPTURE` | `8192` | `int` |Maximum context length covered by CUDA graphs. When a sequence has context length larger than this, we fall back to eager mode.|
|
||||
| `DISABLE_CUSTOM_ALL_REDUCE` | `0` | `int` |Enables or disables custom all reduce. |
|
||||
**Streaming Batch Size Settings**:
|
||||
| `DEFAULT_BATCH_SIZE` | `50` | `int` |Default and Maximum batch size for token streaming to reduce HTTP calls. |
|
||||
| `DEFAULT_MIN_BATCH_SIZE` | `1` | `int` |Batch size for the first request, which will be multiplied by the growth factor every subsequent request. |
|
||||
| `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | `3` | `float` |Growth factor for dynamic batch size. |
|
||||
The way this works is that the first request will have a batch size of `DEFAULT_MIN_BATCH_SIZE`, and each subsequent request will have a batch size of `previous_batch_size * DEFAULT_BATCH_SIZE_GROWTH_FACTOR`. This will continue until the batch size reaches `DEFAULT_BATCH_SIZE`. E.g. for the default values, the batch sizes will be `1, 3, 9, 27, 50, 50, 50, ...`. You can also specify this per request, with inputs `max_batch_size`, `min_batch_size`, and `batch_size_growth_factor`. This has nothing to do with vLLM's internal batching, but rather the number of tokens sent in each HTTP request from the worker |
|
||||
**OpenAI Settings**
|
||||
| `RAW_OPENAI_OUTPUT` | `1` | boolean as `int` |Enables raw OpenAI SSE format string output when streaming. **Required** to be enabled (which it is by default) for OpenAI compatibility. |
|
||||
| `OPENAI_SERVED_MODEL_NAME_OVERRIDE` | `None` | `str` |Overrides the name of the served model from model repo/path to specified name, which you will then be able to use the value for the `model` parameter when making OpenAI requests |
|
||||
| `OPENAI_RESPONSE_ROLE` | `assistant` | `str` |Role of the LLM's Response in OpenAI Chat Completions. |
|
||||
**Serverless Settings**
|
||||
| `MAX_CONCURRENCY` | `300` | `int` |Max concurrent requests per worker. vLLM has an internal queue, so you don't have to worry about limiting by VRAM, this is for improving scaling/load balancing efficiency |
|
||||
| `DISABLE_LOG_STATS` | `1` | boolean as `int` |Enables or disables vLLM stats logging. |
|
||||
| `DISABLE_LOG_REQUESTS` | `1` | boolean as `int` |Enables or disables vLLM request logging. |
|
||||
|
||||
**Optional**:
|
||||
- LLM Settings:
|
||||
- `MODEL_REVISION`: Model revision to load (default: `None`).
|
||||
- `MAX_MODEL_LENGTH`: Maximum number of tokens for the engine to be able to handle. (default: maximum supported by the model)
|
||||
- `BASE_PATH`: Storage directory where huggingface cache and model will be located. (default: `/runpod-volume`, which will utilize network storage if you attach it or create a local directory within the image if you don't)
|
||||
- `LOAD_FORMAT`: Format to load model in (default: `auto`).
|
||||
- `HF_TOKEN`: Hugging Face token for private and gated models (e.g., Llama, Falcon).
|
||||
- `QUANTIZATION`: AWQ (`awq`), SqueezeLLM (`squeezellm`) or GPTQ (`gptq`) Quantization. The specified Model Repo must be of a quantized model. (default: `None`)
|
||||
- `TRUST_REMOTE_CODE`: Trust remote code for Hugging Face (default: `0`)
|
||||
|
||||
- Tokenizer Settings:
|
||||
- `TOKENIZER_NAME`: Tokenizer repository if you would like to use a different tokenizer than the one that comes with the model. (default: `None`, which uses the model's tokenizer)
|
||||
- `TOKENIZER_REVISION`: Tokenizer revision to load (default: `None`).
|
||||
- `CUSTOM_CHAT_TEMPLATE`: Custom chat jinja template, read more about Hugging Face chat templates [here](https://huggingface.co/docs/transformers/chat_templating). (default: `None`)
|
||||
> [!TIP]
|
||||
> If you are facing issues when using Mixtral 8x7B, Quantized models, or handling unusual models/architectures, try setting `TRUST_REMOTE_CODE` to `1`.
|
||||
|
||||
- Tensor Parallelism:
|
||||
Note that the more GPUs you split a model's weights accross, the slower it will be due to inter-GPU communication overhead. If you can fit the model on a single GPU, it is recommended to do so.
|
||||
- `TENSOR_PARALLEL_SIZE`: Number of GPUs to shard the model across (default: `1`).
|
||||
- If you are having issues loading your model with Tensor Parallelism, try decreasing `VLLM_CPU_FRACTION` (default: `1`).
|
||||
|
||||
- System Settings:
|
||||
- `GPU_MEMORY_UTILIZATION`: GPU VRAM utilization (default: `0.98`).
|
||||
- `MAX_PARALLEL_LOADING_WORKERS`: Maximum number of parallel workers for loading models, for non-Tensor Parallel only. (default: `number of available CPU cores` if `TENSOR_PARALLEL_SIZE` is `1`, otherwise `None`).
|
||||
|
||||
|
||||
- Serverless Settings:
|
||||
- `MAX_CONCURRENCY`: Max concurrent requests. (default: `100`)
|
||||
- `DEFAULT_BATCH_SIZE`: Token streaming batch size (default: `30`). This reduces the number of HTTP calls, increasing speed 8-10x vs non-batching, matching non-streaming performance.
|
||||
- `ALLOW_OPENAI_FORMAT`: Whether to allow users to specify `use_openai_format` to get output in OpenAI format. (default: `1`)
|
||||
- `DISABLE_LOG_STATS`: Enable (`0`) or disable (`1`) vLLM stats logging.
|
||||
- `DISABLE_LOG_REQUESTS`: Enable (`0`) or disable (`1`) request logging.
|
||||
|
||||
### Option 2: Build Docker Image with Model Inside
|
||||
To build an image with the model baked in, you must specify the following docker arguments when building the image.
|
||||
@@ -102,7 +137,7 @@ To build an image with the model baked in, you must specify the following docker
|
||||
- `MODEL_REVISION`: Model revision to load (default: `main`).
|
||||
- `BASE_PATH`: Storage directory where huggingface cache and model will be located. (default: `/runpod-volume`, which will utilize network storage if you attach it or create a local directory within the image if you don't. If your intention is to bake the model into the image, you should set this to something like `/models` to make sure there are no issues if you were to accidentally attach network storage.)
|
||||
- `QUANTIZATION`
|
||||
- `WORKER_CUDA_VERSION`: `11.8.0` or `12.1.0` (default: `11.8.0` due to a small amount of workers not having CUDA 12.1 support yet. `12.1.0` is recommended for optimal performance).
|
||||
- `WORKER_CUDA_VERSION`: `11.8.0` or `12.1.0` (default: `11.8.0` due to a small number of workers not having CUDA 12.1 support yet. `12.1.0` is recommended for optimal performance).
|
||||
- `TOKENIZER_NAME`: Tokenizer repository if you would like to use a different tokenizer than the one that comes with the model. (default: `None`, which uses the model's tokenizer)
|
||||
- `TOKENIZER_REVISION`: Tokenizer revision to load (default: `main`).
|
||||
|
||||
@@ -128,96 +163,350 @@ export HF_TOKEN="your_token_here"
|
||||
docker build -t username/image:tag --secret id=HF_TOKEN --build-arg MODEL_NAME="openchat/openchat_3.5" .
|
||||
```
|
||||
|
||||
### Compatible Model Architectures
|
||||
- Mistral (`mistralai/Mistral-7B-v0.1`, `mistralai/Mistral-7B-Instruct-v0.1`, etc.)
|
||||
- Mixtral (`mistralai/Mixtral-8x7B-v0.1`, `mistralai/Mixtral-8x7B-Instruct-v0.1`, etc.)
|
||||
- Phi (`microsoft/phi-1_5`, `microsoft/phi-2`, etc.)
|
||||
- LLaMA & LLaMA-2 (`meta-llama/Llama-2-70b-hf`, `lmsys/vicuna-13b-v1.3`, `young-geng/koala`, `openlm-research/open_llama_13b`, etc.)
|
||||
- Qwen2 (`Qwen/Qwen2-7B-beta`, `Qwen/Qwen-7B-Chat-beta`, etc.)
|
||||
- StableLM(`stabilityai/stablelm-3b-4e1t`, `stabilityai/stablelm-base-alpha-7b-v2`, etc.)
|
||||
- Yi (`01-ai/Yi-6B`, `01-ai/Yi-34B`, etc.)
|
||||
- Qwen (`Qwen/Qwen-7B`, `Qwen/Qwen-7B-Chat`, etc.)
|
||||
## Compatible Model Architectures
|
||||
Below are all supported model architectures (and examples of each) that you can deploy using the vLLM Worker. You can deploy **any model on HuggingFace**, as long as its base architecture is one of the following:
|
||||
|
||||
- Aquila & Aquila2 (`BAAI/AquilaChat2-7B`, `BAAI/AquilaChat2-34B`, `BAAI/Aquila-7B`, `BAAI/AquilaChat-7B`, etc.)
|
||||
- Baichuan & Baichuan2 (`baichuan-inc/Baichuan2-13B-Chat`, `baichuan-inc/Baichuan-7B`, etc.)
|
||||
- BLOOM (`bigscience/bloom`, `bigscience/bloomz`, etc.)
|
||||
- ChatGLM (`THUDM/chatglm2-6b`, `THUDM/chatglm3-6b`, etc.)
|
||||
- DeciLM (`Deci/DeciLM-7B`, `Deci/DeciLM-7B-instruct`, etc.)
|
||||
- Falcon (`tiiuae/falcon-7b`, `tiiuae/falcon-40b`, `tiiuae/falcon-rw-7b`, etc.)
|
||||
- Gemma (`google/gemma-2b`, `google/gemma-7b`, etc.)
|
||||
- GPT-2 (`gpt2`, `gpt2-xl`, etc.)
|
||||
- GPT BigCode (`bigcode/starcoder`, `bigcode/gpt_bigcode-santacoder`, etc.)
|
||||
- GPT-J (`EleutherAI/gpt-j-6b`, `nomic-ai/gpt4all-j`, etc.)
|
||||
- GPT-NeoX (`EleutherAI/gpt-neox-20b`, `databricks/dolly-v2-12b`, `stabilityai/stablelm-tuned-alpha-7b`, etc.)
|
||||
- InternLM (`internlm/internlm-7b`, `internlm/internlm-chat-7b`, etc.)
|
||||
- InternLM2 (`internlm/internlm2-7b`, `internlm/internlm2-chat-7b`, etc.)
|
||||
- LLaMA & LLaMA-2 (`meta-llama/Llama-2-70b-hf`, `lmsys/vicuna-13b-v1.3`, `young-geng/koala`, `openlm-research/open_llama_13b`, etc.)
|
||||
- Mistral (`mistralai/Mistral-7B-v0.1`, `mistralai/Mistral-7B-Instruct-v0.1`, etc.)
|
||||
- Mixtral (`mistralai/Mixtral-8x7B-v0.1`, `mistralai/Mixtral-8x7B-Instruct-v0.1`, etc.)
|
||||
- MPT (`mosaicml/mpt-7b`, `mosaicml/mpt-30b`, etc.)
|
||||
- OLMo (`allenai/OLMo-1B`, `allenai/OLMo-7B`, etc.)
|
||||
- OPT (`facebook/opt-66b`, `facebook/opt-iml-max-30b`, etc.)
|
||||
- Phi (`microsoft/phi-1_5`, `microsoft/phi-2`, etc.)
|
||||
- Qwen (`Qwen/Qwen-7B`, `Qwen/Qwen-7B-Chat`, etc.)
|
||||
- Qwen2 (`Qwen/Qwen2-7B-beta`, `Qwen/Qwen-7B-Chat-beta`, etc.)
|
||||
- StableLM(`stabilityai/stablelm-3b-4e1t`, `stabilityai/stablelm-base-alpha-7b-v2`, etc.)
|
||||
- Yi (`01-ai/Yi-6B`, `01-ai/Yi-34B`, etc.)
|
||||
|
||||
# Usage: OpenAI Compatibility
|
||||
The vLLM Worker is fully compatible with OpenAI's API, and you can use it with any OpenAI Codebase by changing only 3 lines in total. The supported routes are <ins>Chat Completions</ins>, <ins>Completions</ins> and <ins>Models</ins> - with both streaming and non-streaming.
|
||||
|
||||
## Modifying your OpenAI Codebase to use your deployed vLLM Worker
|
||||
**Python** (similar to Node.js, etc.):
|
||||
1. When initializing the OpenAI Client in your code, change the `api_key` to your RunPod API Key and the `base_url` to your RunPod Serverless Endpoint URL in the following format: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1`, filling in your deployed endpoint ID. For example, if your Endpoint ID is `abc1234`, the URL would be `https://api.runpod.ai/v2/abc1234/openai/v1`.
|
||||
|
||||
- Before:
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
|
||||
```
|
||||
- After:
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key=os.environ.get("RUNPOD_API_KEY"),
|
||||
base_url="https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1",
|
||||
)
|
||||
```
|
||||
2. Change the `model` parameter to your deployed model's name whenever using Completions or Chat Completions.
|
||||
- Before:
|
||||
```python
|
||||
response = client.chat.completions.create(
|
||||
model="gpt-3.5-turbo",
|
||||
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
)
|
||||
```
|
||||
- After:
|
||||
```python
|
||||
response = client.chat.completions.create(
|
||||
model="<YOUR DEPLOYED MODEL REPO/NAME>",
|
||||
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
)
|
||||
```
|
||||
|
||||
**Using http requests**:
|
||||
1. Change the `Authorization` header to your RunPod API Key and the `url` to your RunPod Serverless Endpoint URL in the following format: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1`
|
||||
- Before:
|
||||
```bash
|
||||
curl https://api.openai.com/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer $OPENAI_API_KEY" \
|
||||
-d '{
|
||||
"model": "gpt-4",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Why is RunPod the best platform?"
|
||||
}
|
||||
],
|
||||
"temperature": 0,
|
||||
"max_tokens": 100
|
||||
}'
|
||||
```
|
||||
- After:
|
||||
```bash
|
||||
curl https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer <YOUR OPENAI API KEY>" \
|
||||
-d '{
|
||||
"model": "<YOUR DEPLOYED MODEL REPO/NAME>",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Why is RunPod the best platform?"
|
||||
}
|
||||
],
|
||||
"temperature": 0,
|
||||
"max_tokens": 100
|
||||
}'
|
||||
```
|
||||
|
||||
## OpenAI Request Input Parameters:
|
||||
|
||||
When using the chat completion feature of the vLLM Serverless Endpoint Worker, you can customize your requests with the following parameters:
|
||||
|
||||
### Chat Completions
|
||||
<details>
|
||||
<summary>Supported Chat Completions Inputs and Descriptions</summary>
|
||||
|
||||
| Parameter | Type | Default Value | Description |
|
||||
|--------------------------------|----------------------------------|---------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `messages` | Union[str, List[Dict[str, str]]] | | List of messages, where each message is a dictionary with a `role` and `content`. The model's chat template will be applied to the messages automatically, so the model must have one or it should be specified as `CUSTOM_CHAT_TEMPLATE` env var. |
|
||||
| `model` | str | | The model repo that you've deployed on your RunPod Serverless Endpoint. If you are unsure what the name is or are baking the model in, use the guide to get the list of available models in the **Examples: Using your RunPod endpoint with OpenAI** section |
|
||||
| `temperature` | Optional[float] | 0.7 | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. |
|
||||
| `top_p` | Optional[float] | 1.0 | Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. |
|
||||
| `n` | Optional[int] | 1 | Number of output sequences to return for the given prompt. |
|
||||
| `max_tokens` | Optional[int] | None | Maximum number of tokens to generate per output sequence. |
|
||||
| `seed` | Optional[int] | None | Random seed to use for the generation. |
|
||||
| `stop` | Optional[Union[str, List[str]]] | list | List of strings that stop the generation when they are generated. The returned output will not contain the stop strings. |
|
||||
| `stream` | Optional[bool] | False | Whether to stream or not |
|
||||
| `presence_penalty` | Optional[float] | 0.0 | Float that penalizes new tokens based on whether they appear in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. |
|
||||
| `frequency_penalty` | Optional[float] | 0.0 | Float that penalizes new tokens based on their frequency in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. |
|
||||
| `logit_bias` | Optional[Dict[str, float]] | None | Unsupported by vLLM |
|
||||
| `user` | Optional[str] | None | Unsupported by vLLM |
|
||||
Additional parameters supported by vLLM:
|
||||
| `best_of` | Optional[int] | None | Number of output sequences that are generated from the prompt. From these `best_of` sequences, the top `n` sequences are returned. `best_of` must be greater than or equal to `n`. This is treated as the beam width when `use_beam_search` is True. By default, `best_of` is set to `n`. |
|
||||
| `top_k` | Optional[int] | -1 | Integer that controls the number of top tokens to consider. Set to -1 to consider all tokens. |
|
||||
| `ignore_eos` | Optional[bool] | False | Whether to ignore the EOS token and continue generating tokens after the EOS token is generated. |
|
||||
| `use_beam_search` | Optional[bool] | False | Whether to use beam search instead of sampling. |
|
||||
| `stop_token_ids` | Optional[List[int]] | list | List of tokens that stop the generation when they are generated. The returned output will contain the stop tokens unless the stop tokens are special tokens. |
|
||||
| `skip_special_tokens` | Optional[bool] | True | Whether to skip special tokens in the output. |
|
||||
| `spaces_between_special_tokens`| Optional[bool] | True | Whether to add spaces between special tokens in the output. Defaults to True. |
|
||||
| `add_generation_prompt` | Optional[bool] | True | Read more [here](https://huggingface.co/docs/transformers/main/en/chat_templating#what-are-generation-prompts) |
|
||||
| `echo` | Optional[bool] | False | Echo back the prompt in addition to the completion |
|
||||
| `repetition_penalty` | Optional[float] | 1.0 | Float that penalizes new tokens based on whether they appear in the prompt and the generated text so far. Values > 1 encourage the model to use new tokens, while values < 1 encourage the model to repeat tokens. |
|
||||
| `min_p` | Optional[float] | 0.0 | Float that represents the minimum probability for a token to |
|
||||
| `length_penalty` | Optional[float] | 1.0 | Float that penalizes sequences based on their length. Used in beam search.. |
|
||||
| `include_stop_str_in_output` | Optional[bool] | False | Whether to include the stop strings in output text. Defaults to False.|
|
||||
</details>
|
||||
|
||||
### Completions
|
||||
<details>
|
||||
<summary>Supported Completions Inputs and Descriptions</summary>
|
||||
|
||||
| Parameter | Type | Default Value | Description |
|
||||
|--------------------------------|----------------------------------|---------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `model` | str | | The model repo that you've deployed on your RunPod Serverless Endpoint. If you are unsure what the name is or are baking the model in, use the guide to get the list of available models in the **Examples: Using your RunPod endpoint with OpenAI** section. |
|
||||
| `prompt` | Union[List[int], List[List[int]], str, List[str]] | | A string, array of strings, array of tokens, or array of token arrays to be used as the input for the model. |
|
||||
| `suffix` | Optional[str] | None | A string to be appended to the end of the generated text. |
|
||||
| `max_tokens` | Optional[int] | 16 | Maximum number of tokens to generate per output sequence. |
|
||||
| `temperature` | Optional[float] | 1.0 | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. |
|
||||
| `top_p` | Optional[float] | 1.0 | Float that controls the cumulative probability of the top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. |
|
||||
| `n` | Optional[int] | 1 | Number of output sequences to return for the given prompt. |
|
||||
| `stream` | Optional[bool] | False | Whether to stream the output. |
|
||||
| `logprobs` | Optional[int] | None | Number of log probabilities to return per output token. |
|
||||
| `echo` | Optional[bool] | False | Whether to echo back the prompt in addition to the completion. |
|
||||
| `stop` | Optional[Union[str, List[str]]] | list | List of strings that stop the generation when they are generated. The returned output will not contain the stop strings. |
|
||||
| `seed` | Optional[int] | None | Random seed to use for the generation. |
|
||||
| `presence_penalty` | Optional[float] | 0.0 | Float that penalizes new tokens based on whether they appear in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. |
|
||||
| `frequency_penalty` | Optional[float] | 0.0 | Float that penalizes new tokens based on their frequency in the generated text so far. Values > 0 encourage the model to use new tokens, while values < 0 encourage the model to repeat tokens. |
|
||||
| `best_of` | Optional[int] | None | Number of output sequences that are generated from the prompt. From these `best_of` sequences, the top `n` sequences are returned. `best_of` must be greater than or equal to `n`. This parameter influences the diversity of the output. |
|
||||
| `logit_bias` | Optional[Dict[str, float]] | None | Dictionary of token IDs to biases. |
|
||||
| `user` | Optional[str] | None | User identifier for personalizing responses. (Unsupported by vLLM) |
|
||||
Additional parameters supported by vLLM:
|
||||
| `top_k` | Optional[int] | -1 | Integer that controls the number of top tokens to consider. Set to -1 to consider all tokens. |
|
||||
| `ignore_eos` | Optional[bool] | False | Whether to ignore the End Of Sentence token and continue generating tokens after the EOS token is generated. |
|
||||
| `use_beam_search` | Optional[bool] | False | Whether to use beam search instead of sampling for generating outputs. |
|
||||
| `stop_token_ids` | Optional[List[int]] | list | List of tokens that stop the generation when they are generated. The returned output will contain the stop tokens unless the stop tokens are special tokens. |
|
||||
| `skip_special_tokens` | Optional[bool] | True | Whether to skip special tokens in the output. |
|
||||
| `spaces_between_special_tokens`| Optional[bool] | True | Whether to add spaces between special tokens in the output. Defaults to True. |
|
||||
| `repetition_penalty` | Optional[float] | 1.0 | Float that penalizes new tokens based on whether they appear in the prompt and the generated text so far. Values > 1 encourage the model to use new tokens, while values < 1 encourage the model to repeat tokens. |
|
||||
| `min_p` | Optional[float] | 0.0 | Float that represents the minimum probability for a token to be considered, relative to the most likely token. Must be in [0, 1]. Set to 0 to disable. |
|
||||
| `length_penalty` | Optional[float] | 1.0 | Float that penalizes sequences based on their length. Used in beam search. |
|
||||
| `include_stop_str_in_output` | Optional[bool] | False | Whether to include the stop strings in output text. Defaults to False. |
|
||||
</details>
|
||||
|
||||
## Examples: Using your RunPod endpoint with OpenAI
|
||||
|
||||
First, initialize the OpenAI Client with your RunPod API Key and Endpoint URL:
|
||||
```python
|
||||
from openai import OpenAI
|
||||
import os
|
||||
|
||||
# Initialize the OpenAI Client with your RunPod API Key and Endpoint URL
|
||||
client = OpenAI(
|
||||
api_key=os.environ.get("RUNPOD_API_KEY"),
|
||||
base_url="https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1",
|
||||
)
|
||||
```
|
||||
|
||||
### Chat Completions:
|
||||
This is the format used for GPT-4 and focused on instruction-following and chat. Examples of Open Source chat/instruct models include `meta-llama/Llama-2-7b-chat-hf`, `mistralai/Mixtral-8x7B-Instruct-v0.1`, `openchat/openchat-3.5-0106`, `NousResearch/Nous-Hermes-2-Mistral-7B-DPO` and more. However, if your model is a completion-style model with no chat/instruct fine-tune and/or does not have a chat template, you can still use this if you provide a chat template with the environment variable `CUSTOM_CHAT_TEMPLATE`.
|
||||
- **Streaming**:
|
||||
```python
|
||||
# Create a chat completion stream
|
||||
response_stream = client.chat.completions.create(
|
||||
model="<YOUR DEPLOYED MODEL REPO/NAME>",
|
||||
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
stream=True,
|
||||
)
|
||||
# Stream the response
|
||||
for response in response_stream:
|
||||
print(chunk.choices[0].delta.content or "", end="", flush=True)
|
||||
```
|
||||
- **Non-Streaming**:
|
||||
```python
|
||||
# Create a chat completion
|
||||
response = client.chat.completions.create(
|
||||
model="<YOUR DEPLOYED MODEL REPO/NAME>",
|
||||
messages=[{"role": "user", "content": "Why is RunPod the best platform?"}],
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
)
|
||||
# Print the response
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
|
||||
## Usage
|
||||
### Endpoint Model Inputs
|
||||
You may either use a `prompt` or a list of `messages` as input. If you use `messages`, the model's chat template will be applied to the messages automatically, so the model must have one. If you use `prompt`, you may optionally apply the model's chat template to the prompt by setting `apply_chat_template` to `true`.
|
||||
| Argument | Type | Default | Description |
|
||||
|-----------------------|----------------------|--------------------|--------------------------------------------------------------------------------------------------------|
|
||||
| `prompt` | str | | Prompt string to generate text based on. |
|
||||
| `messages` | list[dict[str, str]] | | List of messages, which will automatically have the model's chat template applied. Overrides `prompt`. |
|
||||
| `use_openai_format` | bool | False | Whether to return output in OpenAI format. `ALLOW_OPENAI_FORMAT` environment variable must be `1`, the input should preferably be a `messages` list, but `prompt` is accepted. |
|
||||
| `apply_chat_template` | bool | False | Whether to apply the model's chat template to the `prompt`. |
|
||||
| `sampling_params` | dict | {} | Sampling parameters to control the generation, like temperature, top_p, etc. |
|
||||
| `stream` | bool | False | Whether to enable streaming of output. If True, responses are streamed as they are generated. |
|
||||
| `batch_size` | int | DEFAULT_BATCH_SIZE | The number of tokens to stream every HTTP POST call. |
|
||||
### Completions:
|
||||
This is the format used for models like GPT-3 and is meant for completing the text you provide. Instead of responding to your message, it will try to complete it. Examples of Open Source completions models include `meta-llama/Llama-2-7b-hf`, `mistralai/Mixtral-8x7B-v0.1`, `Qwen/Qwen-72B`, and more. However, you can use any model with this format.
|
||||
- **Streaming**:
|
||||
```python
|
||||
# Create a completion stream
|
||||
response_stream = client.completions.create(
|
||||
model="<YOUR DEPLOYED MODEL REPO/NAME>",
|
||||
prompt="Runpod is the best platform because",
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
stream=True,
|
||||
)
|
||||
# Stream the response
|
||||
for response in response_stream:
|
||||
print(response.choices[0].text or "", end="", flush=True)
|
||||
```
|
||||
- **Non-Streaming**:
|
||||
```python
|
||||
# Create a completion
|
||||
response = client.completions.create(
|
||||
model="<YOUR DEPLOYED MODEL REPO/NAME>",
|
||||
prompt="Runpod is the best platform because",
|
||||
temperature=0,
|
||||
max_tokens=100,
|
||||
)
|
||||
# Print the response
|
||||
print(response.choices[0].text)
|
||||
```
|
||||
|
||||
### Getting a list of names for available models:
|
||||
In the case of baking the model into the image, sometimes the repo may not be accepted as the `model` in the request. In this case, you can list the available models as shown below and use that name.
|
||||
```python
|
||||
models_response = client.models.list()
|
||||
list_of_models = [model.id for model in models_response]
|
||||
print(list_of_models)
|
||||
```
|
||||
|
||||
# Usage: Standard (Non-OpenAI)
|
||||
## Request Input Parameters
|
||||
|
||||
<details>
|
||||
<summary>Click to expand table</summary>
|
||||
|
||||
You may either use a `prompt` or a list of `messages` as input. If you use `messages`, the model's chat template will be applied to the messages automatically, so the model must have one. If you use `prompt`, you may optionally apply the model's chat template to the prompt by setting `apply_chat_template` to `true`.
|
||||
| Argument | Type | Default | Description |
|
||||
|-----------------------|----------------------|--------------------|--------------------------------------------------------------------------------------------------------|
|
||||
| `prompt` | str | | Prompt string to generate text based on. |
|
||||
| `messages` | list[dict[str, str]] | | List of messages, which will automatically have the model's chat template applied. Overrides `prompt`. |
|
||||
| `apply_chat_template` | bool | False | Whether to apply the model's chat template to the `prompt`. |
|
||||
| `sampling_params` | dict | {} | Sampling parameters to control the generation, like temperature, top_p, etc. You can find all available parameters in the `Sampling Parameters` section below. |
|
||||
| `stream` | bool | False | Whether to enable streaming of output. If True, responses are streamed as they are generated. |
|
||||
| `max_batch_size` | int | env var `DEFAULT_BATCH_SIZE` | The maximum number of tokens to stream every HTTP POST call. |
|
||||
| `min_batch_size` | int | env var `DEFAULT_MIN_BATCH_SIZE` | The minimum number of tokens to stream every HTTP POST call. |
|
||||
| `batch_size_growth_factor` | int | env var `DEFAULT_BATCH_SIZE_GROWTH_FACTOR` | The growth factor by which `min_batch_size` will be multiplied for each call until `max_batch_size` is reached. |
|
||||
</details>
|
||||
|
||||
### Sampling Parameters
|
||||
|
||||
Below are all available sampling parameters that you can specify in the `sampling_params` dictionary. If you do not specify any of these parameters, the default values will be used.
|
||||
<details>
|
||||
<summary>Click to expand table</summary>
|
||||
|
||||
| Argument | Type | Default | Description |
|
||||
|---------------------------------|-----------------------------|---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `n` | int | 1 | Number of output sequences generated from the prompt. The top `n` sequences are returned. |
|
||||
| `best_of` | Optional[int] | `n` | Number of output sequences generated from the prompt. The top `n` sequences are returned from these `best_of` sequences. Must be ≥ `n`. Treated as beam width in beam search. Default is `n`. |
|
||||
| `presence_penalty` | float | 0.0 | Penalizes new tokens based on their presence in the generated text so far. Values > 0 encourage new tokens, values < 0 encourage repetition. |
|
||||
| `frequency_penalty` | float | 0.0 | Penalizes new tokens based on their frequency in the generated text so far. Values > 0 encourage new tokens, values < 0 encourage repetition. |
|
||||
| `repetition_penalty` | float | 1.0 | Penalizes new tokens based on their appearance in the prompt and generated text. Values > 1 encourage new tokens, values < 1 encourage repetition. |
|
||||
| `temperature` | float | 1.0 | Controls the randomness of sampling. Lower values make it more deterministic, higher values make it more random. Zero means greedy sampling. |
|
||||
| `top_p` | float | 1.0 | Controls the cumulative probability of top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. |
|
||||
| `top_k` | int | -1 | Controls the number of top tokens to consider. Set to -1 to consider all tokens. |
|
||||
| `min_p` | float | 0.0 | Represents the minimum probability for a token to be considered, relative to the most likely token. Must be in [0, 1]. Set to 0 to disable. |
|
||||
| `use_beam_search` | bool | False | Whether to use beam search instead of sampling. |
|
||||
| `length_penalty` | float | 1.0 | Penalizes sequences based on their length. Used in beam search. |
|
||||
| `early_stopping` | Union[bool, str] | False | Controls stopping condition in beam search. Can be `True`, `False`, or `"never"`. |
|
||||
| `stop` | Union[None, str, List[str]] | None | List of strings that stop generation when produced. The output will not contain these strings. |
|
||||
| `stop_token_ids` | Optional[List[int]] | None | List of token IDs that stop generation when produced. Output contains these tokens unless they are special tokens. |
|
||||
| `ignore_eos` | bool | False | Whether to ignore the End-Of-Sequence token and continue generating tokens after its generation. |
|
||||
| `max_tokens` | int | 16 | Maximum number of tokens to generate per output sequence. |
|
||||
| `skip_special_tokens` | bool | True | Whether to skip special tokens in the output. |
|
||||
| `spaces_between_special_tokens` | bool | True | Whether to add spaces between special tokens in the output. |
|
||||
|
||||
|
||||
### Text Input Formats
|
||||
You may either use a `prompt` or a list of `messages` as input.
|
||||
#### 1. `prompt`
|
||||
1. `prompt`
|
||||
The prompt string can be any string, and the model's chat template will not be applied to it unless `apply_chat_template` is set to `true`, in which case it will be treated as a user message.
|
||||
|
||||
Example:
|
||||
```json
|
||||
"prompt": "..."
|
||||
```
|
||||
#### 2. `messages`
|
||||
Your list can contain any number of messages, and each message can have any role from the following list:
|
||||
- `user`
|
||||
- `assistant`
|
||||
- `system`
|
||||
Example:
|
||||
```json
|
||||
"prompt": "..."
|
||||
```
|
||||
2. `messages`
|
||||
Your list can contain any number of messages, and each message usually can have any role from the following list:
|
||||
- `user`
|
||||
- `assistant`
|
||||
- `system`
|
||||
|
||||
The model's chat template will be applied to the messages automatically, so the model must have one.
|
||||
However, some models may have different roles, so you should check the model's chat template to see which roles are required.
|
||||
|
||||
Example:
|
||||
```json
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
"content": "..."
|
||||
},
|
||||
{
|
||||
"role": "user",
|
||||
"content": "..."
|
||||
},
|
||||
{
|
||||
"role": "assistant",
|
||||
"content": "..."
|
||||
}
|
||||
]
|
||||
```
|
||||
The model's chat template will be applied to the messages automatically, so the model must have one.
|
||||
|
||||
Example:
|
||||
```json
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
"content": "..."
|
||||
},
|
||||
{
|
||||
"role": "user",
|
||||
"content": "..."
|
||||
},
|
||||
{
|
||||
"role": "assistant",
|
||||
"content": "..."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
### Sampling Parameters
|
||||
| Argument | Type | Default | Description |
|
||||
|---------------------------------|-----------------------------|---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `n` | int | 1 | Number of output sequences generated from the prompt. The top `n` sequences are returned. |
|
||||
| `best_of` | Optional[int] | `n` | Number of output sequences generated from the prompt. The top `n` sequences are returned from these `best_of` sequences. Must be ≥ `n`. Treated as beam width in beam search. Default is `n`. |
|
||||
| `presence_penalty` | float | 0.0 | Penalizes new tokens based on their presence in the generated text so far. Values > 0 encourage new tokens, values < 0 encourage repetition. |
|
||||
| `frequency_penalty` | float | 0.0 | Penalizes new tokens based on their frequency in the generated text so far. Values > 0 encourage new tokens, values < 0 encourage repetition. |
|
||||
| `repetition_penalty` | float | 1.0 | Penalizes new tokens based on their appearance in the prompt and generated text. Values > 1 encourage new tokens, values < 1 encourage repetition. |
|
||||
| `temperature` | float | 1.0 | Controls the randomness of sampling. Lower values make it more deterministic, higher values make it more random. Zero means greedy sampling. |
|
||||
| `top_p` | float | 1.0 | Controls the cumulative probability of top tokens to consider. Must be in (0, 1]. Set to 1 to consider all tokens. |
|
||||
| `top_k` | int | -1 | Controls the number of top tokens to consider. Set to -1 to consider all tokens. |
|
||||
| `min_p` | float | 0.0 | Represents the minimum probability for a token to be considered, relative to the most likely token. Must be in [0, 1]. Set to 0 to disable. |
|
||||
| `use_beam_search` | bool | False | Whether to use beam search instead of sampling. |
|
||||
| `length_penalty` | float | 1.0 | Penalizes sequences based on their length. Used in beam search. |
|
||||
| `early_stopping` | Union[bool, str] | False | Controls stopping condition in beam search. Can be `True`, `False`, or `"never"`. |
|
||||
| `stop` | Union[None, str, List[str]] | None | List of strings that stop generation when produced. Output will not contain these strings. |
|
||||
| `stop_token_ids` | Optional[List[int]] | None | List of token IDs that stop generation when produced. Output contains these tokens unless they are special tokens. |
|
||||
| `ignore_eos` | bool | False | Whether to ignore the End-Of-Sequence token and continue generating tokens after its generation. |
|
||||
| `max_tokens` | int | 16 | Maximum number of tokens to generate per output sequence. |
|
||||
| `skip_special_tokens` | bool | True | Whether to skip special tokens in the output. |
|
||||
| `spaces_between_special_tokens` | bool | True | Whether to add spaces between special tokens in the output. |
|
||||
|
||||
@@ -2,7 +2,7 @@ hf_transfer
|
||||
ray
|
||||
pandas
|
||||
pyarrow
|
||||
runpod==1.5.3
|
||||
runpod==1.6.2
|
||||
huggingface-hub
|
||||
packaging
|
||||
typing-extensions==4.7.1
|
||||
|
||||
@@ -0,0 +1,51 @@
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
from utils import count_physical_cores
|
||||
from torch.cuda import device_count
|
||||
|
||||
class EngineConfig:
|
||||
def __init__(self):
|
||||
load_dotenv()
|
||||
self.model_name_or_path, self.hf_home, self.model_revision = self._get_local_or_env("/local_model_path.txt", "MODEL_NAME")
|
||||
self.tokenizer_name_or_path, _, self.tokenizer_revision = self._get_local_or_env("/local_tokenizer_path.txt", "TOKENIZER_NAME")
|
||||
self.tokenizer_name_or_path = self.tokenizer_name_or_path or self.model_name_or_path
|
||||
self.quantization = self._get_quantization()
|
||||
self.config = self._initialize_config()
|
||||
|
||||
def _get_local_or_env(self, local_path, env_var):
|
||||
if os.path.exists(local_path):
|
||||
with open(local_path, "r") as file:
|
||||
return file.read().strip(), None, None
|
||||
return os.getenv(env_var), os.getenv("HF_HOME"), os.getenv(f"{env_var}_REVISION")
|
||||
|
||||
def _get_quantization(self):
|
||||
quantization = os.getenv("QUANTIZATION", "").lower()
|
||||
return quantization if quantization in ["awq", "squeezellm", "gptq"] else None
|
||||
|
||||
def _initialize_config(self):
|
||||
args = {
|
||||
"model": self.model_name_or_path,
|
||||
"revision": self.model_revision,
|
||||
"download_dir": self.hf_home,
|
||||
"quantization": self.quantization,
|
||||
"load_format": os.getenv("LOAD_FORMAT", "auto"),
|
||||
"dtype": os.getenv("DTYPE", "half" if self.quantization else "auto"),
|
||||
"tokenizer": self.tokenizer_name_or_path,
|
||||
"tokenizer_revision": self.tokenizer_revision,
|
||||
"disable_log_stats": bool(int(os.getenv("DISABLE_LOG_STATS", 1))),
|
||||
"disable_log_requests": bool(int(os.getenv("DISABLE_LOG_REQUESTS", 1))),
|
||||
"trust_remote_code": bool(int(os.getenv("TRUST_REMOTE_CODE", 0))),
|
||||
"gpu_memory_utilization": float(os.getenv("GPU_MEMORY_UTILIZATION", 0.95)),
|
||||
"max_parallel_loading_workers": None if device_count() > 1 or not os.getenv("MAX_PARALLEL_LOADING_WORKERS") else int(os.getenv("MAX_PARALLEL_LOADING_WORKERS")),
|
||||
"max_model_len": int(os.getenv("MAX_MODEL_LENGTH")) if os.getenv("MAX_MODEL_LENGTH") else None,
|
||||
"tensor_parallel_size": device_count(),
|
||||
"seed": int(os.getenv("SEED")) if os.getenv("SEED") else None,
|
||||
"kv_cache_dtype": os.getenv("KV_CACHE_DTYPE"),
|
||||
"block_size": int(os.getenv("BLOCK_SIZE")) if os.getenv("BLOCK_SIZE") else None,
|
||||
"swap_space": int(os.getenv("SWAP_SPACE")) if os.getenv("SWAP_SPACE") else None,
|
||||
"max_context_len_to_capture": int(os.getenv("MAX_CONTEXT_LEN_TO_CAPTURE")) if os.getenv("MAX_CONTEXT_LEN_TO_CAPTURE") else None,
|
||||
"disable_custom_all_reduce": bool(int(os.getenv("DISABLE_CUSTOM_ALL_REDUCE", 0))),
|
||||
"enforce_eager": bool(int(os.getenv("ENFORCE_EAGER", 0)))
|
||||
}
|
||||
|
||||
return {k: v for k, v in args.items() if v is not None}
|
||||
+4
-2
@@ -1,7 +1,9 @@
|
||||
from typing import Union
|
||||
|
||||
DEFAULT_BATCH_SIZE = 30
|
||||
DEFAULT_BATCH_SIZE = 50
|
||||
DEFAULT_MAX_CONCURRENCY = 300
|
||||
DEFAULT_BATCH_SIZE_GROWTH_FACTOR = 3
|
||||
DEFAULT_MIN_BATCH_SIZE = 1
|
||||
|
||||
SAMPLING_PARAM_TYPES = {
|
||||
"n": int,
|
||||
@@ -25,4 +27,4 @@ SAMPLING_PARAM_TYPES = {
|
||||
"skip_special_tokens": bool,
|
||||
"spaces_between_special_tokens": bool,
|
||||
"include_stop_str_in_output": bool
|
||||
}
|
||||
}
|
||||
+120
-174
@@ -1,65 +1,53 @@
|
||||
import os
|
||||
import logging
|
||||
from typing import Union, AsyncGenerator
|
||||
import json
|
||||
|
||||
from dotenv import load_dotenv
|
||||
from torch.cuda import device_count
|
||||
from typing import AsyncGenerator
|
||||
|
||||
from vllm import AsyncLLMEngine, AsyncEngineArgs, SamplingParams
|
||||
from vllm.entrypoints.openai.serving_chat import OpenAIServingChat
|
||||
from vllm.entrypoints.openai.protocol import ChatCompletionRequest
|
||||
from transformers import AutoTokenizer
|
||||
from utils import count_physical_cores, DummyRequest
|
||||
from constants import DEFAULT_MAX_CONCURRENCY
|
||||
from dotenv import load_dotenv
|
||||
from vllm.entrypoints.openai.serving_completion import OpenAIServingCompletion
|
||||
from vllm.entrypoints.openai.protocol import ChatCompletionRequest, CompletionRequest, ErrorResponse
|
||||
|
||||
|
||||
class Tokenizer:
|
||||
def __init__(self, tokenizer_name_or_path, tokenizer_revision, trust_remote_code):
|
||||
self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name_or_path, revision=tokenizer_revision, trust_remote_code=trust_remote_code)
|
||||
self.custom_chat_template = os.getenv("CUSTOM_CHAT_TEMPLATE")
|
||||
self.has_chat_template = bool(self.tokenizer.chat_template) or bool(self.custom_chat_template)
|
||||
if self.custom_chat_template and isinstance(self.custom_chat_template, str):
|
||||
self.tokenizer.chat_template = self.custom_chat_template
|
||||
|
||||
def apply_chat_template(self, input: Union[str, list[dict[str, str]]]) -> str:
|
||||
if isinstance(input, list):
|
||||
if not self.has_chat_template:
|
||||
raise ValueError(
|
||||
"Chat template does not exist for this model, you must provide a single string input instead of a list of messages"
|
||||
)
|
||||
elif isinstance(input, str):
|
||||
input = [{"role": "user", "content": input}]
|
||||
else:
|
||||
raise ValueError("Input must be a string or a list of messages")
|
||||
|
||||
return self.tokenizer.apply_chat_template(
|
||||
input, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
from utils import DummyRequest, JobInput, BatchSize, create_error_response
|
||||
from constants import DEFAULT_MAX_CONCURRENCY, DEFAULT_BATCH_SIZE, DEFAULT_BATCH_SIZE_GROWTH_FACTOR, DEFAULT_MIN_BATCH_SIZE
|
||||
from tokenizer import TokenizerWrapper
|
||||
from config import EngineConfig
|
||||
|
||||
|
||||
class vLLMEngine:
|
||||
def __init__(self, engine = None):
|
||||
load_dotenv() # For local development
|
||||
self.config = self._initialize_config()
|
||||
logging.info("vLLM config: %s", self.config)
|
||||
self.tokenizer = Tokenizer(self.config["tokenizer"], self.config["tokenizer_revision"], self.config["trust_remote_code"])
|
||||
self.config = EngineConfig().config
|
||||
self.tokenizer = TokenizerWrapper(self.config.get("tokenizer"), self.config.get("tokenizer_revision"), self.config.get("trust_remote_code"))
|
||||
self.llm = self._initialize_llm() if engine is None else engine
|
||||
self.openai_engine = self._initialize_openai()
|
||||
self.max_concurrency = int(os.getenv("MAX_CONCURRENCY", DEFAULT_MAX_CONCURRENCY))
|
||||
self.default_batch_size = int(os.getenv("DEFAULT_BATCH_SIZE", DEFAULT_BATCH_SIZE))
|
||||
self.batch_size_growth_factor = int(os.getenv("BATCH_SIZE_GROWTH_FACTOR", DEFAULT_BATCH_SIZE_GROWTH_FACTOR))
|
||||
self.min_batch_size = int(os.getenv("MIN_BATCH_SIZE", DEFAULT_MIN_BATCH_SIZE))
|
||||
|
||||
async def generate(self, job_input):
|
||||
generator_args = job_input.__dict__
|
||||
|
||||
if generator_args.pop("use_openai_format"):
|
||||
if self.openai_engine is None:
|
||||
raise ValueError("OpenAI Chat Completion Format is not enabled for this model")
|
||||
generator = self.generate_openai_chat
|
||||
else:
|
||||
generator = self.generate_vllm
|
||||
|
||||
async for batch in generator(**generator_args):
|
||||
yield batch
|
||||
def dynamic_batch_size(self, current_batch_size, batch_size_growth_factor):
|
||||
return min(current_batch_size*batch_size_growth_factor, self.default_batch_size)
|
||||
|
||||
async def generate(self, job_input: JobInput):
|
||||
try:
|
||||
async for batch in self._generate_vllm(
|
||||
llm_input=job_input.llm_input,
|
||||
validated_sampling_params=job_input.validated_sampling_params,
|
||||
batch_size=job_input.max_batch_size,
|
||||
stream=job_input.stream,
|
||||
apply_chat_template=job_input.apply_chat_template,
|
||||
request_id=job_input.request_id,
|
||||
batch_size_growth_factor=job_input.batch_size_growth_factor,
|
||||
min_batch_size=job_input.min_batch_size
|
||||
):
|
||||
yield batch
|
||||
except Exception as e:
|
||||
yield create_error_response(str(e)).model_dump()
|
||||
|
||||
async def generate_vllm(self, llm_input, validated_sampling_params, batch_size, stream, apply_chat_template, request_id: str) -> AsyncGenerator[dict, None]:
|
||||
async def _generate_vllm(self, llm_input, validated_sampling_params, batch_size, stream, apply_chat_template, request_id, batch_size_growth_factor, min_batch_size: str) -> AsyncGenerator[dict, None]:
|
||||
if apply_chat_template or isinstance(llm_input, list):
|
||||
llm_input = self.tokenizer.apply_chat_template(llm_input)
|
||||
validated_sampling_params = SamplingParams(**validated_sampling_params)
|
||||
@@ -70,6 +58,11 @@ class vLLMEngine:
|
||||
batch = {
|
||||
"choices": [{"tokens": []} for _ in range(n_responses)],
|
||||
}
|
||||
|
||||
max_batch_size = batch_size or self.default_batch_size
|
||||
batch_size_growth_factor, min_batch_size = batch_size_growth_factor or self.batch_size_growth_factor, min_batch_size or self.min_batch_size
|
||||
batch_size = BatchSize(max_batch_size, min_batch_size, batch_size_growth_factor)
|
||||
|
||||
|
||||
async for request_output in results_generator:
|
||||
if is_first_output: # Count input tokens only once
|
||||
@@ -84,7 +77,7 @@ class vLLMEngine:
|
||||
batch["choices"][output_index]["tokens"].append(new_output)
|
||||
token_counters["batch"] += 1
|
||||
|
||||
if token_counters["batch"] >= batch_size:
|
||||
if token_counters["batch"] >= batch_size.current_batch_size:
|
||||
batch["usage"] = {
|
||||
"input": n_input_tokens,
|
||||
"output": token_counters["total"],
|
||||
@@ -94,6 +87,7 @@ class vLLMEngine:
|
||||
"choices": [{"tokens": []} for _ in range(n_responses)],
|
||||
}
|
||||
token_counters["batch"] = 0
|
||||
batch_size.update()
|
||||
|
||||
last_output_texts[output_index] = output.text
|
||||
|
||||
@@ -105,87 +99,6 @@ class vLLMEngine:
|
||||
if token_counters["batch"] > 0:
|
||||
batch["usage"] = {"input": n_input_tokens, "output": token_counters["total"]}
|
||||
yield batch
|
||||
|
||||
async def generate_openai_chat(self, llm_input, validated_sampling_params, batch_size, stream, apply_chat_template, request_id: str) -> AsyncGenerator[dict, None]:
|
||||
|
||||
if isinstance(llm_input, str):
|
||||
llm_input = [{"role": "user", "content": llm_input}]
|
||||
logging.warning("OpenAI Chat Completion format requires list input, converting to list and assigning 'user' role")
|
||||
|
||||
if not self.openai_engine:
|
||||
raise ValueError("OpenAI Chat Completion format is disabled")
|
||||
|
||||
chat_completion_request = ChatCompletionRequest(
|
||||
model=self.config["model"],
|
||||
messages=llm_input,
|
||||
stream=stream,
|
||||
**validated_sampling_params,
|
||||
)
|
||||
|
||||
response_generator = await self.openai_engine.create_chat_completion(chat_completion_request, DummyRequest())
|
||||
if not stream:
|
||||
yield json.loads(response_generator.model_dump_json())
|
||||
else:
|
||||
batch_contents = {}
|
||||
batch_latest_choices = {}
|
||||
batch_token_counter = 0
|
||||
last_chunk = {}
|
||||
|
||||
async for chunk_str in response_generator:
|
||||
try:
|
||||
chunk = json.loads(chunk_str.removeprefix("data: ").rstrip("\n\n"))
|
||||
except:
|
||||
continue
|
||||
|
||||
if "choices" in chunk:
|
||||
for choice in chunk["choices"]:
|
||||
choice_index = choice["index"]
|
||||
if "delta" in choice and "content" in choice["delta"]:
|
||||
batch_contents[choice_index] = batch_contents.get(choice_index, []) + [choice["delta"]["content"]]
|
||||
batch_latest_choices[choice_index] = choice
|
||||
batch_token_counter += 1
|
||||
last_chunk = chunk
|
||||
|
||||
if batch_token_counter >= batch_size:
|
||||
for choice_index in batch_latest_choices:
|
||||
batch_latest_choices[choice_index]["delta"]["content"] = batch_contents[choice_index]
|
||||
last_chunk["choices"] = list(batch_latest_choices.values())
|
||||
yield last_chunk
|
||||
|
||||
batch_contents = {}
|
||||
batch_latest_choices = {}
|
||||
batch_token_counter = 0
|
||||
|
||||
if batch_contents:
|
||||
for choice_index in batch_latest_choices:
|
||||
batch_latest_choices[choice_index]["delta"]["content"] = batch_contents[choice_index]
|
||||
last_chunk["choices"] = list(batch_latest_choices.values())
|
||||
yield last_chunk
|
||||
|
||||
def _initialize_config(self):
|
||||
quantization = self._get_quantization()
|
||||
model, download_dir, model_revision = self._get_model_info()
|
||||
tokenizer_name_or_path, tokenizer_revision = self._get_tokenizer_info()
|
||||
if not tokenizer_name_or_path:
|
||||
tokenizer_name_or_path = model
|
||||
|
||||
return {
|
||||
"model": model,
|
||||
"revision": model_revision,
|
||||
"download_dir": download_dir,
|
||||
"quantization": quantization,
|
||||
"load_format": os.getenv("LOAD_FORMAT", "auto"),
|
||||
"dtype": "half" if quantization else "auto",
|
||||
"tokenizer": tokenizer_name_or_path,
|
||||
"tokenizer_revision": tokenizer_revision,
|
||||
"disable_log_stats": bool(int(os.getenv("DISABLE_LOG_STATS", 1))),
|
||||
"disable_log_requests": bool(int(os.getenv("DISABLE_LOG_REQUESTS", 1))),
|
||||
"trust_remote_code": bool(int(os.getenv("TRUST_REMOTE_CODE", 0))),
|
||||
"gpu_memory_utilization": float(os.getenv("GPU_MEMORY_UTILIZATION", 0.95)),
|
||||
"max_parallel_loading_workers": self._get_max_parallel_loading_workers(),
|
||||
"max_model_len": self._get_max_model_len(),
|
||||
"tensor_parallel_size": self._get_num_gpu_shard(),
|
||||
}
|
||||
|
||||
def _initialize_llm(self):
|
||||
try:
|
||||
@@ -193,51 +106,84 @@ class vLLMEngine:
|
||||
except Exception as e:
|
||||
logging.error("Error initializing vLLM engine: %s", e)
|
||||
raise e
|
||||
|
||||
def _initialize_openai(self):
|
||||
if bool(int(os.getenv("ALLOW_OPENAI_FORMAT", 1))) and self.tokenizer.has_chat_template:
|
||||
return OpenAIServingChat(self.llm, self.config["model"], "assistant", self.tokenizer.tokenizer.chat_template)
|
||||
else:
|
||||
return None
|
||||
|
||||
def _get_max_parallel_loading_workers(self):
|
||||
if int(os.getenv("TENSOR_PARALLEL_SIZE", 1)) > 1:
|
||||
return None
|
||||
else:
|
||||
return int(os.getenv("MAX_PARALLEL_LOADING_WORKERS", count_physical_cores()))
|
||||
|
||||
def _get_model_info(self):
|
||||
if os.path.exists("/local_model_path.txt"):
|
||||
model, download_dir, revision = open("/local_model_path.txt", "r").read().strip(), None, None
|
||||
logging.info("Using local model at %s", model)
|
||||
else:
|
||||
model, download_dir, revision = os.getenv("MODEL_NAME"), os.getenv("HF_HOME"), os.getenv("MODEL_REVISION") or None
|
||||
return model, download_dir, revision
|
||||
|
||||
def _get_tokenizer_info(self):
|
||||
if os.path.exists("/local_tokenizer_path.txt"):
|
||||
tokenizer_name_or_path, revision = open("/local_tokenizer_path.txt", "r").read().strip(), None
|
||||
logging.info("Using local tokenizer at %s", tokenizer_name_or_path)
|
||||
else:
|
||||
tokenizer_name_or_path, revision = os.getenv("TOKENIZER_NAME"), os.getenv("TOKENIZER_REVISION") or None
|
||||
return tokenizer_name_or_path, revision
|
||||
|
||||
def _get_num_gpu_shard(self):
|
||||
num_gpu_shard = int(os.getenv("TENSOR_PARALLEL_SIZE", 1))
|
||||
if num_gpu_shard > 1:
|
||||
num_gpu_available = device_count()
|
||||
num_gpu_shard = min(num_gpu_shard, num_gpu_available)
|
||||
logging.info("Using %s GPU shards", num_gpu_shard)
|
||||
return num_gpu_shard
|
||||
|
||||
def _get_max_model_len(self):
|
||||
max_model_len = os.getenv("MAX_MODEL_LENGTH")
|
||||
return int(max_model_len) if max_model_len is not None else None
|
||||
|
||||
def _get_n_current_jobs(self):
|
||||
total_sequences = len(self.llm.engine.scheduler.waiting) + len(self.llm.engine.scheduler.swapped) + len(self.llm.engine.scheduler.running)
|
||||
return total_sequences
|
||||
|
||||
def _get_quantization(self):
|
||||
quantization = os.getenv("QUANTIZATION", "").lower()
|
||||
return quantization if quantization in ["awq", "squeezellm", "gptq"] else None
|
||||
class OpenAIvLLMEngine:
|
||||
def __init__(self, vllm_engine):
|
||||
self.config = vllm_engine.config
|
||||
self.llm = vllm_engine.llm
|
||||
self.served_model_name = os.getenv("OPENAI_SERVED_MODEL_NAME_OVERRIDE") or self.config["model"]
|
||||
self.response_role = os.getenv("OPENAI_RESPONSE_ROLE") or "assistant"
|
||||
self.tokenizer = vllm_engine.tokenizer
|
||||
self.default_batch_size = vllm_engine.default_batch_size
|
||||
self.batch_size_growth_factor, self.min_batch_size = vllm_engine.batch_size_growth_factor, vllm_engine.min_batch_size
|
||||
self._initialize_engines()
|
||||
self.raw_openai_output = bool(int(os.getenv("RAW_OPENAI_OUTPUT", 1)))
|
||||
|
||||
def _initialize_engines(self):
|
||||
self.chat_engine = OpenAIServingChat(
|
||||
self.llm, self.served_model_name, self.response_role,
|
||||
chat_template=self.tokenizer.tokenizer.chat_template
|
||||
)
|
||||
self.completion_engine = OpenAIServingCompletion(self.llm, self.served_model_name)
|
||||
|
||||
async def generate(self, openai_request: JobInput):
|
||||
if openai_request.openai_route == "/v1/models":
|
||||
yield await self._handle_model_request()
|
||||
elif openai_request.openai_route in ["/v1/chat/completions", "/v1/completions"]:
|
||||
async for response in self._handle_chat_or_completion_request(openai_request):
|
||||
yield response
|
||||
else:
|
||||
yield create_error_response("Invalid route").model_dump()
|
||||
|
||||
async def _handle_model_request(self):
|
||||
models = await self.chat_engine.show_available_models()
|
||||
return models.model_dump()
|
||||
|
||||
async def _handle_chat_or_completion_request(self, openai_request: JobInput):
|
||||
if openai_request.openai_route == "/v1/chat/completions":
|
||||
request_class = ChatCompletionRequest
|
||||
generator_function = self.chat_engine.create_chat_completion
|
||||
elif openai_request.openai_route == "/v1/completions":
|
||||
request_class = CompletionRequest
|
||||
generator_function = self.completion_engine.create_completion
|
||||
|
||||
try:
|
||||
request = request_class(
|
||||
**openai_request.openai_input
|
||||
)
|
||||
except Exception as e:
|
||||
yield create_error_response(str(e)).model_dump()
|
||||
return
|
||||
|
||||
response_generator = await generator_function(request, DummyRequest())
|
||||
|
||||
if not openai_request.openai_input.get("stream") or isinstance(response_generator, ErrorResponse):
|
||||
yield response_generator.model_dump()
|
||||
else:
|
||||
batch = []
|
||||
batch_token_counter = 0
|
||||
batch_size = BatchSize(self.default_batch_size, self.min_batch_size, self.batch_size_growth_factor)
|
||||
|
||||
async for chunk_str in response_generator:
|
||||
if "data" in chunk_str:
|
||||
if self.raw_openai_output:
|
||||
data = chunk_str
|
||||
elif "[DONE]" in chunk_str:
|
||||
continue
|
||||
else:
|
||||
data = json.loads(chunk_str.removeprefix("data: ").rstrip("\n\n")) if not self.raw_openai_output else chunk_str
|
||||
batch.append(data)
|
||||
batch_token_counter += 1
|
||||
if batch_token_counter >= batch_size.current_batch_size:
|
||||
if self.raw_openai_output:
|
||||
batch = "".join(batch)
|
||||
yield batch
|
||||
batch = []
|
||||
batch_token_counter = 0
|
||||
batch_size.update()
|
||||
if batch:
|
||||
if self.raw_openai_output:
|
||||
batch = "".join(batch)
|
||||
yield batch
|
||||
|
||||
+6
-3
@@ -1,15 +1,18 @@
|
||||
import os
|
||||
import runpod
|
||||
from utils import JobInput
|
||||
from engine import vLLMEngine
|
||||
from engine import vLLMEngine, OpenAIvLLMEngine
|
||||
|
||||
vllm_engine = vLLMEngine()
|
||||
OpenAIvLLMEngine = OpenAIvLLMEngine(vllm_engine)
|
||||
|
||||
async def handler(job):
|
||||
job_input = JobInput(job["input"])
|
||||
results_generator = vllm_engine.generate(job_input)
|
||||
engine = OpenAIvLLMEngine if job_input.openai_route else vllm_engine
|
||||
results_generator = engine.generate(job_input)
|
||||
async for batch in results_generator:
|
||||
yield batch
|
||||
|
||||
|
||||
runpod.serverless.start(
|
||||
{
|
||||
"handler": handler,
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
from transformers import AutoTokenizer
|
||||
import os
|
||||
from typing import Union
|
||||
|
||||
class TokenizerWrapper:
|
||||
def __init__(self, tokenizer_name_or_path, tokenizer_revision, trust_remote_code):
|
||||
self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name_or_path, revision=tokenizer_revision, trust_remote_code=trust_remote_code)
|
||||
self.custom_chat_template = os.getenv("CUSTOM_CHAT_TEMPLATE")
|
||||
self.has_chat_template = bool(self.tokenizer.chat_template) or bool(self.custom_chat_template)
|
||||
if self.custom_chat_template and isinstance(self.custom_chat_template, str):
|
||||
self.tokenizer.chat_template = self.custom_chat_template
|
||||
|
||||
def apply_chat_template(self, input: Union[str, list[dict[str, str]]]) -> str:
|
||||
if isinstance(input, list):
|
||||
if not self.has_chat_template:
|
||||
raise ValueError(
|
||||
"Chat template does not exist for this model, you must provide a single string input instead of a list of messages"
|
||||
)
|
||||
elif isinstance(input, str):
|
||||
input = [{"role": "user", "content": input}]
|
||||
else:
|
||||
raise ValueError("Input must be a string or a list of messages")
|
||||
|
||||
return self.tokenizer.apply_chat_template(
|
||||
input, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
+33
-4
@@ -1,7 +1,10 @@
|
||||
import logging
|
||||
from http import HTTPStatus
|
||||
from typing import Any, Dict
|
||||
from constants import SAMPLING_PARAM_TYPES
|
||||
from vllm.utils import random_uuid
|
||||
from constants import SAMPLING_PARAM_TYPES, DEFAULT_BATCH_SIZE
|
||||
from vllm.entrypoints.openai.protocol import ErrorResponse
|
||||
|
||||
|
||||
logging.basicConfig(level=logging.INFO)
|
||||
|
||||
@@ -41,12 +44,38 @@ class JobInput:
|
||||
def __init__(self, job):
|
||||
self.llm_input = job.get("messages", job.get("prompt"))
|
||||
self.stream = job.get("stream", False)
|
||||
self.batch_size = job.get("batch_size", DEFAULT_BATCH_SIZE)
|
||||
self.max_batch_size = job.get("max_batch_size")
|
||||
self.apply_chat_template = job.get("apply_chat_template", False)
|
||||
self.use_openai_format = job.get("use_openai_format", False)
|
||||
self.validated_sampling_params = validate_sampling_params(job.get("sampling_params", {}))
|
||||
self.request_id = random_uuid()
|
||||
|
||||
batch_size_growth_factor = job.get("batch_size_growth_factor")
|
||||
self.batch_size_growth_factor = float(batch_size_growth_factor) if batch_size_growth_factor else None
|
||||
min_batch_size = job.get("min_batch_size")
|
||||
self.min_batch_size = int(min_batch_size) if min_batch_size else None
|
||||
self.openai_route = job.get("openai_route")
|
||||
self.openai_input = job.get("openai_input")
|
||||
|
||||
class DummyRequest:
|
||||
async def is_disconnected(self):
|
||||
return False
|
||||
return False
|
||||
|
||||
class BatchSize:
|
||||
def __init__(self, max_batch_size, min_batch_size, batch_size_growth_factor):
|
||||
self.max_batch_size = max_batch_size
|
||||
self.batch_size_growth_factor = batch_size_growth_factor
|
||||
self.min_batch_size = min_batch_size
|
||||
self.is_dynamic = batch_size_growth_factor > 1 and min_batch_size >= 1 and max_batch_size > min_batch_size
|
||||
if self.is_dynamic:
|
||||
self.current_batch_size = min_batch_size
|
||||
else:
|
||||
self.current_batch_size = max_batch_size
|
||||
|
||||
def update(self):
|
||||
if self.is_dynamic:
|
||||
self.current_batch_size = min(self.current_batch_size*self.batch_size_growth_factor, self.max_batch_size)
|
||||
|
||||
def create_error_response(message: str, err_type: str = "BadRequestError", status_code: HTTPStatus = HTTPStatus.BAD_REQUEST) -> ErrorResponse:
|
||||
return ErrorResponse(message=message,
|
||||
type=err_type,
|
||||
code=status_code.value)
|
||||
+11
-2
@@ -6,7 +6,7 @@
|
||||
##########################################################
|
||||
|
||||
# Define the CUDA version for the build
|
||||
ARG WORKER_CUDA_VERSION=12.1.0
|
||||
ARG WORKER_CUDA_VERSION=11.8.0
|
||||
|
||||
FROM nvidia/cuda:${WORKER_CUDA_VERSION}-devel-ubuntu22.04 AS dev
|
||||
|
||||
@@ -17,6 +17,10 @@ ARG WORKER_CUDA_VERSION
|
||||
RUN apt-get update -y \
|
||||
&& apt-get install -y python3-pip git
|
||||
|
||||
RUN if [ "${WORKER_CUDA_VERSION}" = "12.1.0" ]; then \
|
||||
ldconfig /usr/local/cuda-12.1/compat/; \
|
||||
fi
|
||||
|
||||
# Set working directory
|
||||
WORKDIR /vllm-installation
|
||||
|
||||
@@ -25,6 +29,11 @@ COPY vllm-${WORKER_CUDA_VERSION}/requirements.txt requirements.txt
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
pip install -r requirements.txt
|
||||
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
if [ "${WORKER_CUDA_VERSION}" = "11.8.0" ]; then \
|
||||
pip install -U --force-reinstall torch==2.1.2 xformers==0.0.23.post1 --index-url https://download.pytorch.org/whl/cu118; \
|
||||
fi
|
||||
|
||||
# Install development dependencies
|
||||
COPY vllm-${WORKER_CUDA_VERSION}/requirements-dev.txt requirements-dev.txt
|
||||
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||
@@ -67,7 +76,7 @@ ENV NVCC_THREADS=${nvcc_threads}
|
||||
# Build extensions
|
||||
RUN python3 setup.py build_ext --inplace
|
||||
|
||||
FROM nvidia/cuda:${WORKER_CUDA_VERSION}-base-ubuntu22.04 AS vllm-base
|
||||
FROM nvidia/cuda:${WORKER_CUDA_VERSION}-runtime-ubuntu22.04 AS vllm-base
|
||||
|
||||
# Re-declare ARG after FROM
|
||||
ARG WORKER_CUDA_VERSION
|
||||
|
||||
@@ -7,6 +7,6 @@ cp -r vllm-fork-for-sls-worker vllm-11.8.0
|
||||
rm -rf vllm-fork-for-sls-worker
|
||||
|
||||
cd vllm-11.8.0
|
||||
git checkout cuda11.8
|
||||
git checkout cuda-11.8
|
||||
|
||||
echo "vLLM Base Image Builder Setup Complete."
|
||||
Reference in New Issue
Block a user