Update README.md
Remove formatting of tables
This commit is contained in:
@@ -172,7 +172,6 @@ The vLLM Worker is fully compatible with OpenAI's API, and you can use it with a
|
||||
**Python** (similar to Node.js, etc.):
|
||||
1. When initializing the OpenAI Client in your code, change the `api_key` to your RunPod API Key and the `base_url` to your RunPod Serverless Endpoint URL in the following format: `https://api.runpod.ai/v2/<YOUR ENDPOINT ID>/openai/v1`, filling in your deployed endpoint ID.
|
||||
|
||||
|
||||
- Before:
|
||||
```python
|
||||
from openai import OpenAI
|
||||
@@ -254,7 +253,7 @@ When using the chat completion feature of the vLLM Serverless Endpoint Worker, y
|
||||
<summary>Click to expand table</summary>
|
||||
|
||||
| Parameter | Type | Default Value | Description |
|
||||
| ------------------- | -------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
|--------------------------------|----------------------------------|---------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `messages` | Union[str, List[Dict[str, str]]] | | List of messages, where each message is a dictionary with a `role` and `content`. The model's chat template will be applied to the messages automatically, so the model must have one or it should be specified as `CUSTOM_CHAT_TEMPLATE` env var. |
|
||||
| `model` | str | | The model repo that you've deployed on your RunPod Serverless Endpoint. If you are unsure what the name is or are baking the model in, use the guide to get the list of available models in the **Examples: Using your RunPod endpoint with OpenAI** section |
|
||||
| `temperature` | Optional[float] | 0.7 | Float that controls the randomness of the sampling. Lower values make the model more deterministic, while higher values make the model more random. Zero means greedy sampling. |
|
||||
@@ -270,7 +269,7 @@ When using the chat completion feature of the vLLM Serverless Endpoint Worker, y
|
||||
| `user` | Optional[str] | None | Unsupported by vLLM |
|
||||
Additional parameters supported by vLLM:
|
||||
| `best_of` | Optional[int] | None | Number of output sequences that are generated from the prompt. From these `best_of` sequences, the top `n` sequences are returned. `best_of` must be greater than or equal to `n`. This is treated as the beam width when `use_beam_search` is True. By default, `best_of` is set to `n`. |
|
||||
| `-----------------------------` | O-----------------] | ----1 | I---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------. |
|
||||
| `top_k` | Optional[int] | -1 | Integer that controls the number of top tokens to consider. Set to -1 to consider all tokens. |
|
||||
| `ignore_eos` | Optional[bool] | False | Whether to ignore the EOS token and continue generating tokens after the EOS token is generated. |
|
||||
| `use_beam_search` | Optional[bool] | False | Whether to use beam search instead of sampling. |
|
||||
| `stop_token_ids` | Optional[List[int]] | list | List of tokens that stop the generation when they are generated. The returned output will contain the stop tokens unless the stop tokens are special tokens. |
|
||||
@@ -289,7 +288,7 @@ When using the chat completion feature of the vLLM Serverless Endpoint Worker, y
|
||||
<summary>Click to expand table</summary>
|
||||
|
||||
| Parameter | Type | Default Value | Description |
|
||||
| ------------------- | ------------------------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
|--------------------------------|----------------------------------|---------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `model` | str | | The model repo that you've deployed on your RunPod Serverless Endpoint. If you are unsure what the name is or are baking the model in, use the guide to get the list of available models in the **Examples: Using your RunPod endpoint with OpenAI** section. |
|
||||
| `prompt` | Union[List[int], List[List[int]], str, List[str]] | | A string, array of strings, array of tokens, or array of token arrays to be used as the input for the model. |
|
||||
| `suffix` | Optional[str] | None | A string to be appended to the end of the generated text. |
|
||||
@@ -309,7 +308,7 @@ When using the chat completion feature of the vLLM Serverless Endpoint Worker, y
|
||||
| `user` | Optional[str] | None | User identifier for personalizing responses. (Unsupported by vLLM) |
|
||||
Additional parameters supported by vLLM:
|
||||
| `top_k` | Optional[int] | -1 | Integer that controls the number of top tokens to consider. Set to -1 to consider all tokens. |
|
||||
| `-----------------------------` | O-----------------] | F---e | W----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------. |
|
||||
| `ignore_eos` | Optional[bool] | False | Whether to ignore the End Of Sentence token and continue generating tokens after the EOS token is generated. |
|
||||
| `use_beam_search` | Optional[bool] | False | Whether to use beam search instead of sampling for generating outputs. |
|
||||
| `stop_token_ids` | Optional[List[int]] | list | List of tokens that stop the generation when they are generated. The returned output will contain the stop tokens unless the stop tokens are special tokens. |
|
||||
| `skip_special_tokens` | Optional[bool] | True | Whether to skip special tokens in the output. |
|
||||
@@ -409,7 +408,7 @@ print(list_of_models)
|
||||
|
||||
You may either use a `prompt` or a list of `messages` as input. If you use `messages`, the model's chat template will be applied to the messages automatically, so the model must have one. If you use `prompt`, you may optionally apply the model's chat template to the prompt by setting `apply_chat_template` to `true`.
|
||||
| Argument | Type | Default | Description |
|
||||
| -------------------------- | -------------------- | ------------------------------------------ | --------------------------------------------------------------------------------------------------------------- |
|
||||
|-----------------------|----------------------|--------------------|--------------------------------------------------------------------------------------------------------|
|
||||
| `prompt` | str | | Prompt string to generate text based on. |
|
||||
| `messages` | list[dict[str, str]] | | List of messages, which will automatically have the model's chat template applied. Overrides `prompt`. |
|
||||
| `apply_chat_template` | bool | False | Whether to apply the model's chat template to the `prompt`. |
|
||||
@@ -464,7 +463,7 @@ Below are all available sampling parameters that you can specify in the `samplin
|
||||
<summary>Click to expand table</summary>
|
||||
|
||||
| Argument | Type | Default | Description |
|
||||
| ------------------------------- | --------------------------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
|---------------------------------|-----------------------------|---------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `n` | int | 1 | Number of output sequences generated from the prompt. The top `n` sequences are returned. |
|
||||
| `best_of` | Optional[int] | `n` | Number of output sequences generated from the prompt. The top `n` sequences are returned from these `best_of` sequences. Must be ≥ `n`. Treated as beam width in beam search. Default is `n`. |
|
||||
| `presence_penalty` | float | 0.0 | Penalizes new tokens based on their presence in the generated text so far. Values > 0 encourage new tokens, values < 0 encourage repetition. |
|
||||
@@ -483,4 +482,3 @@ Below are all available sampling parameters that you can specify in the `samplin
|
||||
| `max_tokens` | int | 16 | Maximum number of tokens to generate per output sequence. |
|
||||
| `skip_special_tokens` | bool | True | Whether to skip special tokens in the output. |
|
||||
| `spaces_between_special_tokens` | bool | True | Whether to add spaces between special tokens in the output. |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user