From e32626ca9d6b719ed7439c78b19088d4dd56a1f8 Mon Sep 17 00:00:00 2001 From: Marut Pandya Date: Mon, 5 Aug 2024 14:19:38 -0700 Subject: [PATCH] Update README.md --- README.md | 9 +-------- 1 file changed, 1 insertion(+), 8 deletions(-) diff --git a/README.md b/README.md index 1f35f13..d3c611c 100644 --- a/README.md +++ b/README.md @@ -97,7 +97,7 @@ Below is a summary of the available RunPod Worker images, categorized by image s | `TOKENIZER_MODE` | 'auto' | ['auto', 'slow'] | The tokenizer mode. | | `TRUST_REMOTE_CODE` | False | `bool` | Trust remote code from Hugging Face. | | `DOWNLOAD_DIR` | None | `str` | Directory to download and load the weights. | -| `LOAD_FORMAT` | 'auto' | ['auto', 'pt', 'safetensors', 'npcache', 'dummy', 'tensorizer', 'bitsandbytes'] | The format of the model weights to load. | +| `LOAD_FORMAT` | 'auto' | ['auto', 'pt', 'safetensors', 'npcache', 'dummy', 'bitsandbytes'] | The format of the model weights to load. | | `DTYPE` | 'auto' | ['auto', 'half', 'float16', 'bfloat16', 'float', 'float32'] | Data type for model weights and activations. | | `KV_CACHE_DTYPE` | 'auto' | ['auto', 'fp8', 'fp8_e5m2', 'fp8_e4m3'] | Data type for KV cache storage. | | `QUANTIZATION_PARAM_PATH` | None | `str` | Path to the JSON file containing the KV cache scaling factors. | @@ -109,14 +109,11 @@ Below is a summary of the available RunPod Worker images, categorized by image s | `TENSOR_PARALLEL_SIZE` | 1 | `int` | Number of tensor parallel replicas. | | `MAX_PARALLEL_LOADING_WORKERS` | None | `int` | Load model sequentially in multiple batches. | | `RAY_WORKERS_USE_NSIGHT` | False | `bool` | If specified, use nsight to profile Ray workers. | -| `BLOCK_SIZE` | 16 | [8, 16, 32] | Token block size for contiguous chunks of tokens. | | `ENABLE_PREFIX_CACHING` | False | `bool` | Enables automatic prefix caching. | | `DISABLE_SLIDING_WINDOW` | False | `bool` | Disables sliding window, capping to sliding window size. | | `USE_V2_BLOCK_MANAGER` | False | `bool` | Use BlockSpaceMangerV2. | | `NUM_LOOKAHEAD_SLOTS` | 0 | `int` | Experimental scheduling config necessary for speculative decoding. | | `SEED` | 0 | `int` | Random seed for operations. | -| `SWAP_SPACE` | 4 | `int` | CPU swap space size (GiB) per GPU. | -| `GPU_MEMORY_UTILIZATION` | 0.90 | `float` | The fraction of GPU memory to be used for the model executor. | | `NUM_GPU_BLOCKS_OVERRIDE` | None | `int` | If specified, ignore GPU profiling result and use this number of GPU blocks. | | `MAX_NUM_BATCHED_TOKENS` | None | `int` | Maximum number of batched tokens per iteration. | | `MAX_NUM_SEQS` | 256 | `int` | Maximum number of sequences per iteration. | @@ -125,10 +122,6 @@ Below is a summary of the available RunPod Worker images, categorized by image s | `QUANTIZATION` | None | [*QUANTIZATION_METHODS, None] | Method used to quantize the weights. | | `ROPE_SCALING` | None | `dict` | RoPE scaling configuration in JSON format. | | `ROPE_THETA` | None | `float` | RoPE theta. Use with rope_scaling. | -| `ENFORCE_EAGER` | False | `bool` | Always use eager-mode PyTorch. | -| `MAX_CONTEXT_LEN_TO_CAPTURE` | None | `int` | Maximum context length covered by CUDA graphs. | -| `MAX_SEQ_LEN_TO_CAPTURE` | 8192 | `int` | Maximum sequence length covered by CUDA graphs. | -| `DISABLE_CUSTOM_ALL_REDUCE` | False | `bool` | See ParallelConfig. | | `TOKENIZER_POOL_SIZE` | 0 | `int` | Size of tokenizer pool to use for asynchronous tokenization. | | `TOKENIZER_POOL_TYPE` | 'ray' | `str` | Type of tokenizer pool to use for asynchronous tokenization. | | `TOKENIZER_POOL_EXTRA_CONFIG` | None | `dict` | Extra config for tokenizer pool. |