Add RPC longhaul paging
This commit is contained in:
+37
-2
@@ -16,6 +16,41 @@ llama-cli \
|
||||
|
||||
`--longhaul-cache` is the expert cache budget in GiB. It is required when `--longhaul` is used. The same options are accepted by `llama-server`.
|
||||
|
||||
## RPC mode
|
||||
|
||||
RPC longhaul keeps the paging loop on the accelerator host so expert payloads do
|
||||
not cross the network during generation. Build both hosts with `-DGGML_RPC=ON`.
|
||||
Place the same GGUF shard files on both hosts, preserving their basenames, then
|
||||
start the Metal host with the directory containing those shards:
|
||||
|
||||
```sh
|
||||
ggml-rpc-server \
|
||||
--device MTL0 \
|
||||
--longhaul-root /path/to/model-directory \
|
||||
--host 192.168.1.20
|
||||
```
|
||||
|
||||
Start the client normally:
|
||||
|
||||
```sh
|
||||
llama-server \
|
||||
--model /local/path/model.gguf \
|
||||
--rpc 192.168.1.20:50052 \
|
||||
--longhaul \
|
||||
--longhaul-cache 2 \
|
||||
--n-gpu-layers 99
|
||||
```
|
||||
|
||||
RPC longhaul currently requires all repeating layers on one remote Metal device.
|
||||
The server validates each shard basename, size, GGUF metadata layout digest, and
|
||||
every registered expert range before decoding. An older server, a server without
|
||||
`--longhaul-root`, or a shard mismatch is a fatal model-load error; the client
|
||||
does not fall back to sending expert slices over RPC.
|
||||
|
||||
`--longhaul-root` grants connected RPC clients read access to registered byte
|
||||
ranges in files in that directory. The RPC server remains experimental and
|
||||
insecure and must not be exposed to an untrusted network.
|
||||
|
||||
Longhaul may reduce `--ubatch-size` so that every expert selected by one graph segment can be present in the cache at the same time. The effective value is logged during context creation. If one token selects more experts than the cache has slots, the routed MoE computation is split into multiple stages and the partial results are summed. This permits smaller caches at the cost of additional graph work.
|
||||
|
||||
The normal startup warmup is skipped automatically in longhaul mode. Routed expert weights are not read until the first real decode.
|
||||
@@ -24,9 +59,9 @@ The normal startup warmup is skipped automatically in longhaul mode. Routed expe
|
||||
|
||||
Longhaul currently requires:
|
||||
|
||||
- macOS with the Metal backend
|
||||
- the Metal backend, either local on macOS or on one RPC host
|
||||
- Qwen3.5 MoE or Laguna architecture
|
||||
- all repeating model layers assigned to Metal
|
||||
- all repeating model layers assigned to local Metal or one RPC Metal device
|
||||
- text generation without embeddings or LoRA adapters
|
||||
|
||||
Longhaul does not restrict the GGUF quantization type. Individual tensor types must still be supported by Metal.
|
||||
|
||||
Reference in New Issue
Block a user