Files
longhaul.cpp/docs/longhaul.md
2026-08-05 21:13:25 -05:00

5.6 KiB

Longhaul MoE loading

Longhaul mode keeps the non-expert model weights resident and loads routed MoE expert slices from the GGUF file as they are needed. The expert cache has a fixed budget and uses least-recently-used replacement independently for each layer.

This mode is intended for running a model whose full expert weights do not fit in unified memory. It trades throughput and latency for a smaller resident model allocation.

Usage

llama-cli \
    --model /path/to/model.gguf \
    --longhaul \
    --longhaul-cache 2 \
    --n-gpu-layers 0

--longhaul-cache is the expert cache budget in GiB. It is required when --longhaul is used. The same options are accepted by llama-server.

The command above runs every repeating layer on CPU and is supported on macOS, Linux, and Windows. On macOS, the existing all-Metal mode remains available by using --n-gpu-layers 99 instead. Longhaul does not silently change device placement: an unsupported GPU or mixed CPU/GPU repeating-layer placement fails with guidance to use --n-gpu-layers 0.

Longhaul may reduce --ubatch-size so that every expert selected by one graph segment can be present in the cache at the same time. The effective value is logged during context creation. If one token selects more experts than the cache has slots, the routed MoE computation is split into multiple stages and the partial results are summed. This permits smaller caches at the cost of additional graph work.

The normal startup warmup is skipped automatically in longhaul mode. Routed expert weights are not read until the first real decode.

Current scope

Longhaul currently requires:

  • all repeating layers on CPU, or all repeating layers on Metal
  • Qwen3.5 MoE, DeepSeek V4, Laguna, or Inkling architecture
  • text generation without embeddings or LoRA adapters

CPU mode is available wherever the CPU backend is supported. Metal mode requires macOS. Longhaul does not restrict the GGUF quantization type; individual tensor types must still be supported by the selected compute backend.

MTP/speculative decoding, tensor validation during loading, vocabulary-only loading, and mixed CPU/GPU repeating-layer placement are not supported. Both single-file and split GGUF models are supported; routed expert tensors are read from the shard that owns each tensor.

The cache budget covers the compact routed-expert tensors. It does not include dense weights, attention weights, the KV cache, graph allocations, or temporary staging buffers. CPU expert caches use the standard directly writable tensor layout so individual slots can be replaced; optimized CPU buffer selection continues to apply to all non-streamed weights.

File-cache control is best effort and is outside the explicit cache budget. macOS disables caching for streamed files where supported, and Linux advises the kernel to discard completed reads. Windows may retain file data in its system cache. Windows reads from one GGUF shard are serialized to preserve positional read correctness, while POSIX systems retain concurrent pread operations.

This implementation synchronizes at each routed MoE layer to discover the selected experts, populate missing cache slots, and continue execution with cache-local expert IDs. Storage speed and expert reuse therefore have a large effect on generation speed.

Expert IDs are planned as a batch at each synchronization point. Experts already needed by that batch are protected from eviction, duplicate IDs are loaded only once, and independent expert slices are read concurrently where the platform supports positional reads. CPU and shared Metal buffers are populated directly; private Metal buffers use a staged fallback.

DeepSeek V4 keeps its shared expert, router weights, learned routing bias, and hash-routing tables resident. Longhaul streams the routed gate, up, and down expert banks. DeepSeek V4 Flash selects six routed experts per token; caches with fewer than six slots use the existing staged MoE path. The startup log is the authoritative source for slot capacity because the bytes per slot depend on the GGUF tensor types and shard layout. MTP tensors remain outside the current scope.

Inkling keeps its two shared experts resident and streams only the routed 256-expert banks. Inkling-Small selects six routed experts per token. In the seven-shard Q8_0 model, one cache slot across all 40 MoE layers uses about 0.996 GiB, so a 6 GiB cache provides six slots and executes each MoE layer in one stage. The non-routed tensors, including both shared experts, require about 6.02 GiB in addition to the expert cache:

llama-cli \
    --model /path/to/Inkling-Small-Q8_0-00001-of-00007.gguf \
    --longhaul \
    --longhaul-cache 6 \
    --n-gpu-layers 99 \
    --flash-attn off \
    --jinja

Budgets from 1 through 5 GiB are supported by splitting the six selected experts into multiple stages. Add --n-gpu-layers 0 to run entirely on CPU. Inkling currently forces flash attention off because its relative-position bias is evaluated by the unfused attention path. This support covers text and tool-call generation from the language-model GGUF; multimodal projection files are not loaded by longhaul. Use --jinja with llama-cli, llama-completion, or llama-server so the model's embedded TML chat template is accepted.

Benchmarking prompt processing

llama-bench accepts the longhaul load mode and cache budget:

llama-bench \
    --model /path/to/model.gguf \
    --load-mode longhaul \
    --longhaul-cache 2 \
    --n-gpu-layers 0 \
    --n-prompt 2048 \
    --n-gen 0

Use --no-warmup --repetitions 1 in separate processes to measure a cold expert cache. Leave warmup enabled to measure steady-state cache reuse.