feat: add Ling 3.0 LongHaul support

This commit is contained in:
Owen Qwen
2026-08-09 09:57:32 -05:00
parent b26b176575
commit d74649d438
20 changed files with 820 additions and 19 deletions
+10 -1
View File
@@ -31,7 +31,7 @@ The normal startup warmup is skipped automatically in longhaul mode. Routed expe
Longhaul currently requires:
- all repeating layers on CPU, or all repeating layers on Metal
- Qwen3.5 MoE, Laguna, or Inkling architecture
- Qwen3.5 MoE, Laguna, Inkling, or Ling 3.0 (`bailingmoe3`) architecture
- text generation without embeddings or LoRA adapters
CPU mode is available wherever the CPU backend is supported. Metal mode requires
@@ -63,6 +63,15 @@ once, and independent expert slices are read concurrently where the platform
supports positional reads. CPU and shared Metal buffers are populated directly;
private Metal buffers use a staged fallback.
Ling 3.0 support covers its hybrid KDA/MLA transformer and sigmoid-routed MoE.
For Ling-3.0-flash, the router, score-correction bias, and shared expert remain
resident while the 512 routed experts are streamed. Each token selects eight
routed experts, so cache budgets with fewer than eight slots execute an MoE
layer in multiple stages. AtomicChat's Ling GGUFs can be loaded directly, and
the HF converter also recognizes `BailingMoeV3ForCausalLM` checkpoints. The
converter omits the auxiliary MTP tensor block because Longhaul does not support
MTP tensors.
Inkling keeps its two shared experts resident and streams only the routed
256-expert banks. Inkling-Small selects six routed experts per token. In the
seven-shard Q8_0 model, one cache slot across all 40 MoE layers uses about