feat: add Ling 3.0 LongHaul support
This commit is contained in:
+10
-1
@@ -31,7 +31,7 @@ The normal startup warmup is skipped automatically in longhaul mode. Routed expe
|
||||
Longhaul currently requires:
|
||||
|
||||
- all repeating layers on CPU, or all repeating layers on Metal
|
||||
- Qwen3.5 MoE, Laguna, or Inkling architecture
|
||||
- Qwen3.5 MoE, Laguna, Inkling, or Ling 3.0 (`bailingmoe3`) architecture
|
||||
- text generation without embeddings or LoRA adapters
|
||||
|
||||
CPU mode is available wherever the CPU backend is supported. Metal mode requires
|
||||
@@ -63,6 +63,15 @@ once, and independent expert slices are read concurrently where the platform
|
||||
supports positional reads. CPU and shared Metal buffers are populated directly;
|
||||
private Metal buffers use a staged fallback.
|
||||
|
||||
Ling 3.0 support covers its hybrid KDA/MLA transformer and sigmoid-routed MoE.
|
||||
For Ling-3.0-flash, the router, score-correction bias, and shared expert remain
|
||||
resident while the 512 routed experts are streamed. Each token selects eight
|
||||
routed experts, so cache budgets with fewer than eight slots execute an MoE
|
||||
layer in multiple stages. AtomicChat's Ling GGUFs can be loaded directly, and
|
||||
the HF converter also recognizes `BailingMoeV3ForCausalLM` checkpoints. The
|
||||
converter omits the auxiliary MTP tensor block because Longhaul does not support
|
||||
MTP tensors.
|
||||
|
||||
Inkling keeps its two shared experts resident and streams only the routed
|
||||
256-expert banks. Inkling-Small selects six routed experts per token. In the
|
||||
seven-shard Q8_0 model, one cache slot across all 40 MoE layers uses about
|
||||
|
||||
Reference in New Issue
Block a user