Adding v3
This commit is contained in:
@@ -26,6 +26,12 @@ The new research architecture is available as `v2-32x1m`:
|
|||||||
- router z-loss to keep router logits stable
|
- router z-loss to keep router logits stable
|
||||||
- optional phased expert training so only a subset of experts are trainable/routeable at a time
|
- optional phased expert training so only a subset of experts are trainable/routeable at a time
|
||||||
|
|
||||||
|
The early v3 preset is `v3-2l-32x1m`:
|
||||||
|
|
||||||
|
- same 32 experts/layer and 1M params/expert as v2
|
||||||
|
- 2 transformer/MoE layers instead of 1
|
||||||
|
- intended as the first quality-focused depth increase while keeping the v2 expert shape intact
|
||||||
|
|
||||||
The default still uses `n_layers=1`, which means the model has one MoE block. If you raise `--n-layers`, each layer gets its own full expert set.
|
The default still uses `n_layers=1`, which means the model has one MoE block. If you raise `--n-layers`, each layer gets its own full expert set.
|
||||||
|
|
||||||
## Install
|
## Install
|
||||||
@@ -67,6 +73,25 @@ verysimplemoe-train \
|
|||||||
--block-size 256
|
--block-size 256
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Train early v3: 2 layers, 32 experts/layer, 1M params/expert
|
||||||
|
|
||||||
|
Recommended GPU command for a 50M-token FineWeb EDU run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
verysimplemoe-train \
|
||||||
|
--arch v3-2l-32x1m \
|
||||||
|
--dataset-name HuggingFaceFW/fineweb-edu \
|
||||||
|
--dataset-config sample-10BT \
|
||||||
|
--out-dir checkpoints/verysimplemoe-v3-2l-32x1m-fineweb-edu-50m \
|
||||||
|
--max-steps 6104 \
|
||||||
|
--batch-size 8 \
|
||||||
|
--grad-accum-steps 4 \
|
||||||
|
--block-size 256 \
|
||||||
|
--train-experts-per-phase 16 \
|
||||||
|
--expert-phase-steps 500 \
|
||||||
|
--save-every 1000
|
||||||
|
```
|
||||||
|
|
||||||
Phased expert training means:
|
Phased expert training means:
|
||||||
|
|
||||||
- only 16 of the 32 experts are eligible for routing in a given phase
|
- only 16 of the 32 experts are eligible for routing in a given phase
|
||||||
|
|||||||
@@ -66,6 +66,17 @@ ARCH_PRESETS: dict[str, dict[str, int | float | str]] = {
|
|||||||
"router_noise_std": 0.1,
|
"router_noise_std": 0.1,
|
||||||
"router_z_loss_coef": 1e-4,
|
"router_z_loss_coef": 1e-4,
|
||||||
},
|
},
|
||||||
|
"v3-2l-32x1m": {
|
||||||
|
"arch": "v3-2l-32x1m",
|
||||||
|
"n_layers": 2,
|
||||||
|
"d_model": 500,
|
||||||
|
"n_heads": 10,
|
||||||
|
"n_experts": 32,
|
||||||
|
"active_experts": 4,
|
||||||
|
"expert_hidden_size": 1000,
|
||||||
|
"router_noise_std": 0.1,
|
||||||
|
"router_z_loss_coef": 1e-4,
|
||||||
|
},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user