Adding v3

This commit is contained in:
2026-06-01 00:22:39 -05:00
parent 285bbd7aac
commit fcd6f6742c
2 changed files with 36 additions and 0 deletions
+25
View File
@@ -26,6 +26,12 @@ The new research architecture is available as `v2-32x1m`:
- router z-loss to keep router logits stable
- optional phased expert training so only a subset of experts are trainable/routeable at a time
The early v3 preset is `v3-2l-32x1m`:
- same 32 experts/layer and 1M params/expert as v2
- 2 transformer/MoE layers instead of 1
- intended as the first quality-focused depth increase while keeping the v2 expert shape intact
The default still uses `n_layers=1`, which means the model has one MoE block. If you raise `--n-layers`, each layer gets its own full expert set.
## Install
@@ -67,6 +73,25 @@ verysimplemoe-train \
--block-size 256
```
## Train early v3: 2 layers, 32 experts/layer, 1M params/expert
Recommended GPU command for a 50M-token FineWeb EDU run:
```bash
verysimplemoe-train \
--arch v3-2l-32x1m \
--dataset-name HuggingFaceFW/fineweb-edu \
--dataset-config sample-10BT \
--out-dir checkpoints/verysimplemoe-v3-2l-32x1m-fineweb-edu-50m \
--max-steps 6104 \
--batch-size 8 \
--grad-accum-steps 4 \
--block-size 256 \
--train-experts-per-phase 16 \
--expert-phase-steps 500 \
--save-every 1000
```
Phased expert training means:
- only 16 of the 32 experts are eligible for routing in a given phase
+11
View File
@@ -66,6 +66,17 @@ ARCH_PRESETS: dict[str, dict[str, int | float | str]] = {
"router_noise_std": 0.1,
"router_z_loss_coef": 1e-4,
},
"v3-2l-32x1m": {
"arch": "v3-2l-32x1m",
"n_layers": 2,
"d_model": 500,
"n_heads": 10,
"n_experts": 32,
"active_experts": 4,
"expert_hidden_size": 1000,
"router_noise_std": 0.1,
"router_z_loss_coef": 1e-4,
},
}