91 lines
2.4 KiB
Markdown
91 lines
2.4 KiB
Markdown
# 1B Dense GPT Pretraining Skeleton
|
|
|
|
This repo contains:
|
|
|
|
- `model.py` — GPT-style dense decoder-only Transformer (~1.02B params with a 128k vocab)
|
|
- `train_tokenizer.py` — optional 128k byte-level BPE tokenizer training on FineWeb-EDU
|
|
- `train.py` — DDP pretraining on FineWeb-EDU with gradient accumulation and checkpoint resume/init
|
|
- `generate.py` — quick text generation from a checkpoint
|
|
- `requirements.txt`
|
|
|
|
## Model defaults
|
|
|
|
The default architecture targets a custom 128k tokenizer:
|
|
|
|
- vocab: tokenizer length, expected `128000`
|
|
- layers: `16`
|
|
- hidden size: `2048`
|
|
- query heads: `16`
|
|
- KV heads: `8` (GQA)
|
|
- SwiGLU intermediate size: `5632`
|
|
- tied input/output embeddings
|
|
|
|
This is about **1.02B trainable parameters**.
|
|
|
|
## Install
|
|
|
|
```bash
|
|
pip install -r requirements.txt
|
|
```
|
|
|
|
## Train tokenizer (optional)
|
|
|
|
If you do not already have a custom 128k tokenizer:
|
|
|
|
```bash
|
|
python train_tokenizer.py --out_dir tokenizers/fwe_128k --vocab_size 128000
|
|
```
|
|
|
|
Then pass `--tokenizer_path tokenizers/fwe_128k` to training/generation.
|
|
|
|
## Train on all GPUs
|
|
|
|
On the training box:
|
|
|
|
```bash
|
|
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
|
|
--tokenizer_path /path/to/custom-128k-tokenizer \
|
|
--out_dir runs/one_b_fineweb_edu \
|
|
--total_tokens 10000000000 \
|
|
--block_size 4096 \
|
|
--micro_batch_size 1 \
|
|
--grad_accum_steps 64 \
|
|
--precision bf16
|
|
```
|
|
|
|
The script streams `HuggingFaceFW/fineweb-edu` with config `sample-10BT` by default.
|
|
|
|
## Resume same run
|
|
|
|
Restores model, optimizer, scaler, RNG, step, and token count:
|
|
|
|
```bash
|
|
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
|
|
--tokenizer_path /path/to/custom-128k-tokenizer \
|
|
--out_dir runs/one_b_fineweb_edu \
|
|
--resume runs/one_b_fineweb_edu/ckpt_last.pt
|
|
```
|
|
|
|
For exact streamed data-position resume, keep `--num_workers 0`.
|
|
|
|
## Start a new run from checkpoint weights
|
|
|
|
Loads model weights/config only and starts a fresh optimizer/schedule in a new output dir:
|
|
|
|
```bash
|
|
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
|
|
--tokenizer_path /path/to/custom-128k-tokenizer \
|
|
--out_dir runs/continued \
|
|
--init_from runs/one_b_fineweb_edu/ckpt_last.pt
|
|
```
|
|
|
|
## Generate
|
|
|
|
```bash
|
|
python generate.py \
|
|
--checkpoint runs/one_b_fineweb_edu/ckpt_last.pt \
|
|
--tokenizer_path /path/to/custom-128k-tokenizer \
|
|
--prompt "The purpose of education is" \
|
|
--max_new_tokens 128
|
|
```
|