main
1B Dense GPT Pretraining Skeleton
This repo contains:
model.py— GPT-style dense decoder-only Transformer (~1.02B params with a 128k vocab)train_tokenizer.py— optional 128k byte-level BPE tokenizer training on FineWeb-EDUtrain.py— DDP pretraining on FineWeb-EDU with gradient accumulation and checkpoint resume/initgenerate.py— quick text generation from a checkpointrequirements.txt
Model defaults
The default architecture targets a custom 128k tokenizer:
- vocab: tokenizer length, expected
128000 - layers:
16 - hidden size:
2048 - query heads:
16 - KV heads:
8(GQA) - SwiGLU intermediate size:
5632 - tied input/output embeddings
This is about 1.02B trainable parameters.
Install
pip install -r requirements.txt
Train tokenizer (optional)
If you do not already have a custom 128k tokenizer:
python train_tokenizer.py --out_dir tokenizers/fwe_128k --vocab_size 128000
Then pass --tokenizer_path tokenizers/fwe_128k to training/generation.
Train on all GPUs
On the training box:
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
--tokenizer_path /path/to/custom-128k-tokenizer \
--out_dir runs/one_b_fineweb_edu \
--total_tokens 10000000000 \
--block_size 4096 \
--micro_batch_size 1 \
--grad_accum_steps 64 \
--precision bf16
The script streams HuggingFaceFW/fineweb-edu with config sample-10BT by default.
Resume same run
Restores model, optimizer, scaler, RNG, step, and token count:
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
--tokenizer_path /path/to/custom-128k-tokenizer \
--out_dir runs/one_b_fineweb_edu \
--resume runs/one_b_fineweb_edu/ckpt_last.pt
For exact streamed data-position resume, keep --num_workers 0.
Start a new run from checkpoint weights
Loads model weights/config only and starts a fresh optimizer/schedule in a new output dir:
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
--tokenizer_path /path/to/custom-128k-tokenizer \
--out_dir runs/continued \
--init_from runs/one_b_fineweb_edu/ckpt_last.pt
Generate
python generate.py \
--checkpoint runs/one_b_fineweb_edu/ckpt_last.pt \
--tokenizer_path /path/to/custom-128k-tokenizer \
--prompt "The purpose of education is" \
--max_new_tokens 128
Languages
Python
100%