2026-06-06 17:06:08 -05:00
2026-06-06 17:06:08 -05:00
2026-06-06 17:06:08 -05:00
2026-06-06 17:06:08 -05:00
2026-06-06 17:06:08 -05:00
2026-06-06 17:06:08 -05:00
2026-06-06 17:06:08 -05:00

1B Dense GPT Pretraining Skeleton

This repo contains:

  • model.py — GPT-style dense decoder-only Transformer (~1.02B params with a 128k vocab)
  • train_tokenizer.py — optional 128k byte-level BPE tokenizer training on FineWeb-EDU
  • train.py — DDP pretraining on FineWeb-EDU with gradient accumulation and checkpoint resume/init
  • generate.py — quick text generation from a checkpoint
  • requirements.txt

Model defaults

The default architecture targets a custom 128k tokenizer:

  • vocab: tokenizer length, expected 128000
  • layers: 16
  • hidden size: 2048
  • query heads: 16
  • KV heads: 8 (GQA)
  • SwiGLU intermediate size: 5632
  • tied input/output embeddings

This is about 1.02B trainable parameters.

Install

pip install -r requirements.txt

Train tokenizer (optional)

If you do not already have a custom 128k tokenizer:

python train_tokenizer.py --out_dir tokenizers/fwe_128k --vocab_size 128000

Then pass --tokenizer_path tokenizers/fwe_128k to training/generation.

Train on all GPUs

On the training box:

torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
  --tokenizer_path /path/to/custom-128k-tokenizer \
  --out_dir runs/one_b_fineweb_edu \
  --total_tokens 10000000000 \
  --block_size 4096 \
  --micro_batch_size 1 \
  --grad_accum_steps 64 \
  --precision bf16

The script streams HuggingFaceFW/fineweb-edu with config sample-10BT by default.

Resume same run

Restores model, optimizer, scaler, RNG, step, and token count:

torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
  --tokenizer_path /path/to/custom-128k-tokenizer \
  --out_dir runs/one_b_fineweb_edu \
  --resume runs/one_b_fineweb_edu/ckpt_last.pt

For exact streamed data-position resume, keep --num_workers 0.

Start a new run from checkpoint weights

Loads model weights/config only and starts a fresh optimizer/schedule in a new output dir:

torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
  --tokenizer_path /path/to/custom-128k-tokenizer \
  --out_dir runs/continued \
  --init_from runs/one_b_fineweb_edu/ckpt_last.pt

Generate

python generate.py \
  --checkpoint runs/one_b_fineweb_edu/ckpt_last.pt \
  --tokenizer_path /path/to/custom-128k-tokenizer \
  --prompt "The purpose of education is" \
  --max_new_tokens 128
S
Description
No description provided
Readme
40 KiB
Languages
Python 100%