# 1B Dense GPT Pretraining Skeleton This repo contains: - `model.py` — GPT-style dense decoder-only Transformer (~1.02B params with a 128k vocab) - `train_tokenizer.py` — optional 128k byte-level BPE tokenizer training on FineWeb-EDU - `train.py` — DDP pretraining on FineWeb-EDU with gradient accumulation and checkpoint resume/init - `generate.py` — quick text generation from a checkpoint - `requirements.txt` ## Model defaults The default architecture targets a custom 128k tokenizer: - vocab: tokenizer length, expected `128000` - layers: `16` - hidden size: `2048` - query heads: `16` - KV heads: `8` (GQA) - SwiGLU intermediate size: `5632` - tied input/output embeddings This is about **1.02B trainable parameters**. ## Install ```bash pip install -r requirements.txt ``` ## Train tokenizer (optional) If you do not already have a custom 128k tokenizer: ```bash python train_tokenizer.py --out_dir tokenizers/fwe_128k --vocab_size 128000 ``` Then pass `--tokenizer_path tokenizers/fwe_128k` to training/generation. ## Train on all GPUs On the training box: ```bash torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \ --tokenizer_path /path/to/custom-128k-tokenizer \ --out_dir runs/one_b_fineweb_edu \ --total_tokens 10000000000 \ --block_size 4096 \ --micro_batch_size 1 \ --grad_accum_steps 64 \ --precision bf16 ``` The script streams `HuggingFaceFW/fineweb-edu` with config `sample-10BT` by default. ## Resume same run Restores model, optimizer, scaler, RNG, step, and token count: ```bash torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \ --tokenizer_path /path/to/custom-128k-tokenizer \ --out_dir runs/one_b_fineweb_edu \ --resume runs/one_b_fineweb_edu/ckpt_last.pt ``` For exact streamed data-position resume, keep `--num_workers 0`. ## Start a new run from checkpoint weights Loads model weights/config only and starts a fresh optimizer/schedule in a new output dir: ```bash torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \ --tokenizer_path /path/to/custom-128k-tokenizer \ --out_dir runs/continued \ --init_from runs/one_b_fineweb_edu/ckpt_last.pt ``` ## Generate ```bash python generate.py \ --checkpoint runs/one_b_fineweb_edu/ckpt_last.pt \ --tokenizer_path /path/to/custom-128k-tokenizer \ --prompt "The purpose of education is" \ --max_new_tokens 128 ```