Inital commit

This commit is contained in:
2026-06-06 17:06:08 -05:00
commit 09ccc0aa32
6 changed files with 944 additions and 0 deletions
+90
View File
@@ -0,0 +1,90 @@
# 1B Dense GPT Pretraining Skeleton
This repo contains:
- `model.py` — GPT-style dense decoder-only Transformer (~1.02B params with a 128k vocab)
- `train_tokenizer.py` — optional 128k byte-level BPE tokenizer training on FineWeb-EDU
- `train.py` — DDP pretraining on FineWeb-EDU with gradient accumulation and checkpoint resume/init
- `generate.py` — quick text generation from a checkpoint
- `requirements.txt`
## Model defaults
The default architecture targets a custom 128k tokenizer:
- vocab: tokenizer length, expected `128000`
- layers: `16`
- hidden size: `2048`
- query heads: `16`
- KV heads: `8` (GQA)
- SwiGLU intermediate size: `5632`
- tied input/output embeddings
This is about **1.02B trainable parameters**.
## Install
```bash
pip install -r requirements.txt
```
## Train tokenizer (optional)
If you do not already have a custom 128k tokenizer:
```bash
python train_tokenizer.py --out_dir tokenizers/fwe_128k --vocab_size 128000
```
Then pass `--tokenizer_path tokenizers/fwe_128k` to training/generation.
## Train on all GPUs
On the training box:
```bash
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
--tokenizer_path /path/to/custom-128k-tokenizer \
--out_dir runs/one_b_fineweb_edu \
--total_tokens 10000000000 \
--block_size 4096 \
--micro_batch_size 1 \
--grad_accum_steps 64 \
--precision bf16
```
The script streams `HuggingFaceFW/fineweb-edu` with config `sample-10BT` by default.
## Resume same run
Restores model, optimizer, scaler, RNG, step, and token count:
```bash
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
--tokenizer_path /path/to/custom-128k-tokenizer \
--out_dir runs/one_b_fineweb_edu \
--resume runs/one_b_fineweb_edu/ckpt_last.pt
```
For exact streamed data-position resume, keep `--num_workers 0`.
## Start a new run from checkpoint weights
Loads model weights/config only and starts a fresh optimizer/schedule in a new output dir:
```bash
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
--tokenizer_path /path/to/custom-128k-tokenizer \
--out_dir runs/continued \
--init_from runs/one_b_fineweb_edu/ckpt_last.pt
```
## Generate
```bash
python generate.py \
--checkpoint runs/one_b_fineweb_edu/ckpt_last.pt \
--tokenizer_path /path/to/custom-128k-tokenizer \
--prompt "The purpose of education is" \
--max_new_tokens 128
```