Inital commit
This commit is contained in:
@@ -0,0 +1,90 @@
|
||||
# 1B Dense GPT Pretraining Skeleton
|
||||
|
||||
This repo contains:
|
||||
|
||||
- `model.py` — GPT-style dense decoder-only Transformer (~1.02B params with a 128k vocab)
|
||||
- `train_tokenizer.py` — optional 128k byte-level BPE tokenizer training on FineWeb-EDU
|
||||
- `train.py` — DDP pretraining on FineWeb-EDU with gradient accumulation and checkpoint resume/init
|
||||
- `generate.py` — quick text generation from a checkpoint
|
||||
- `requirements.txt`
|
||||
|
||||
## Model defaults
|
||||
|
||||
The default architecture targets a custom 128k tokenizer:
|
||||
|
||||
- vocab: tokenizer length, expected `128000`
|
||||
- layers: `16`
|
||||
- hidden size: `2048`
|
||||
- query heads: `16`
|
||||
- KV heads: `8` (GQA)
|
||||
- SwiGLU intermediate size: `5632`
|
||||
- tied input/output embeddings
|
||||
|
||||
This is about **1.02B trainable parameters**.
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
## Train tokenizer (optional)
|
||||
|
||||
If you do not already have a custom 128k tokenizer:
|
||||
|
||||
```bash
|
||||
python train_tokenizer.py --out_dir tokenizers/fwe_128k --vocab_size 128000
|
||||
```
|
||||
|
||||
Then pass `--tokenizer_path tokenizers/fwe_128k` to training/generation.
|
||||
|
||||
## Train on all GPUs
|
||||
|
||||
On the training box:
|
||||
|
||||
```bash
|
||||
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
|
||||
--tokenizer_path /path/to/custom-128k-tokenizer \
|
||||
--out_dir runs/one_b_fineweb_edu \
|
||||
--total_tokens 10000000000 \
|
||||
--block_size 4096 \
|
||||
--micro_batch_size 1 \
|
||||
--grad_accum_steps 64 \
|
||||
--precision bf16
|
||||
```
|
||||
|
||||
The script streams `HuggingFaceFW/fineweb-edu` with config `sample-10BT` by default.
|
||||
|
||||
## Resume same run
|
||||
|
||||
Restores model, optimizer, scaler, RNG, step, and token count:
|
||||
|
||||
```bash
|
||||
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
|
||||
--tokenizer_path /path/to/custom-128k-tokenizer \
|
||||
--out_dir runs/one_b_fineweb_edu \
|
||||
--resume runs/one_b_fineweb_edu/ckpt_last.pt
|
||||
```
|
||||
|
||||
For exact streamed data-position resume, keep `--num_workers 0`.
|
||||
|
||||
## Start a new run from checkpoint weights
|
||||
|
||||
Loads model weights/config only and starts a fresh optimizer/schedule in a new output dir:
|
||||
|
||||
```bash
|
||||
torchrun --standalone --nproc_per_node=$(nvidia-smi -L | wc -l) train.py \
|
||||
--tokenizer_path /path/to/custom-128k-tokenizer \
|
||||
--out_dir runs/continued \
|
||||
--init_from runs/one_b_fineweb_edu/ckpt_last.pt
|
||||
```
|
||||
|
||||
## Generate
|
||||
|
||||
```bash
|
||||
python generate.py \
|
||||
--checkpoint runs/one_b_fineweb_edu/ckpt_last.pt \
|
||||
--tokenizer_path /path/to/custom-128k-tokenizer \
|
||||
--prompt "The purpose of education is" \
|
||||
--max_new_tokens 128
|
||||
```
|
||||
Reference in New Issue
Block a user