A lightweight, highly optimized framework for training an autoregressive language model (GPT) from absolute scratch on a local GPU. This repository is specifically tuned for training on the HuggingFace TinyStories dataset using Byte-Pair Encoding (BPE), Automatic Mixed Precision (AMP bfloat16), and PyTorch TensorFloat32 (TF32) for maximum speed. Based on the work of LLM From Scratch.
The easiest way to run the entire framework end-to-end is by using the provided PowerShell batch script.
.\run_experiment.ps1 -DataPath "data\TinyStoriesV2-GPT4-train.txt" -MaxSteps 40000 -BatchSize 64 -NLayer 8 -NHead 8 -NEmbd 512 -BlockSize 256 -MaxLR 7e-4Warning
Because run_experiment.ps1 is a PowerShell script, you must use PowerShell parameter syntax (a single hyphen, e.g., -MaxLR). Do not use Python argparse syntax (double hyphens, e.g., --max_lr), or the script will crash.
- Trains the Model: Executes
train.pywith your specified hyperparameters. - Generates the Loss Chart: Automatically runs
plottraining.pyto create a smoothed, high-resolution.pnggraph of your training vs validation loss. - Generates Text Samples: Automatically runs
generate.py10 separate times using the newly trained model and saves all the text outputs into a.txtfile named after your model (e.g.,checkpoint_8L-8H-512D_..._samples.txt).
The main training loop. It streams tokenized datasets directly from the hard drive into VRAM to prevent Out-Of-Memory (OOM) RAM crashes.
Arguments:
(Note: Use --flag when calling train.py directly, but use -Flag when calling via run_experiment.ps1)
--data_path/-DataPath(default:data/TinyStories-valid.txt): The path to your raw text file. (The tokenizer will automatically cache it as a.ptfile next to it).--max_steps/-MaxSteps(default:40000): Total number of training iterations.--batch_size/-BatchSize(default:64): Number of sequences processed in parallel. Set this as high as your GPU VRAM allows.--n_layer/-NLayer(default:8): Number of transformer blocks (Depth).--n_head/-NHead(default:8): Number of attention heads.--n_embd/-NEmbd(default:512): Dimensionality of the embeddings (Width). Must be perfectly divisible byn_head!--block_size/-BlockSize(default:256): Maximum context window size.--max_lr/-MaxLR(default:7e-4): Maximum learning rate for the AdamW optimizer.
Example usage:
python train.py --data_path data/TinyStories.txt --max_steps 5000 --batch_size 32 --max_lr 1e-3Generates a loss_curve_*.png chart showing raw training loss, smoothed training loss, and smoothed validation loss.
It reads directly from loss_log.json and embeds all hyperparameter metadata (Params, Batch Size, Block Size, Time, etc.) directly into the chart's subtitle. It runs automatically in the background without hanging your terminal.
Loads a saved model checkpoint (.pt) and its corresponding tokenizer, and generates autoregressive text.
Arguments:
checkpoint(Required): Path to your saved.ptfile.--prompt(default:"Once upon a time"): The starting seed string.--max_new_tokens(default:200): Maximum number of words to generate.--temperature(default:0.8): Sampling temperature (0.0 = greedy/deterministic, >1.0 = chaotic/random).--top_k(default:40): Restricts the model to only sample from the topkmost likely next tokens.--seed(default:None): Sets a manual PyTorch seed for perfectly reproducible text generation.
Example usage:
python generate.py checkpoint_8L-8H-512D_final.pt --prompt "Lily found a box" --temperature 0.7Handles training a custom ByteLevel BPE Tokenizer from scratch and provides memory-efficient chunked iterators to feed PyTorch tensors directly to the GPU. You rarely need to run this manually as train.py handles it automatically.
TinyStories:https://huggingface.co/datasets/roneneldan/TinyStories
TinyStoriesV2: https://huggingface.co/datasets/roneneldan/TinyStoriesV2