Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LLM From Scratch: TinyStories Edition

A lightweight, highly optimized framework for training an autoregressive language model (GPT) from absolute scratch on a local GPU. This repository is specifically tuned for training on the HuggingFace TinyStories dataset using Byte-Pair Encoding (BPE), Automatic Mixed Precision (AMP bfloat16), and PyTorch TensorFloat32 (TF32) for maximum speed. Based on the work of LLM From Scratch.

Quickstart: The Experiment Pipeline

The easiest way to run the entire framework end-to-end is by using the provided PowerShell batch script.

.\run_experiment.ps1 -DataPath "data\TinyStoriesV2-GPT4-train.txt" -MaxSteps 40000 -BatchSize 64 -NLayer 8 -NHead 8 -NEmbd 512 -BlockSize 256 -MaxLR 7e-4

Warning

Because run_experiment.ps1 is a PowerShell script, you must use PowerShell parameter syntax (a single hyphen, e.g., -MaxLR). Do not use Python argparse syntax (double hyphens, e.g., --max_lr), or the script will crash.

What run_experiment.ps1 does:

  1. Trains the Model: Executes train.py with your specified hyperparameters.
  2. Generates the Loss Chart: Automatically runs plottraining.py to create a smoothed, high-resolution .png graph of your training vs validation loss.
  3. Generates Text Samples: Automatically runs generate.py 10 separate times using the newly trained model and saves all the text outputs into a .txt file named after your model (e.g., checkpoint_8L-8H-512D_..._samples.txt).

Core Scripts & Parameters

1. train.py

The main training loop. It streams tokenized datasets directly from the hard drive into VRAM to prevent Out-Of-Memory (OOM) RAM crashes.

Arguments: (Note: Use --flag when calling train.py directly, but use -Flag when calling via run_experiment.ps1)

  • --data_path / -DataPath (default: data/TinyStories-valid.txt): The path to your raw text file. (The tokenizer will automatically cache it as a .pt file next to it).
  • --max_steps / -MaxSteps (default: 40000): Total number of training iterations.
  • --batch_size / -BatchSize (default: 64): Number of sequences processed in parallel. Set this as high as your GPU VRAM allows.
  • --n_layer / -NLayer (default: 8): Number of transformer blocks (Depth).
  • --n_head / -NHead (default: 8): Number of attention heads.
  • --n_embd / -NEmbd (default: 512): Dimensionality of the embeddings (Width). Must be perfectly divisible by n_head!
  • --block_size / -BlockSize (default: 256): Maximum context window size.
  • --max_lr / -MaxLR (default: 7e-4): Maximum learning rate for the AdamW optimizer.

Example usage:

python train.py --data_path data/TinyStories.txt --max_steps 5000 --batch_size 32 --max_lr 1e-3

2. plottraining.py

Generates a loss_curve_*.png chart showing raw training loss, smoothed training loss, and smoothed validation loss.

It reads directly from loss_log.json and embeds all hyperparameter metadata (Params, Batch Size, Block Size, Time, etc.) directly into the chart's subtitle. It runs automatically in the background without hanging your terminal.

3. generate.py

Loads a saved model checkpoint (.pt) and its corresponding tokenizer, and generates autoregressive text.

Arguments:

  • checkpoint (Required): Path to your saved .pt file.
  • --prompt (default: "Once upon a time"): The starting seed string.
  • --max_new_tokens (default: 200): Maximum number of words to generate.
  • --temperature (default: 0.8): Sampling temperature (0.0 = greedy/deterministic, >1.0 = chaotic/random).
  • --top_k (default: 40): Restricts the model to only sample from the top k most likely next tokens.
  • --seed (default: None): Sets a manual PyTorch seed for perfectly reproducible text generation.

Example usage:

python generate.py checkpoint_8L-8H-512D_final.pt --prompt "Lily found a box" --temperature 0.7

4. tokenizer.py

Handles training a custom ByteLevel BPE Tokenizer from scratch and provides memory-efficient chunked iterators to feed PyTorch tensors directly to the GPU. You rarely need to run this manually as train.py handles it automatically.

Recommended Data Sets

TinyStories:https://huggingface.co/datasets/roneneldan/TinyStories

TinyStoriesV2: https://huggingface.co/datasets/roneneldan/TinyStoriesV2

About

A lightweight framework for training an autoregressive language model (GPT) from absolute scratch on a local GPU.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages