Annotated study fork of LLMs-from-scratch — GPT-2 architecture in PyTorch, tokenisation, pre-training loop, instruction fine-tuning, RLHF pipeline notes, and extended implementation commentary.
deep-learning pytorch transformer neural-networks educational language-model attention-mechanism tokenization llm-training-from-scratch pedagogical-resource
-
Updated
Aug 2, 2025 - Jupyter Notebook