tinyGPT
A GPT-2 124M model written from scratch in PyTorch: the model, data pipeline, distributed training loop and evaluation harness. Trained on 10B tokens on a single 4xA100 node in about five hours.
Things I've built to understand how they work, mostly from scratch.
A GPT-2 124M model written from scratch in PyTorch: the model, data pipeline, distributed training loop and evaluation harness. Trained on 10B tokens on a single 4xA100 node in about five hours.
tinyGPT with rotary position embeddings replacing the learned position table. It reaches a validation loss of 3.04 against 3.29 for the published GPT-2 checkpoint, on a different training set, and 0.316 on HellaSwag against 0.296. Measuring loss at longer contexts shows RoPE does not extrapolate to unseen lengths out of the box.
Supervised fine-tuning of the pre-trained tinyGPT checkpoint with LoRA. The layer, the injection and the freezing are written from scratch, with no PEFT library. Training 0.36% of the parameters lifts ARC-Easy accuracy from 52.69 to 58.08. The same fine-tuning is also done on Qwen3-4B with Hugging Face PEFT.
A small automatic differentiation engine, worked through from the ground up in a single notebook.
An alert system for camera setups that runs entirely locally and in real time, on the most portable setup possible.