-
PyTorch Transforms: From Raw Data to the Tensor Your Model Actually Sees
-
Models From First Principles 07: PACS — Building an Optimizer From Gradient Statistics
-
Models From First Principles 04: HRM — Hierarchical Reasoning With Fast and Slow Recurrent State
-
Models From First Principles 03: SICQL — Building a Model From Q, V and Policy Networks
-
10: Build a Small GPT-Style Language Model From Scratch
-
PyTorch Performance Debugging: CUDA OOM, Slow Training, GPU Utilization and torch.compile
-
PyTorch Model Not Learning? A Systematic Debugging Guide
-
PyTorch Attention Shapes: Q, K, V, Multi-Head Attention Masks and Transformer Dimension Errors
-
PyTorch CNN Shape Errors: Conv2d Output Sizes, Channels, Flatten Bugs and How to Debug Them
-
PyTorch DataLoader Performance: num_workers, pin_memory, Prefetching and Why Your GPU Is Waiting
-
The Most Important Idea in PyTorch: Recursive Composition
-
Build a Neural Network From Scratch in PyTorch Without nn.Module
-
PyTorch Autograd Debugging: requires_grad, detach, backward() and NaN Gradients
-
PyTorch Tensor Shapes: Broadcasting, Reshape, View, Permute and the Errors That Waste Your Time
-
00: What Are We Actually Doing?
-
Writing Neural Networks with PyTorch