Pre-trained / 15 chapters
How foundation models learn
Fifteen focused chapters connect the data, mathematical, architectural, and systems decisions behind a base language model.
- 01Read ↗
What a Language Model Learns
Token representations, autoregressive factorization, teacher forcing, cross-entropy, and the limits of next-token prediction.
- 02Read ↗
From Unicode to Byte-Pair Encoding
Code points, UTF-8 bytes, merge learning, vocabulary construction, compression, and exact encode–decode behavior.
- 03Read ↗
Corpus Design and Data Lineage
Source registration, parsing, filtering, deduplication, decontamination, mixture weights, document boundaries, and immutable shards.
- 04Read ↗
Batches, Tensor Shapes, and Shifted Targets
Turn token streams into input–target windows, track batch and sequence dimensions, mask invalid positions, and prevent split leakage.
- 05Read ↗
The Transformer Residual Stream
Follow embeddings through encoder, decoder, and encoder–decoder families while keeping every tensor contract explicit.
- 06Read ↗
Queries, Keys, Values, and Causal Masks
Build scaled dot-product attention, understand multi-head projections, apply causal and padding masks, and verify the softmax axis.
- 07Read ↗
Residual Blocks, Normalization, and MLPs
Assemble pre-norm attention and feed-forward branches, preserve gradient paths, and reason about width, depth, and activation choice.
- 08Read ↗
Position and Context Length
Compare learned positions, sinusoidal encodings, rotary embeddings, relative bias, extrapolation, and the real cost of longer context.
- 09Read ↗
KV Caches and Efficient Attention
Separate full-sequence training from incremental decoding, account for cache memory, and test fused kernels against a clear reference.
- 10Read ↗
Scaling Laws and Training Budgets
Allocate parameters, clean tokens, context, batch size, compute, and wall-clock budget without confusing size with capability.
- 11Read ↗
Mixture-of-Experts Routing
Map tokens to experts with top-k routing, distinguish total from active parameters, and trace dispatch and combine operations.
- 12Read ↗
Expert Capacity, Balance, and Parallelism
Handle overflow, token dropping, auxiliary losses, shared experts, expert parallel communication, and routing failure modes.
- 13Read ↗
Initialization and Optimization
Set residual-aware initialization, AdamW parameter groups, warmup and decay schedules, gradient accumulation, and clipping.
- 14Read ↗
Distributed Training and Throughput
Use mixed precision, fused kernels, compilation, data parallelism, all-reduce, and exact global-token accounting.
- 15Read ↗
Checkpoints, Evaluation, and Reproducibility
Package weights with optimizer and data state, compare fixed-token runs, audit contamination, and preserve the evidence behind every result.