These notes cover the Transformer architecture and generation techniques:
- Self-Attention: Query, key, value projections, scaled dot-product attention, and attention score computation.
- Multi-Head Attention: Parallel attention heads, concatenation, and why multiple heads capture different relationships.
- Positional Encoding: Sinusoidal encodings, learned positions, and injecting sequence order into attention.
- Transformer Blocks: Layer normalization, residual connections, feed-forward layers, and encoder-decoder structure.
- Sampling Strategies: Greedy decoding, temperature scaling, top-k, top-p (nucleus sampling), and beam search.