Transformer & Sampling

deep-learning

Handwritten notes on the Transformer architecture: self-attention, multi-head attention, positional encoding, and sampling strategies for text generation.

These notes cover the Transformer architecture and generation techniques:

  • Self-Attention: Query, key, value projections, scaled dot-product attention, and attention score computation.
  • Multi-Head Attention: Parallel attention heads, concatenation, and why multiple heads capture different relationships.
  • Positional Encoding: Sinusoidal encodings, learned positions, and injecting sequence order into attention.
  • Transformer Blocks: Layer normalization, residual connections, feed-forward layers, and encoder-decoder structure.
  • Sampling Strategies: Greedy decoding, temperature scaling, top-k, top-p (nucleus sampling), and beam search.