Attention Is All You Need

Vaswani and colleagues introduce the Transformer, an encoder-decoder architecture built around attention rather than recurrence or convolution, and report its parallel training advantages, translation results, and transfer to constituency parsing.

7 sections ยท 21 lessons

Course outline

Motivation

  1. RNN Limitations
  2. Attention-Only Model

Core Architecture

  1. Encoder-Decoder Stack
  2. Encoder Layers
  3. Decoder Masking

Attention Mechanisms

  1. Attention Basics
  2. Dot-Product Attention
  3. Multi-Head Attention
  4. Attention Types

Supporting Components

  1. Feed-Forward Networks
  2. Embeddings
  3. Positional Encoding

Advantages of Self-Attention

  1. Complexity Analysis
  2. Long-Range Dependencies

Training Details

  1. Batching Strategy
  2. Training Efficiency
  3. Learning Rate Schedule
  4. Regularization

Results and Impact

  1. Translation Results
  2. Ablation Studies
  3. Generalization

Start learning with Wondering