Attention Is All You Need
Vaswani and colleagues introduce the Transformer, an encoder-decoder architecture built around attention rather than recurrence or convolution, and report its parallel training advantages, translation results, and transfer to constituency parsing.
7 sections ยท 21 lessons
Course outline
Motivation
- RNN Limitations
- Attention-Only Model
Core Architecture
- Encoder-Decoder Stack
- Encoder Layers
- Decoder Masking
Attention Mechanisms
- Attention Basics
- Dot-Product Attention
- Multi-Head Attention
- Attention Types
Supporting Components
- Feed-Forward Networks
- Embeddings
- Positional Encoding
Advantages of Self-Attention
- Complexity Analysis
- Long-Range Dependencies
Training Details
- Batching Strategy
- Training Efficiency
- Learning Rate Schedule
- Regularization
Results and Impact
- Translation Results
- Ablation Studies
- Generalization