How Large Language Models (LLMs) Work
Master the mechanics of transformers and self-attention to see how modern AI processes language. You will learn exactly how these models predict and generate human-like text. * **Architectural flow** of transformer encoders and decoders * **Self-attention mechanisms** for contextual token relationships * **Matrix operations** driving parallelized sequence processing * **Probability distributions** for next-token generation logic
10 sections ยท 27 lessons
Course outline
Foundational Concepts
- Text to Numbers
- Layered Processing
- Tokenization
- Embeddings
Sequential Processing Problems
- Sequential Limitations
Transformer Architecture
- Parallel Processing
- Query-Key-Value
- Attention Weights
Self-Attention Mechanics
- Self-Attention
- Causal Masking
- Multi-Head Attention
Model Architectures
- Architecture Types
- Model Scale
Pre-Training Process
- Pre-Training
- Next-Word Prediction
Fine-Tuning and Alignment
- Fine-Tuning: Specializing Pre-Trained Models
- RLHF Overview
- RLHF Stages
Text Generation
- Autoregressive Generation
- Softmax Transformation
- Temperature: Controlling Output Randomness
- Sampling Strategies
Generation Control
- Creativity vs Precision
Model Limitations
- Hallucinations
- Mitigating Hallucinations
- Context Windows
- Context Management