Building & Scaling LLM Infrastructure

Understand how large language models are actually trained and served at scale: what the hardware physically constrains, why serving systems are designed the way they are, how to size and cost a deployment from first principles, and where the security boundaries sit. Written for a learner with strong mathematics and general computing ability but no hands-on experience of GPUs, clusters, orchestration or inference servers — so every tool and piece of infrastructure is named and explained before it is used, and every quantitative claim is shown as an explicit calculation with units. Build from the physical machine upward through memory and bandwidth arithmetic, distributed training, inference mechanics, serving at scale, reliability, and security, ending able to size a deployment, read a vendor benchmark critically and diagnose a degraded system.

14 sections · 78 lessons

Course outline

Orientation: What This Machinery Actually Is

  1. The Shape of the Problem
  2. A Tour of the Physical Stack
  3. Training versus Inference Are Different Businesses
  4. The Vocabulary You Will Meet

Accelerator Hardware

  1. What a GPU Provides That a CPU Does Not
  2. Tensor Cores and Matrix Multiplication
  3. The Memory Hierarchy
  4. Reading an Accelerator Spec Sheet
  5. Alternatives to GPUs

The Arithmetic That Governs Everything

  1. Counting Parameters and Weight Memory
  2. FLOPs per Token
  3. Arithmetic Intensity
  4. The Roofline Model
  5. Why Batching Changes the Physics

Numerical Precision

  1. Floating Point: Range versus Precision
  2. Mixed-Precision Training
  3. Quantization for Inference
  4. What Quantization Actually Costs

Interconnect and Collective Communication

  1. Why the Network Is the Bottleneck
  2. Collective Operations
  3. Topology Awareness

Distributed Training

  1. Memory Accounting for a Training Step
  2. Data Parallelism
  3. Sharding Optimizer State and Parameters
  4. Tensor Parallelism
  5. Pipeline Parallelism
  6. Sequence and Context Parallelism
  7. Composing 3D Parallelism
  8. Activation Checkpointing

Running a Training Job

  1. The Data Pipeline
  2. Checkpointing
  3. Failures, Stragglers and Silent Corruption
  4. Monitoring a Run
  5. Schedulers and Orchestration
  6. Fine-Tuning Infrastructure

Inference Fundamentals

  1. Prefill and Decode
  2. The KV Cache
  3. Why the KV Cache Limits Concurrency
  4. Latency Metrics That Matter
  5. The Throughput-Latency Trade-off

Inference Optimization

  1. Static Batching and Its Waste
  2. Continuous Batching
  3. Paged Attention
  4. Prefix Caching
  5. Speculative Decoding
  6. Chunked Prefill and Disaggregation
  7. Attention Kernel Optimisation
  8. Structured Output and Constrained Decoding

Serving at Scale

  1. What an Inference Engine Does
  2. Model Loading and Cold Start
  3. Replicas, Routing and Load Balancing
  4. Autoscaling Under Bursty Load
  5. Queueing, Admission Control and Load Shedding
  6. Multi-Tenancy and Fairness
  7. Streaming, Timeouts and Cancellation
  8. Serving Many Models

Reliability and Observability

  1. What to Measure
  2. Diagnosing Slowness
  3. Failure Modes and Graceful Degradation
  4. Deployment and Rollback
  5. Capacity Planning
  6. The Economics

Security: Model, Data and Platform

  1. Threat Modelling an LLM System
  2. Authentication, Authorization and Tenant Isolation
  3. Rate Limiting and Economic Abuse
  4. Model Extraction and Training-Data Leakage
  5. Supply Chain: Weights, Data and Dependencies
  6. Logging, Audit and Privacy

Security: The Application and Agent Layer

  1. Prompt Injection Is Structural
  2. Trust Boundaries in Agentic Systems
  3. Tool and Code Execution Sandboxing
  4. Output Handling as Untrusted Input
  5. Guardrails and Content Filtering
  6. Red-Teaming and Continuous Evaluation

Putting It Together

  1. Sizing a Deployment End to End
  2. Reading a Vendor Benchmark Critically
  3. Diagnosing a Degraded System
  4. Build, Buy or Call an API

Start learning with Wondering