Inference Engineering

Philip Kiely's guide to running generative AI models in production, explaining how workload requirements, model architecture, GPUs, inference engines, techniques such as quantization and speculative decoding, other modalities, and autoscaling determine the speed, cost, and reliability of inference.

5 sections ยท 16 lessons

Course outline

Before You Optimize

  1. Why Constraints Make Inference Fast
  2. When Dedicated GPUs Beat Token APIs
  3. The Smallest Model That Passes Your Evals
  4. TTFT, TPS, and the Slow Tail
  5. Should Lumen Go Dedicated?

Where Time Goes

  1. Prefill, Decode, and the KV Cache
  2. Compute-Bound or Memory-Bound?
  3. Picking an Engine and Benchmarking It

Optimization Techniques

  1. Quantization Without Losing Quality
  2. When Speculative Decoding Pays
  3. Prefix Caching and Cache-Aware Routing
  4. Splitting a Model Across GPUs
  5. When to Disaggregate Prefill and Decode
  6. Tuning a Code Editor's Model

Beyond Text

  1. Serving Voice: Transcription and Speech
  2. Faster Image and Video Generation

Production

  1. Autoscaling and Cold Starts
  2. Reliability and Safe Rollouts
  3. Surviving a Viral Launch

Start learning with Wondering