Google Scale Abstractions

Build practical intuition for the canonical Google-era abstractions behind large-scale storage and data processing: GFS, MapReduce, Bigtable, and the design tradeoffs they made visible. Designed for engineers and data-platform builders who already know distributed-systems basics and want sharper abstraction judgment.

3 sections ยท 10 lessons

Course outline

Scale Forces

  1. Workloads Outgrow Machines
  2. Commodity Failure Budget
  3. Locality Beats Bandwidth
  4. Design a Failure Budget

Batch Stack

  1. GFS Shares Disks
  2. MapReduce Hides Recovery
  3. Movement Shapes Jobs
  4. Stragglers Set Tails
  5. Diagnose a Slow Job

Structured Storage

  1. Bigtable Sparse Map
  2. Tablets Move Load
  3. Beyond One Cluster
  4. Choose the Next Abstraction

Start learning with Wondering