Lab · Interactive

On call

A day of production traffic in two and a half minutes. Keep the system inside its SLO through a traffic spike, a crashed node, a flash sale and a cold cache, and spend as little as you can doing it.

503 · Service unavailable

This viewport is below minimum capacity.

On call runs a live architecture diagram, a metrics panel, a pager and six scaling controls side by side. A phone can't fit all of them without hiding something you need to watch, so the load balancer is routing you away. Open it on a tablet or desktop.

GET /lab HTTP/1.1
Viewport: <640px

HTTP/1.1 503 Service Unavailable
X-Min-Viewport: 640px
Retry-After: larger-display

What the simulation models

It is small, but the parts that make real systems fall over are there. Each tier is a pool with a fixed capacity. Latency stays flat until about 80% utilisation and then bends sharply upward, the way queueing delay does. Past 100% a tier builds a backlog, and once the wait passes the one-second client timeout the excess is dropped. Overload doesn't just fail the extra requests: it makes every request slow.

That is why the rate limiter is worth a dollar an hour. Shedding the traffic you can't serve keeps the requests you can serve fast. Rejected requests still count against the SLO, but it is a few percent instead of everything.

  • Cache. Absorbs up to 90% of reads once warm. A cold cache is the same as no cache, which the 21:00 deploy will remind you of.
  • Read replicas. Split read load with the primary. They take longest to provision because they copy data first.
  • Queue and workers. Writes behind a queue are acknowledged at once and drained by workers. The workers back off when the primary has no headroom, so a burst becomes a backlog instead of an outage.

The p99 shown is derived rather than sampled. Each path a request can take (cache hit, primary read, replica read, sync write, queued write) has a mean latency and a backlog wait. Its tail is treated as exponential, and p99 is the latency that 1% of all requests exceed. The error budget is 1% of the day's projected traffic.