Incident: An internal system responsible for automatically scaling backend server capacity in one of our compute clusters stopped replacing capacity that had been cycled out during routine maintenance. This caused a gradual reduction in available capacity over several hours. A subsequent deployment was activated with insufficient capacity, causing all new requests for that compute cluster to fail. To mitigate the impact, we reverted to a previous release revision, which still had sufficient capacity.
Impact: For approximately 30 minutes, customers whose traffic was handled by this compute cluster saw full downtime; other customers were unaffected. No customer data was lost.
Moving forward: We have added additional monitoring to detect this type of capacity-scaling failure much earlier, and have added safeguards to prevent deployments from shifting traffic before sufficient healthy capacity is confirmed.
Our metric considers a weighted average of uptime experienced by users at each data center. The number of minutes of downtime shown reflects this weighted average.