Asana full outage for some users (2 hours, 25% outage)

Incident Report for Asana

Postmortem

We’ve been working to add a caching layer to our update pipeline, tuning it carefully and rolling it out gradually. On Wednesday, August 26 we enabled it for most use cases, and it initially performed well.

On Monday, August 31, a combination of unrelated infrastructure changes and peak traffic pushed the cache past its scaling limits. Once that threshold was crossed, the cache became unusable, and many pods serving read traffic for the Asana application could no longer serve it.

We mitigated the incident by reverting the system to use the previous, non-cached code path. The revert was successful, but recovery took longer than we would expect for this class of issue.

A fuller analysis is underway. We will follow up with root causes, action items, and improvements, including why recovery took as long as it did.

We were fully down for about 25% of our users, for 2 hours, 15 minutes.

Posted Sep 02, 2026 - 21:38 UTC

Resolved

User-facing symptoms have recovered. We'll continue to monitor, and will prioritize a retrospective to understand and prevent similar incidents in the future.
Posted Aug 31, 2026 - 17:18 UTC

Monitoring

We've applied a fix are seeing signs of recovery.
Posted Aug 31, 2026 - 16:47 UTC

Update

We have made configuration changes and see partial recovery, but we continue to see some elevated errors.
Posted Aug 31, 2026 - 16:24 UTC

Update

We are continuing to investigate the issue; we have reverted recent changes, and are working to identify the source of the errors.
Posted Aug 31, 2026 - 15:55 UTC

Investigating

We are investigating alerts for slow performance and application errors.
Posted Aug 31, 2026 - 15:14 UTC
This incident affected: US (App, API, Mobile).