Resolved -
Fully resolved. The root cause was our infrastructure was using old hand crafted modules instead of our standard module library. They hadn't been updated in a long time. The problem is the terraform source code in that module was not pinned to a commit, resulting to complete deletion of our primary load balancer. We also did not have delete protection enabled on the load balancer (our standard modules have this on by default).
We could have resolved the downtime sooner, but with only a few users at this time we took the opportunity to rebuild our infra with our standard modules, ensuring delete protection is now enabled.
We're also looking at adding some product features to warn users who might find themselves in a similar situation.
Jul 28, 18:27 UTC
Monitoring -
All systems restored, monitoring for any remaining issues
Jul 28, 17:23 UTC
Update -
Main API is restored, but few remaining backend microservices still to fix before pipelines and deploys will work.
Jul 28, 17:11 UTC
Update -
Still working on a fix, should be resolved very soon
Jul 28, 15:51 UTC
Identified -
The issue has been identified and a fix is being implemented.
Jul 28, 14:22 UTC