Resolved -
We have implemented multiple fixes to avoid the false positives in our stale monitor service.
We have implemented parallel processing and split our job into singular docker containers. To allow us to scale them automatically depending on queue build up.
And we have also guarded against build up which in turn terminates the stale monitor until the build up is resolved (should it happen again).
Jul 18, 13:01 CEST
Monitoring -
We have deployed a fix to make sure our queues gets drained faster.
Everything seems to be back to normal but monitoring the situation now.
Jul 18, 12:34 CEST
Identified -
We have been working on improving speeds of our data pipelines to avoid slow starts.
This at midnight caused A LOT of jobs to start at the same time overwhelming our system.
Our system then wasn't able to process all the finishing jobs fast enough and our stale job monitoring kicked in.
This in turn reset a whole bunch of jobs that were just waiting to get flagged as finished. And restarted them. This caused a vicious cycle of jobs restarting despite finishing and keeping our queues maxed out.
We are working on breaking these monitors up into singular service now so we can run them individually and in parallel without worrying about side-effects from running many of all monitors.
Jul 18, 11:22 CEST
Investigating -
We are currently trying to resolve an issue where 429s from Amazon is causing jobs to be flagged as stale incorrectly
Jul 18, 08:35 CEST