Sharing a technical write-up of our recent incident where our service repeatedly crashed due to a poison pill event in the async workers.
Interesting from a general incident coordination angle, but also in terms of mitigations some of which apply generally to web apps, such as splitting work by category/type for improved reliability.
1 comment
[ 4.3 ms ] story [ 12.8 ms ] threadSharing a technical write-up of our recent incident where our service repeatedly crashed due to a poison pill event in the async workers.
Interesting from a general incident coordination angle, but also in terms of mitigations some of which apply generally to web apps, such as splitting work by category/type for improved reliability.
Hope you enjoy, feedback always welcome!