by Serguey Shinder
Subscription renewals were failing for about a third of our customers from twenty past nine on a Tuesday until twenty to two. I owned the pricing service, and within ten minutes I was certain it was not us. Error rate was two hundredths of a per cent. Latency was forty milliseconds. Nothing had shipped from my team in six days.
So I proved it. I spent ninety minutes building the proof, and it was good work. Graphs with the incident window shaded. A trace showing a healthy request answered correctly at the height of the failures. A short written summary with timestamps. I posted it into the incident channel at about eleven, and it was accurate, and it was the only thing I produced in ninety minutes during which a third of our customers could not renew.
The uncomfortable part is what else was on that channel when I posted. Two other teams had spent the same morning doing exactly what I had done. There were three careful demonstrations of innocence in that channel and not one hypothesis about the failure. We had, between us, spent the best part of four engineer hours establishing that the outage was not being caused by any of the people in the room, which everybody suspected already and which brought nobody a step closer to a fix.
It ended when somebody from the support team, who had not been invited, asked whether anyone had noticed that every failing renewal she had seen was on an account with a promotional rate. That was it. A change to promotions had gone out the previous evening. From her sentence to a resolved incident was thirty five minutes.
She asked a different question from the rest of us. We were all answering whose is this, because that is what the room was implicitly asking, and it is a question every one of us could make progress on immediately using tools we already had open. She asked what the failures have in common, which no dashboard answers, because the answer lives in the intersection rather than in any one system.
Two things changed after that. When an incident starts, anybody who owns a service says one sentence, covering what they can rule out and what they cannot, and then says what they are going to go and look at. No documents. And whoever runs the incident asks what the failing requests share before asking which component is broken.
The distinction I carry is personal rather than procedural. Not being at fault is a fact about me. The outage is a fact about the customer. When I notice the pull to establish the first, it is a reliable signal that I have stopped working on the second, and the pull is strongest exactly when I am innocent, because that is when the proof is easiest and most satisfying to build.
– Serguey Asael Shinder
Leave a Reply