by Serguey Shinder
Support tickets used to be sorted by a person. Whoever was on the morning shift read each one and sent it to one of five teams, and it took about forty minutes a day. We replaced that with a classifier, measured it against six months of history, and it agreed with the humans ninety two per cent of the time. That was better than the humans agreed with each other, which we checked, and we were pleased.
What we did not measure was what the eight per cent did next.
A customer wrote in about a payment that had failed during a renewal. The classifier read the word integration, which appeared in the sentence describing where she had seen the error, and sent it to the integrations team. It sat on their board for nine days, behind their own work, because it was not their ticket and nobody on that team had any reason to care about it. She had cancelled by the time anybody read it properly.
Here is what the morning shift had been doing that we never wrote in any specification. When a person could not tell which team owned a ticket, they got up and asked. Twenty seconds, three or four times a day. The value of that step was never its accuracy. It was that it contained a state called I do not know, and that state had a fast and social route out of it.
Our classifier had no such state. It returned a label for every ticket, with a confidence number attached, and the confidence was highest exactly where it should have been lowest, on tickets worded like the common ones. A system that always answers has not removed the uncertainty. It has removed the signal that uncertainty was present, which is a different thing and considerably worse, because the ticket now looks handled.
The fix was to give it permission to abstain. Anything below a threshold goes to a human queue, and we tuned the threshold until about twelve per cent of tickets land there, which is roughly the rate at which the morning shift used to get up and ask. Every routed ticket now carries a one line reason and a bounce button, and we watch the bounce rate per team, because a receiving team knows within seconds what our evaluation could never see.
Nothing here is an argument against the tool. We still do not sort tickets by hand. But when I replace a step a person was doing, the question I ask first is no longer how often they were right. It is what they did on the occasions when they did not know, because that behaviour is usually holding up more of the process than anybody realises, and it is never in the description of the job.
– Serguey Asael Shinder
Leave a Reply