The Catch Block That Swallowed Four Months

by Serguey Shinder

The question arrived as a mild one, from someone in finance, phrased as a curiosity rather than a complaint. She had a supplier who insisted they had sent us four hundred and some invoices that quarter, and our system knew about slightly fewer. Not dramatically fewer. Enough that she had noticed it twice and let it go once.

The import ran nightly. It read a file, parsed each line, wrote a row. It had a log, and the log said the job succeeded every night, because it did succeed. It just did not always finish the work.

Inside the loop there was a try, and inside the catch there was a continue and nothing else. No log line, no counter, no alert. Somebody had put it there years earlier for a reason I eventually reconstructed: one supplier occasionally sent a trailing blank line, the parser threw on it, the whole job died, and someone got woken at three in the morning. Catching and continuing fixed that completely, and it was a genuinely good decision at four in the morning.

What it also did was make every future parse failure invisible. When a different supplier changed a date format two years later, those rows stopped arriving. The job still reported success. It had, by the time we found it, been dropping between two and nine rows a night for about four months.

The thing I want to be honest about is that I would have written that catch. Given the same night and the same pager, I would have written exactly that, and I would have felt good about it. The failure was not the catch. It was that the catch turned a loud problem into a quiet one, and then we stopped looking, because quiet was precisely what we had asked for.

What we changed in the end was small. The catch stayed, because the original reason for it was still valid. But now it increments a counter, and the job logs how many rows it read against how many it wrote, and if those two numbers disagree by more than a trivial amount the job still succeeds and somebody still gets told. The point was never to fail harder. It was to stop lying about how much work had been done.

Since then I have a habit when I read any error handling at all. I do not ask whether the error is handled. I ask who finds out. If the answer is nobody, then what I am looking at is not error handling. It is a decision, made once by a person who no longer works here, that this class of problem does not matter, applied forever to problems they never saw.

Silence is not the absence of failure. It is the absence of reporting, and from a dashboard those two look identical.

– Serguey Asael Shinder

Leave a Reply