The Regular Expression I Finally Understood

by Serguey Shinder

There was a validation rule in our intake code that nobody could read. Seventy-odd characters of regular expression, written before I joined, checking the shape of a partner reference. It had been correct for four years and every one of us treated it as a black box, which meant that when the format changed slightly we were all rather afraid of it.

So I gave it to a model and asked for an explanation and a readable rewrite. What came back was genuinely good. The explanation walked through it piece by piece, the prefix, the numeric block, the optional suffix. The rewrite was in extended form, spread over eight lines with a comment beside each part, and it came with a handful of tests. For the first time in four years I read that rule and understood it. My reviewer, who was equally relieved, approved it in about eight minutes.

Two weeks later we started getting duplicate accounts for the same partner reference. The new expression had no anchor at the end. The old one did. References with trailing whitespace passed, and so did references with a slash and a note appended, which one partner had been doing for years and which the old rule had silently rejected the whole time, sending them down a correction path we did not know existed.

The explanation had been accurate about everything it discussed. The anchor was simply not one of the things it discussed, and my attention had followed the shape of the explanation exactly. Anything it did not name became invisible to me. I had checked the parts it listed against the parts it listed, which feels like verification and is closer to reading a summary of a book and then reviewing the summary.

The second half of it is about readability buying trust. I had never been able to read the original. The moment I could read the replacement, I felt I had inspected it, when what had actually happened was that I had inspected a different thing that was easier to look at.

There was a check available that cost nothing and that I did not do until afterwards. We had four hundred thousand historic references in a table. Running every one of them through both expressions and diffing the results took twenty minutes, found the discrepancy immediately, and would have found it before the merge just as easily. Adding random junk to the end of a few thousand of them would have found it too.

Any rewrite of something I could not previously read is a behaviour change until proven otherwise, and the proof is old and new, same inputs, compared. Equivalence is cheap to test and impossible to see, and it does not become more visible because the new version is prettier.

– Serguey Asael Shinder

Leave a Reply