We Took Payment Twice And Both Attempts Were Correct

by Serguey Shinder

On a Monday in March we took twenty nine payments twice. The customers noticed before we did, which is the worst order for these things to happen in, and by the time the first support message arrived the duplicate charge had already cleared on most of the cards.

The payment code was not careless. It called the gateway, waited for an answer, and wrote the result to the order. What it did not do was decide anything about timeouts, because the HTTP client we used had a retry policy built into it, set to two attempts, configured once in 2019 by somebody sensible who was thinking about a flaky internal service. The gateway that morning was slow rather than broken. It accepted the first request, took longer than our thirty second timeout to say so, and then received a second identical request from a client that had quietly given up on the first.

Our logs said we charged each customer once, and that cost us most of a day. We wrote the log line after the response came back, which meant the attempt that timed out produced no line at all, so our own record of the morning was a tidy list of single successful payments. Everything we could see agreed with us. The truth existed only on the gateway's own dashboard, which nobody thought to open for three hours, precisely because our logs looked clean.

Two changes fixed it and neither was clever. Every payment now carries an idempotency key that we generate before the first attempt and store on the order, so the gateway recognises a repeat and answers with the original result instead of taking the money again. Their API had supported that since years before we integrated. Nobody had read that far down the page. And automatic retries are now off by default in our client, switched on per call only for operations somebody has looked at and decided are safe to repeat, which turns out to be most reads and almost no writes.

We moved the log line too. It is written before the request goes out, with the key in it, and updated afterwards with whatever came back or did not. A record that describes only the calls that finished is a record of your successes.

What I carry from it is a distinction I did not used to make. A timeout is not a failure. It is the absence of news, and those are entirely different facts that our code had been treating as one. Whenever you cross a boundary you can succeed and still not be told, and every retry sitting on top of that assumption is a decision to do the thing again on the strength of not having heard back. Ask what happens if it already worked, before you ask how many times to try.

– Serguey Asael Shinder

Leave a Reply