by Serguey Shinder
The advisory landed on a Tuesday and concerned a library we used in a document renderer. It was a real vulnerability and the fix was a version bump, one line in one file, the sort of change I would normally make between two meetings.
That component had last been released in 2019. It was still running, still serving customers, still perfectly healthy in every dashboard we had. Nobody had built it from source in four years.
The repository was where it should be. The build failed in eleven seconds. What followed took five working days, and almost none of it was the change.
The dependency manifest specified ranges rather than exact versions, so what it resolved to in 2019 and what it resolved to that Tuesday were different sets of packages, and the new set did not compile together. Two dependencies came from an internal mirror decommissioned in 2021 during a cost exercise. The build required a compiler version no longer packaged for any operating system we run, so it had to be reconstructed inside a container image built from an archived base that itself needed patching. And the final publishing step used a credential belonging to a person who left in 2020, in a pipeline definition nobody had read since.
We shipped the one line change on the Monday.
What I took from that week is a distinction I had never made. We had spent real effort deciding how long we would support our software. We had spent none at all on how long we would be able to change it, and those are separate properties with separate decay rates. Running is passive. A container image will run happily forever and will not degrade in any way that anybody can see. Building reaches out into the world on every attempt, to registries, mirrors, package indexes, certificate authorities, base images and the continued willingness of other people to host things, and every one of those is somebody else's decision. It rots continuously and silently, because nothing anywhere tells you that your build no longer works until the day you need it to.
The question that actually determines whether an old system is an asset or a liability is not whether it still works. It is whether you could ship a fix to it by Friday, and almost nobody measures that, because the answer feels obviously yes right up until it is obviously no.
We now rebuild every service that is still in production on a quarterly schedule, changed or not, and the job exists only to prove that we can. The first time we ran it, three others failed. Versions are pinned exactly, dependencies are mirrored somewhere we control, and the toolchain is recorded with the release rather than assumed.
Everything still runs. That was always going to be true, and it is the reason nobody looks.
– Serguey Asael Shinder
Leave a Reply