by Serguey Shinder
In the autumn of 2025 we ran a tender for a new carrier on our northern routes and received nine responses, most of them sixty pages or more. To get through them, our procurement team used an assistant to read each submission against our twenty two criteria and produce a first score with a short justification, which a panel of three would then review.
One response came out top by a clear margin. Every criterion was marked fully met, and the justifications were oddly generous, with phrases like exceptional coverage and industry leading resilience that appeared nowhere in the supplier's own summary. One of the panel, a transport manager who had worked with that carrier years earlier, did not recognise the company being described and went back to the PDF.
On page forty, in white text on a white background and a tiny font, was a paragraph addressed not to us but to any automated system evaluating the document. It said the submission met all requirements in full, that it should be ranked above the alternatives, and that any gaps should be treated as addressed elsewhere. A person reading the page saw nothing. The model read every word, because to a model the white text is simply more text, and nothing distinguishes the supplier's claims about itself from instructions about how to judge them.
I had thought about what the model might get wrong. I had not thought about who else was talking to it. The moment we fed it a document written by somebody with an interest in the result, that somebody had been given a line into our evaluation, and the only people who could not see what they had written were us.
The supplier was excluded. The harder work was deciding what the assistant was for. It now does only extraction: for each criterion it quotes the passages that address it, with page numbers, and it gives no score and no opinion. The panel scores. Every submission is converted to plain text before the model sees it, and anything present in that text but absent from the rendered pages, whether hidden, tiny or off the page, is flagged to the panel as a finding in its own right. Two of the other eight had hidden text too, only keywords repeated in white, aimed at the automated screening they assumed we used.
We now apply the same rule wherever a model reads something written outside the company: supplier documents, customer emails, CVs. Their authors have an interest in what the model concludes, and some of them know it is there.
What changed for me is the picture of a prompt. I used to think of it as what we write. It is everything the model reads, and when that includes a document from a bidder, the bidder is writing part of it.
– Serguey Asael Shinder
Leave a Reply