Our Model Answered Questions About Sales Using Its Own Idea Of A Sale

by Serguey Shinder

In the spring of 2025 we gave our regional managers a tool that answered questions about our own figures. You typed a question in ordinary English, a model wrote a query against our reporting database, ran it, and gave you a table and a sentence. It was popular within a week. Managers who had waited days for an analyst could now ask which branches had grown fastest in the first quarter and have an answer before their coffee went cold.

In June one of those answers went into a board pack. Our finance director read it and asked why the branch sales figures were seven per cent higher than the ones in her own report for the same quarter.

Both numbers came from the same database. The difference was the word sales. To finance, a sale is an invoiced line, net of credit notes, without VAT. The model had found a table called orders and a column that looked like value, and summed it. That included orders later cancelled, quotes our system stores as orders with a particular status, and it ignored credit notes entirely, because they live somewhere else. Every one of those choices was reasonable from the names on the columns. None of them was ours.

When we looked further it was worse in a quieter way. The same question worded slightly differently sometimes produced a slightly different query, so two managers asking about the same branch could leave with two numbers, each looking authoritative, neither labelled with what it actually measured. The model had not made an error in the usual sense. It had written correct queries for its best guess at what we meant, and we had never told it.

So we told it, and then we stopped letting it guess. Finance wrote down twenty three measures in plain words with a single tested query behind each one: invoiced sales, gross margin, returns, and so on, each with an owner. The model can now only choose from those measures and from a list of ways to slice them. It cannot write a query against the raw tables at all. Every answer states which definitions it used, in a sentence anyone can read. And a question it cannot map to a defined measure comes back as not yet defined, which goes onto a list that finance works through every month.

The tool answers fewer questions than it used to. The ones it answers now agree with the finance report, and the not yet defined list has turned out to be one of the more useful documents in the company.

A model can find the column. Only the business can say what the word means.

– Serguey Asael Shinder

Leave a Reply