The pilot answered fluently, the committee nodded, and three weeks later someone checked a figure against the ledger. It was wrong. It had been wrong on the day of the demonstration. The model was blamed, the vendor was changed, and the next pilot was built on the same foundation as the first: a copy of the business, taken on a Tuesday, asked questions on a Friday.
The pilot is a ritual
Europe is the world champion of the proof of concept. Every large organisation on the continent holds a portfolio of pilots that will never ship, supervised by a transformation office that will be reorganised before anyone has to say why.[1] This is not incompetence. A pilot confers every political benefit of action and none of its exposure. Shipping requires that a named person bear a consequence, and the pilot is designed so that nobody does.
So when a pilot fails, the diagnosis is chosen for its comfort. “The model hallucinated” flatters everyone: the vendor, whose product was simply ahead of its time; the buyer, who was simply prudent; the integrator, who will simply try a larger model next quarter.
The copy was wrong
Look at what the pilot was built on. An extract. A warehouse load. A vector index refreshed nightly, or weekly, or once, in March, by the intern. Every one of these is a copy of the business, and a copy is stale by construction: it is right at the instant it is taken and drifts from then on. The contract was renewed on Thursday; the copy says it lapsed. The part was reclassified; the copy holds the old code. The model read the copy and answered correctly about a world that had already moved.
A large share of what gets filed as hallucination is exactly this: inconsistent, stale or partially replicated sources feeding the wrong context to the system.[2] The utterance was fluent because the model is very good at fluency. The figure was wrong because nothing between the question and the answer ever touched the record.
Plausibility is the adversary
An answer that is wrong and looks wrong is a minor inconvenience, caught by the least attentive reviewer. An answer that is wrong and looks right is a structural hazard, and generative systems are optimised, precisely and by design, to look right.[3] Knowledge has always consisted in the capacity to exhibit, on demand and before a hostile examiner, the chain by which one arrived at a proposition. Strip away the chain and what remains is an opinion with good grammar.
This is why capability curves have stopped predicting adoption. The institutions declining to deploy are not waiting for a better model. They are waiting for something that can be held.[4]
What a pilot that ships does differently
Four conditions, and a roadmap that lists them among its features has already misread them.[5]
It goes to the data. The software reads across the systems the organisation already runs and computes on what it reads. Nothing is migrated, so there is no copy to be stale. The figure in the answer is the figure in the system of record, at the moment the question was asked.
The model translates; it does not compute. The same question, put twice, returns the same answer, to the letter and on every occasion. This is obtainable, and the method is no mystery: the model is confined to turning an intention into a plan expressed in the vocabulary of the ontology, and the plan runs without it.[5:1] A figure that a model produced is a probability. A figure that a plan produced is a fact you can re-run.
Every figure reaches the record it came from. Provenance is not a feature. It is the condition under which anyone will act on the answer who will later be asked to justify the action.
A named human bears the decision. The machine proposes. A person, named and timestamped, approves, and an error before that approval is a draft, thrown away. A decision with no name at the foot of it is weather.[5:2]
The questions to ask before the next pilot
Put these to the vendor, in writing, before the demonstration.
- Where does this figure come from, and can I follow it back to the row?
- If I ask the same question tomorrow with the same parameters, do I get the same answer, byte for byte?
- What happens between the answer and the change to my system? Who sees it first?
- Whose name is on the decision?
- Is any of this a copy of my data, and when was it taken?
A vendor who answers all five with a mechanism you can verify from your own network has built a machine of consequence. A vendor who answers with a roadmap has built a pilot.
Monarch is built on the four conditions above, inside your infrastructure, on the systems you already run. See it on your data.