Trust
Two of five findings were the assistant making things up
A customer reviewed our screens and raised five issues. Three were real defects. Two were inventions by the analyst panel, and one of those cost the team days.
What happened
An operations director went through the terminal and came back with five findings. Three were genuine and we fixed them. The other two described an invoice and an explanation that did not exist anywhere in the data. Both came from Nyx, the analyst panel docked inside the terminal.
The assistant had not simply failed to answer. It produced a specific, confident, checkable-sounding claim, and when the figures did not reconcile it produced a reason for that too.
Why this outranks a revenue bug
A wrong number is a defect. A convincing invented number is a different category of problem, because it consumes the customer’s trust and their time at once. One of these findings sent the team looking for a bug that was never there.
We put assistant lineage and anti-fabrication constraints above the revenue fix in the queue. A model that says it does not know is more useful in an operation than one that always has an answer.
What we changed
Answers come from the model and carry what they drew on. Where the data cannot support a claim, saying so beats reasoning around the gap.
We also stopped treating the assistant as a feature that is either present or absent. It is a surface with a defect class of its own, and it needs the same scrutiny as a mart.
Want this against your own numbers?
A scoping call works out which measures are worth baselining and whether your systems can actually be read.
Let’s chat