|
EXPLAINER |
AI FOR ACCOUNTANTS · PART 8 OF 16
Why an AI model invents numbers - and how to stop it
The model is not lying - it is completing a pattern, and that distinction changes how you build a prompt.
When an AI model writes a number that is not in the source file, it is not malfunctioning. It is doing exactly what it was built to do: produce the most plausible next token given everything that came before. The problem is that 'plausible' is a statistical property, not an accounting one. A figure that fits the sentence structure and the surrounding context will be generated confidently, whether or not it matches the trial balance.
The mechanism that causes this matters because it tells you where the risk is highest. Language models build a probability distribution over possible next words - or numbers - based on patterns learned during training. When you paste a reconciliation and ask for a variance summary, the model has seen thousands of variance summaries. It knows what one looks like. If the actual variance is ambiguous in the text you gave it, the model will fill the gap with something that looks right, formatted correctly, denominated correctly, signed correctly. The error is invisible at a glance because everything around the invented figure is accurate.
In accounting work this bites hardest in three places. First, totals: ask a model to summarize an aging schedule and it may calculate or recall a total rather than read the one you provided. Second, prior-period comparatives: the model may substitute a figure from its training data when you have not supplied last period's numbers. Third, footnote cross-references: when a workpaper refers to a schedule, the model may quote a figure from that schedule that it is extrapolating rather than reading. None of these look like hallucinations on the surface - they look like ordinary accounting prose.
The fix is a constraint on what the model is allowed to do when it cannot find the answer in the text you gave it. There are two levers. The first is a retrieval instruction: tell the model it may only state a figure if it can quote the cell or line it came from, using language like 'cite the exact line from the document or do not state the number.' The second is an uncertainty instruction: tell it to write 'not found in the provided file' rather than infer. Models follow these instructions reliably - they do not follow them automatically. If you do not write the instruction, the model defaults to fluency.
In practice, this means treating the prompt as a source-citation rule rather than a question. A prompt that says 'summarize the cash reconciliation' will produce a summary. A prompt that says 'summarize the cash reconciliation; for each figure you state, identify the row it comes from; if you cannot find a figure in the document, write not found' will produce a summary that is checkable. The difference is not the model - it is the constraint you put on the output before it runs. That constraint is the practitioner's job, and right now very few prompts in accounting workflows include it.
One more thing worth knowing: a model that is wrong with high confidence is not a sign of a bad model. Confidence in generated text is not a signal of accuracy - it is a signal of fluency. The model sounds certain because certainty is the pattern it learned from professional documents. Treating model output the way you treat a staff preparer's first draft - assume it needs a tick-mark, not a signature - is the correct posture. The review step does not go away; it just moves to a different place in the workflow.
WORKED EXAMPLE
In practice
A senior associate pastes a five-page bank reconciliation workpaper for the month just ended into a chat interface and asks the model to summarize outstanding items and the ending reconciled balance.
What came back. The model returned a clean summary with four reconciling items, each citing its line label. The ending reconciled balance was cited correctly from the 'Adjusted bank balance' line. One item - a deposit in transit - was listed with the wrong dollar amount; the model had read a subtotal row rather than the individual item line.
How it was checked. The associate ticked each cited line label against the physical workpaper and found the deposit-in-transit figure on row 14 did not match the model's output, correcting it before the file was signed off.
A constructed example. The prompt is usable as written; the figures show the shape of a result, not a measured one.
WHEN TO USE IT
| WHEN NOT TO
|
WHAT TO TAKE FROM THIS
| Always instruct the model to cite the source row or write 'not found' - it will not do this by default. | |
| Check totals and prior-period figures first; those are where invented numbers hide most easily. | |
| Model confidence in tone is not evidence of accuracy - fluency and correctness are different things. |
SPONSORED
QUESTIONS THIS ANSWERS
Why does an AI model make up numbers that look real?
The model predicts the most statistically likely next token. A number that fits the format and context of surrounding text will be generated even if it does not appear in the source document.
How do I stop an AI from hallucinating figures in a workpaper?
Add two instructions to your prompt: tell the model to cite the exact line a figure comes from, and tell it to write 'not found' rather than infer. Without those constraints, the default is fluent completion, not accuracy.
Is a confident-sounding AI answer more likely to be correct?
No. Confidence in generated text reflects fluency, not accuracy. Professional documents sound certain, so models trained on them produce certain-sounding output regardless of whether the underlying figure is real.
What the accounting job market is actually asking for.
GO DEEPER
Go deeper
IN THIS SERIES
Previously: How to evaluate an AI feature before it touches client data
Next: What a token is and why it sets the price of AI work (coming)
An explainer, not a study: it carries no statistics on purpose. Examples are illustrative.
How accountants are using AI, automation and smarter workflows to close faster, audit cleaner, and free up time for real work.
Accounting Stack · Audit Friendly Data · Accounting & Finance Jobs
Audit Friendly · modernaccounting.ai