|
EXPLAINER |
AI FOR ACCOUNTANTS · PART 11 OF 16
Why the same prompt gives a different answer twice
The variation is not a bug or a bad day - it is the mechanism, and ignoring it is how errors slip past review.
Ask an LLM to classify a set of expense lines and run the same prompt twice. You may get two different category assignments for the same transaction. This is not a glitch. It is sampling - the process by which a language model chooses its next word by drawing from a probability distribution rather than always picking the single most likely option. The model is, by design, a generator of plausible continuations, not a calculator of deterministic answers.
The mechanism matters more in accounting than in almost any other field because accounting work depends on repeatability. When a staff accountant posts the same accrual twice in a week, the entries match. When an LLM re-runs a workpaper task twice, they may not. The parameter that controls how much randomness the model introduces is called temperature. At a higher temperature the model ranges more widely across possible outputs; at a lower temperature it converges, though it rarely reaches zero variation entirely. Most hosted tools expose this as a setting, but most default configurations are not set to zero, and even near-zero temperatures can still produce small divergences when the model is choosing between two nearly equal options.
Where this bites hardest is in any workflow that treats a single LLM output as a finished product without a verification step. Consider a reconciliation where a model is asked to match open invoices to bank credits and explain any gaps. Run it Monday and run it again Wednesday - or hand the same file to two team members using the same assistant - and the explanations for the same unmatched item may differ in ways that are hard to audit. The real risk is not that the answer is obviously wrong. It is that two plausible-sounding but inconsistent answers both pass a quick read.
The fix is structural, not stylistic. First, treat LLM output as a first draft that must be checked against the source document, not as a finished conclusion. Second, where consistency is non-negotiable - transaction coding, disclosure language, tick-mark explanations - run the prompt once, record the output, and version it like any other workpaper. Do not re-run and compare; pick one and sign off on it. Third, for any task where the answer should be deterministic, use a deterministic tool. A formula, a rule engine, or a structured extract from a system of record will not drift between runs. Reserve the LLM for the genuinely ambiguous work: drafting explanations, summarizing variance narratives, flagging items that need human judgment. That is the boundary worth drawing.
WORKED EXAMPLE
In practice
A senior accountant is using an LLM assistant to draft explanations for five reconciling items on the bank reconciliation for the month just ended. She pastes the open-item detail into the chat and asks for a plain-English explanation of each gap.
What came back. The model returned plausible one-sentence explanations for items 1, 2, 4, and 5, and correctly flagged item 3 as too ambiguous to explain without more information. However, for item 2, it described the check as 'likely lost in transit' - a specific conclusion not supported by the data, which only indicates the check has not cleared yet.
How it was checked. The accountant compared each explanation against the bank statement and GL detail side by side, removed the 'lost in transit' characterization for item 2, and replaced it with 'outstanding check not yet presented to bank,' which the source data supports.
A constructed example. The prompt is usable as written; the figures show the shape of a result, not a measured one.
WHEN TO USE IT
| WHEN NOT TO
|
WHAT TO TAKE FROM THIS
| Record and version the first LLM output; re-running produces drift, not confirmation. | |
| Set temperature to its lowest available value for any coding or classification task. | |
| Use deterministic tools for deterministic work; LLMs belong on ambiguous drafting tasks. |
SPONSORED
QUESTIONS THIS ANSWERS
Can I just set temperature to zero to fix this?
Lowering temperature reduces variation significantly, but most models can still produce small divergences when two options are nearly equally probable. Near-zero is better; truly zero is not guaranteed.
Does this affect every kind of accounting task equally?
No. Tasks with a single correct answer - matching a payment to an invoice, applying a fixed tax rate - are most exposed because any drift is wrong. Open-ended tasks like drafting a variance explanation have more tolerance for variation.
If two team members run the same prompt and get different answers, which one is right?
Neither output is authoritative by virtue of being an LLM output. Both need to be checked against the source document. The answer that matches the supporting evidence is the one that goes in the workpaper.
SOURCES
Where this comes from
What the accounting job market is actually asking for.
GO DEEPER
Go deeper
IN THIS SERIES
Previously: Why structured output beats a paragraph in accounting work
Next: Agents versus automation: the difference that matters for a close (coming)
An explainer, not a study: it carries no statistics on purpose. Examples are illustrative.
How accountants are using AI, automation and smarter workflows to close faster, audit cleaner, and free up time for real work.
Accounting Stack · Audit Friendly Data · Accounting & Finance Jobs
Audit Friendly · modernaccounting.ai