Audit Friendly MODERN ACCOUNTING
EXPLAINER

AI FOR ACCOUNTANTS · PART 11 OF 16

Why the same prompt gives a different answer twice

 

The variation is not a bug or a bad day - it is the mechanism, and ignoring it is how errors slip past review.

Ask an LLM to classify a set of expense lines and run the same prompt twice. You may get two different category assignments for the same transaction. This is not a glitch. It is sampling - the process by which a language model chooses its next word by drawing from a probability distribution rather than always picking the single most likely option. The model is, by design, a generator of plausible continuations, not a calculator of deterministic answers.

The mechanism matters more in accounting than in almost any other field because accounting work depends on repeatability. When a staff accountant posts the same accrual twice in a week, the entries match. When an LLM re-runs a workpaper task twice, they may not. The parameter that controls how much randomness the model introduces is called temperature. At a higher temperature the model ranges more widely across possible outputs; at a lower temperature it converges, though it rarely reaches zero variation entirely. Most hosted tools expose this as a setting, but most default configurations are not set to zero, and even near-zero temperatures can still produce small divergences when the model is choosing between two nearly equal options.

Where this bites hardest is in any workflow that treats a single LLM output as a finished product without a verification step. Consider a reconciliation where a model is asked to match open invoices to bank credits and explain any gaps. Run it Monday and run it again Wednesday - or hand the same file to two team members using the same assistant - and the explanations for the same unmatched item may differ in ways that are hard to audit. The real risk is not that the answer is obviously wrong. It is that two plausible-sounding but inconsistent answers both pass a quick read.

The fix is structural, not stylistic. First, treat LLM output as a first draft that must be checked against the source document, not as a finished conclusion. Second, where consistency is non-negotiable - transaction coding, disclosure language, tick-mark explanations - run the prompt once, record the output, and version it like any other workpaper. Do not re-run and compare; pick one and sign off on it. Third, for any task where the answer should be deterministic, use a deterministic tool. A formula, a rule engine, or a structured extract from a system of record will not drift between runs. Reserve the LLM for the genuinely ambiguous work: drafting explanations, summarizing variance narratives, flagging items that need human judgment. That is the boundary worth drawing.

WORKED EXAMPLE

In practice

A senior accountant is using an LLM assistant to draft explanations for five reconciling items on the bank reconciliation for the month just ended. She pastes the open-item detail into the chat and asks for a plain-English explanation of each gap.

THE PROMPT
Below is a list of five unreconciled items from our bank reconciliation for the month just ended. For each item, write one sentence explaining the most likely reason it has not cleared, based on the description, amount, and counterparty shown. Do not invent facts not present in the data. Flag any item where the description is too ambiguous to explain confidently. Item 1: $4,200 - vendor ACH - Acme Supplies - posted in bank, not in GL Item 2: $750 - check #4471 - issued in GL, not cleared in bank Item 3: $12,000 - wire - unknown counterparty - posted in bank only Item 4: $340 - bank fee - posted in bank, not in GL Item 5: $8,900 - customer payment - ref INV-2241 - posted in GL, not in bank

What came back. The model returned plausible one-sentence explanations for items 1, 2, 4, and 5, and correctly flagged item 3 as too ambiguous to explain without more information. However, for item 2, it described the check as 'likely lost in transit' - a specific conclusion not supported by the data, which only indicates the check has not cleared yet.

How it was checked. The accountant compared each explanation against the bank statement and GL detail side by side, removed the 'lost in transit' characterization for item 2, and replaced it with 'outstanding check not yet presented to bank,' which the source data supports.

A constructed example. The prompt is usable as written; the figures show the shape of a result, not a measured one.

WHEN TO USE IT

 
Drafting variance narratives where tone and clarity matter more than a single correct figure.
 
Summarizing a long client email thread before a planning call.
 
Generating a first-pass list of questions about an unfamiliar account balance.
 
Explaining a reconciling item in plain English for a non-accountant reviewer.

WHEN NOT TO

 
Coding transactions to a chart of accounts where every posting must match a rule.
 
Producing figures that will flow directly into a financial statement without recalculation.
 
Any task where two team members running the same prompt must get identical output.
The pitfall: The model gives a confident, well-written explanation that contradicts what the supporting document actually says - and it reads too smoothly to trigger a second look.

WHAT TO TAKE FROM THIS

 
Record and version the first LLM output; re-running produces drift, not confirmation.
 
Set temperature to its lowest available value for any coding or classification task.
 
Use deterministic tools for deterministic work; LLMs belong on ambiguous drafting tasks.
FREE, EVERY WEEK
Get the next explainer in your inbox
Tuesdays: one Build with AI workflow with the prompt, the steps and the verification rule. Thursdays: the explainer and the week's tools.

SPONSORED

Gusto: full-service payroll that posts clean journal entries to your ledger
Federal, state and local filing in all 50 states, native two-way sync to QuickBooks Online and Xero, and published per-employee pricing with no contract. Audit Friendly scored it 74/100 in a fact-checked review; we earn a referral fee if you sign up, and the score is not affected.

QUESTIONS THIS ANSWERS

Can I just set temperature to zero to fix this?

Lowering temperature reduces variation significantly, but most models can still produce small divergences when two options are nearly equally probable. Near-zero is better; truly zero is not guaranteed.

Does this affect every kind of accounting task equally?

No. Tasks with a single correct answer - matching a payment to an invoice, applying a fixed tax rate - are most exposed because any drift is wrong. Open-ended tasks like drafting a variance explanation have more tolerance for variation.

If two team members run the same prompt and get different answers, which one is right?

Neither output is authoritative by virtue of being an LLM output. Both need to be checked against the source document. The answer that matches the supporting evidence is the one that goes in the workpaper.

SOURCES

Where this comes from

Read the insights wire →

What the accounting job market is actually asking for.

GO DEEPER

Go deeper

Skills that pay → Which named tools move the number
Software in postings → Every tool ranked by live demand
AF Stack Designer → Design your accounting stack in ten minutes
The insights wire → Cross-dataset findings, refreshed hourly
Accounting and finance job board → The Audit Friendly job board, every posting verified live

IN THIS SERIES

Previously: Why structured output beats a paragraph in accounting work

Next: Agents versus automation: the difference that matters for a close (coming)

An explainer, not a study: it carries no statistics on purpose. Examples are illustrative.

How accountants are using AI, automation and smarter workflows to close faster, audit cleaner, and free up time for real work.

Accounting Stack · Audit Friendly Data · Accounting & Finance Jobs

Audit Friendly · modernaccounting.ai