|
EXPLAINER |
AI FOR ACCOUNTANTS · PART 9 OF 16
What a token is and why it sets the price of every AI task you run
It is not the task that drives the cost - it is the volume of text the model has to read and write to complete it.
A token is roughly four characters of text - a fragment of a word, sometimes a whole short word. When you send a document to an AI model, the model does not read it the way a person does. It converts everything - your prompt, the document, its own reply - into tokens, processes them, and bills you for the total count. The input tokens are what you sent; the output tokens are what came back. Both sides of that exchange cost money, and the rates differ.
For accounting work, that mechanic matters more than it does in most fields because the documents are long and the data is dense. A multi-page audit workpaper converted to plain text can run to an enormous token count before you have typed a single instruction. A multi-entity bank reconciliation with hundreds of line items is similar. The moment you paste either of those into a prompt, you have already consumed a meaningful chunk of the model's context window - the ceiling on how much it can hold at once - and you have set the floor on what that task will cost. If the model also needs to produce a detailed commentary or a structured output, the output token count adds to the bill.
Some accounting tasks have a naturally low token footprint. Classifying a single transaction, drafting a one-paragraph client email, or checking whether one account description matches a chart-of-accounts category - these stay small because both the input and the expected output are short. The math is different for tasks where the source material is inherently large: reading a full general ledger to spot anomalies, summarizing a thick client file, or extracting data from a dense tax return. Those tasks are not necessarily bad uses of AI, but the cost is proportional to the page count, and 'summarize this' is not a free instruction.
Context window size is the related constraint. Models have a maximum number of tokens they can process in one call. If your document exceeds that limit, you either truncate it - and risk losing the figures that matter - or you split it into chunks and run multiple calls, which multiplies cost and introduces the risk that something important falls in a gap between chunks. A reconciliation that spans two chunks, for instance, might have its opening balance in one call and its closing entry in another, and the model in each call sees only half the picture.
The practical adjustment is to be deliberate about what you paste. Extracting only the relevant columns from a reconciliation before passing it to an AI call, or summarizing a workpaper section manually before asking the model to analyze it, reduces input tokens without reducing output quality - often improves it, because a cleaner input produces a more focused response. It also helps to match task size to model tier: a small, inexpensive model handles transaction classification well; a larger, more expensive one earns its cost only when the reasoning task is genuinely complex. The question to ask before running a workflow is not 'can the AI do this' but 'how many tokens does this require, and does the output justify them.'
WORKED EXAMPLE
In practice
A senior accountant is preparing a month-end bank reconciliation workpaper for a mid-size client. The reconciliation covers this month's activity, exported as a plain-text CSV, and she wants an AI assistant to flag any items older than 30 days that remain uncleared.
What came back. The model returned a filtered list of uncleared items over 30 days, formatted as requested. A handful of the rows were duplicates that had already been flagged as timing differences in a separate note - the model had no way to know that from the CSV alone and listed them anyway.
How it was checked. The accountant compared the model's output list against the 'timing differences' tab in the workpaper and removed the items already documented, then agreed the remainder to the aging schedule.
A constructed example. The prompt is usable as written; the figures show the shape of a result, not a measured one.
WHEN TO USE IT
| WHEN NOT TO
|
WHAT TO TAKE FROM THIS
| Check the token count of your largest recurring documents before committing to an AI workflow. | |
| Tasks that require the model to read a full document cost far more than tasks with short, structured inputs. | |
| Splitting large files across multiple calls saves on context limits but adds cost and creates gap risk. |
SPONSORED
QUESTIONS THIS ANSWERS
What is a token in AI billing?
Roughly four characters of text. Models count every token they receive as input and every token they produce as output, and vendors charge for both.
Why do large accounting documents cost more to process with AI?
Because the full text of the document - every line item, heading and label - is converted to tokens before the model begins working. A long workpaper or reconciliation sets a high input-token floor before any analysis begins.
What happens when a document exceeds the context window?
The model cannot process what it cannot see. You must either truncate the document, risking lost data, or split it into multiple calls, which multiplies cost and can leave critical figures in the gap between chunks.
SOURCES
Where this comes from
What the accounting job market is actually asking for.
GO DEEPER
Go deeper
IN THIS SERIES
Previously: Why a model invents a number and how to stop it
Next: Why structured output beats a paragraph in accounting work (coming)
An explainer, not a study: it carries no statistics on purpose. Examples are illustrative.
How accountants are using AI, automation and smarter workflows to close faster, audit cleaner, and free up time for real work.
Accounting Stack · Audit Friendly Data · Accounting & Finance Jobs
Audit Friendly · modernaccounting.ai