Start a conversation Contact
← Research

Automate the close, not the sign-off

AI can move accounting time from routine classification to client work and quality assurance. The value appears only when uncertain entries stay visible, review effort is measured and a person still signs off the close.

A controller reaches the last week of the month with a queue of invoices, receipts and bank transactions waiting to be classified. An accounting assistant can now suggest a ledger code for each item. The close looks ready to speed up, so the team is tempted to let the system run until the books are closed.

That turns one decision into two. The system can perform a repeatable classification task, while a controller still owns whether the ledger is fit to sign. If those boundaries are not designed together, the organisation gets faster data entry and a harder review problem.

Automate the close, not the sign-off. Let AI classify routine transactions and route uncertainty to an accountable reviewer. Measure the capacity and cycle-time gain alongside corrections and review effort. The commercial value is the client work and quality assurance the team can take on after routine handling moves, provided the control evidence remains intact.

Capacity appears after routine classification moves

There is field evidence for the capacity mechanism, although it is not a promise about every accounting estate. Choi and Xie surveyed 277 accountants and analysed more than 200,000 transaction records from an AI-enabled platform serving 79 small and medium-sized businesses. Greater use of generative AI was associated with more client support, more granular ledgers and a shorter month-end close. The authors also observed time moving away from routine data entry towards business communication and quality assurance. Choi and Xie, Human + AI in Accounting: Early Evidence from the Field

Their reported estimates give a useful pilot hypothesis. A one standard deviation increase in AI use was associated with an 18 per cent increase in weekly client support, with a gap of up to 59 per cent between the lowest and highest use groups. About 9 per cent of accountant time was reallocated, ledger granularity rose by 12 per cent and monthly close time fell by 7.5 days in the observed platform. Stanford Graduate School of Business summary of the study

Those figures describe an association in one platform and a set of small and medium-sized businesses. They do not establish a return for a new vendor or a different chart of accounts. They do identify the thing a CFO or managing partner should measure: whether classification work released enough attention to change client capacity, close timing or quality work.

Confidence should route the review

Classification is a good boundary for selective automation because each suggestion can carry its source transaction, proposed account and confidence. The reviewer does not need to inspect every high-confidence item in the same way. They do need to see what was uncertain, what the system recommended and what changed before sign-off.

Choi and Xie also report a framed field experiment in which AI improved classification accuracy on average, while relying on recommendations that did not match the consensus answer could increase errors. The practical lesson is not to set a universal confidence threshold from the paper. It is to make uncertainty and disagreement part of the queue, then test how much review each bucket needs. The study’s working-paper record

An accounting close that keeps the human boundary visible Fig. 01
  1. 01 Capture Collect the source document, bank line or invoice with its period and entity.
  2. 02 Classify Suggest account, tax treatment and dimensions with confidence and provenance.
  3. 03 Reconcile Match entries and surface missing, duplicated or contradictory evidence.
  4. 04 Review Route low-confidence or non-consensus items to a named accountant with a reason code.
  5. 05 Sign off Keep the close approval with the accountable controller or partner and preserve the audit trail.

Design the exception path before the pilot

The buyer for this decision is a CFO, controller or managing partner who owns close reliability and the team’s client capacity. Start with the current close for one entity or service line. Record how transactions arrive, who classifies them, when reconciliations happen and which evidence a reviewer uses to accept an entry.

Set the automation boundary in writing. Routine items can be auto-classified only when the source is present, the account and dimensions are in the allowed set, and the confidence bucket has passed the pilot rule. A low-confidence item, a new supplier or a non-consensus suggestion should enter an exception queue. The queue needs an owner, due date, reason code and a link back to the source record.

Use a control group or a staged rollout. Compare the AI path with the existing process for the same close periods, entities and transaction types. The point is to learn whether the queue reduces work while preserving accuracy, not to celebrate a faster screen.

The economic mechanism is capacity plus cycle time. If an accountant spends fewer minutes typing routine ledger entries, that time can become client communication, quality assurance or additional accounts. If the review queue grows faster than the classification work disappears, the benefit is only a shift in where the work waits. Read capacity measures with corrected entries, reconciliation breaks and close duration.

Protect the sign-off

An AI suggestion is not an approval. Preserve the original transaction, source document, model version, confidence, reviewer identity and final ledger entry. Do not allow a suggestion to overwrite evidence or erase the reason a person changed it. A controller should be able to reconstruct the path from source to close without relying on a model’s current output.

Ask a supplier to demonstrate the unhappy path. Feed the system a duplicated invoice, a missing receipt, a new supplier and a transaction whose account is ambiguous. Check that each case is held or routed, that the owner can find it, and that a later correction updates the reporting trail. Test what happens when the model or chart of accounts changes during a close period.

The control is especially important when the system disagrees with a reviewer. Record the disagreement and its resolution. If the same category repeatedly needs a person to repair it, narrow the automation boundary or improve the source data before adding more volume.

What the Inference Institute can help decide

The useful engagement begins with the close workflow and the economic constraint. We map the transaction path, identify where evidence is lost, define confidence and exception rules, and specify a pilot that a controller can audit.

The output is a decision pack with a baseline, control group, observation window, review boundary, measurement plan and stopping rule. It makes the intended gain explicit: more clients per accountant, faster reporting, better quality assurance or a shorter close. It also gives the finance owner a way to reject automation that increases correction work.

What this does not tell you

The Choi and Xie field results come from one AI-enabled platform and are primarily observational. Their experiment concerns transaction classification, not every accounting judgement, reporting standard or organisation. The businesses studied are mainly small and medium-sized US firms, so a multinational close or a regulated audit process may have different data, controls and review obligations.

The evidence also does not show that a particular model, threshold or vendor will improve profit. It shows why the boundary is testable. A pilot still needs its own baseline, control period and accountable sign-off. If the organisation cannot preserve provenance or measure corrections, it is not ready to widen the automation.

The controller should make one decision before enabling the feature: which transactions may move without routine review, and what evidence will force a person to intervene. That decision turns an AI bookkeeping suggestion into a measurable operating change. It gives the finance team a chance to buy back capacity while keeping responsibility for the books where it belongs.

Filed under · Method · Accounting · Financial operations · Human oversight Inference Institute · 15 Sept 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.