Automate the close, not the sign-off
AI-assisted classification may release accounting capacity, but review and correction work determine the result. Test routine handling against the existing close process while preserving evidence and accountable sign-off.
Keep development and test groups separate before comparing performance. No measured results are shown.
Routine transaction classification can consume attention needed for reconciliations, client work and close review. An AI assistant may help organise that workload, but a faster entry process does not by itself establish that the ledger is ready to sign.
Separate permission to classify a transaction from authority to approve the close. The first can be bounded by source evidence and a tested routing rule. The second remains an accountable judgement about the completed record and unresolved exceptions.
Test routine classification while retaining accountable close approval. Measure handling time, review effort, corrections and completed work together. The opportunity is additional useful capacity or a shorter close, provided the organisation preserves the evidence needed to review its accounts.
Capacity appears after routine classification moves
There is field evidence for the capacity mechanism, although it is not a promise about every accounting estate. Choi and Xie surveyed 277 accountants and analysed more than 200,000 transaction records from an AI-enabled platform serving 79 small and medium-sized businesses. Greater use of generative AI was associated with more client support, more granular ledgers and a shorter month-end close. The authors also observed time moving away from routine data entry towards business communication and quality assurance. Choi and Xie, Human + AI in Accounting: Early Evidence from the Field
Their reported estimates give a useful pilot hypothesis. A one standard deviation increase in AI use was associated with an 18 per cent increase in weekly client support, with a gap of up to 59 per cent between the lowest and highest use groups. About 9 per cent of accountant time was reallocated, ledger granularity rose by 12 per cent and monthly close time fell by 7.5 days in the observed platform. Stanford Graduate School of Business summary of the study
The findings describe an association in one platform serving small and medium-sized businesses. They do not establish the return from another vendor or chart of accounts. Use them to define a local test of whether classification reduces total handling effort and improves the work the finance team can complete.
Confidence should route the review
Classification is a good boundary for selective automation because each suggestion can carry its source transaction, proposed account and confidence. The reviewer does not need to inspect every high-confidence item in the same way. They do need to see what was uncertain, what the system recommended and what changed before sign-off.
Choi and Xie also report a framed field experiment in which AI improved classification accuracy on average, while relying on recommendations that did not match the consensus answer could increase errors. The practical lesson is not to set a universal confidence threshold from the paper. It is to make uncertainty and disagreement part of the queue, then test how much review each bucket needs. The study’s working-paper record
- 01 Capture Collect the source document, bank line or invoice with its period and entity.
- 02 Classify Suggest account, tax treatment and dimensions with confidence and provenance.
- 03 Reconcile Match entries and surface missing, duplicated or contradictory evidence.
- 04 Review Route low-confidence or non-consensus items to a named accountant with a reason code.
- 05 Sign off Keep the close approval with the accountable controller or partner and preserve the audit trail.
Design the exception path before the pilot
The buyer for this decision is a CFO, controller or managing partner who owns close reliability and the team’s client capacity. Start with the current close for one entity or service line. Record how transactions arrive, who classifies them, when reconciliations happen and which evidence a reviewer uses to accept an entry.
Set the automation boundary in writing. Routine items can be auto-classified only when the source is present, the account and dimensions are in the allowed set, and the confidence bucket has passed the pilot rule. A low-confidence item, a new supplier or a non-consensus suggestion should enter an exception queue. The queue needs an owner, due date, reason code and a link back to the source record.
Use a control group or a staged comparison covering comparable periods, entities and transaction types. Count the work transferred to reviewers and later corrections, as well as the entries processed automatically. Agree the acceptance criteria before observing the result.
The economic mechanism is capacity plus cycle time. If an accountant spends fewer minutes typing routine ledger entries, that time can become client communication, quality assurance or additional accounts. If the review queue grows faster than the classification work disappears, the benefit is only a shift in where the work waits. Read capacity measures with corrected entries, reconciliation breaks and close duration.
Protect the sign-off
An AI suggestion is not an approval. Preserve the original transaction, source document, model version, confidence, reviewer identity and final ledger entry. Do not allow a suggestion to overwrite evidence or erase the reason a person changed it. A controller should be able to reconstruct the path from source to close without relying on a model’s current output.
Ask a supplier to demonstrate the unhappy path. Feed the system a duplicated invoice, a missing receipt, a new supplier and a transaction whose account is ambiguous. Check that each case is held or routed, that the owner can find it, and that a later correction updates the reporting trail. Test what happens when the model or chart of accounts changes during a close period.
The control is especially important when the system disagrees with a reviewer. Record the disagreement and its resolution. If the same category repeatedly needs a person to repair it, narrow the automation boundary or improve the source data before adding more volume.
What the Inference Institute can help decide
The Institute can assess the close workflow and the constraint the finance owner wants to address. Architecture and risk advice can identify where source evidence is lost, specify confidence and exception boundaries, and define a pilot the client’s delivery team can operate and the controller can review.
The output is a decision pack with a baseline, control group, observation window, review boundary, measurement plan and stopping rule. It makes the intended gain explicit: more clients per accountant, faster reporting, better quality assurance or a shorter close. It also gives the finance owner a way to reject automation that increases correction work.
What this does not tell you
The Choi and Xie field results come from one AI-enabled platform and are primarily observational. Their experiment concerns transaction classification, not every accounting judgement, reporting standard or organisation. The businesses studied are mainly small and medium-sized US firms, so a multinational close or a regulated audit process may have different data, controls and review obligations.
The evidence also does not show that a particular model, threshold or vendor will improve profit. It shows why the boundary is testable. A pilot still needs its own baseline, control period and accountable sign-off. If the organisation cannot preserve provenance or measure corrections, it is not ready to widen the automation.
Before enabling routine classification, the controller should approve the eligible transactions and the evidence that triggers intervention. Review the resulting capacity, correction burden and close quality over the stated observation window. Expand only where the comparison supports an operating benefit with a usable audit trail.