An accounting firm runs a promising AI pilot. A task that used to take an hour now produces an output in fifteen minutes. The demonstration is quick, the team is impressed, and the time-saving calculation looks obvious.
Then the work reaches review.
The reviewer has to find the source behind each conclusion, correct inconsistent formatting, investigate false positives, and reconstruct decisions that the system did not explain. A few files go back for rework. Several exceptions require senior judgment. The preparation step became faster, but the complete workflow did not create much usable capacity.
This is the measurement problem at the center of many accounting AI pilots. Gross preparation time is easy to see. Review, correction, exception handling, training, and evidence reconstruction are spread across several people and often disappear into ordinary work. A pilot can therefore report hours saved while shifting cost and risk downstream.
The right question is not only, "How quickly did the system produce something?" It is, "How much reviewable work did the firm accept, at what total human cost, and with what evidence?"
AI adoption is moving faster than AI measurement
The measurement gap is not limited to accounting firms. The 2026 Thomson Reuters AI in Professional Services Report drew on more than 1,500 professionals across legal, tax and accounting, risk, fraud, and government work. It found that 40 percent said their organizations were using generative AI, up from 22 percent the year before. Yet only 18 percent said their organizations tracked return on investment, while another 40 percent did not know whether ROI was measured.
Among the organizations that did measure ROI, internal cost savings and employee usage were much more common than client satisfaction or business outcomes. That is understandable. Time and usage are easier to collect than quality, acceptance, or trust. They are also incomplete.
The 2025 Intuit QuickBooks Accountant Technology Survey offers the optimistic side of the picture. In a vendor-commissioned survey of 700 US accounting professionals, 81 percent said AI had positively affected productivity. The same survey also reported technology friction: firms used an average of eight applications for core operations, 41 percent cited integration difficulty, and 33 percent cited staff training burdens.
Those findings do not cancel each other out. They show why a firm needs its own workflow evidence. AI can produce real gains, while integration, training, and review still determine whether those gains survive in practice.
The unit of value is accepted work
Accounting firms do not deliver generated output. They deliver work that has cleared the firm's quality requirements and approval process.
For a bookkeeping workflow, that might mean a reconciliation that ties to source records, explains open items, contains the required workpaper support, and has been approved by the firm's reviewer. For a cleanup engagement, it might mean a completed batch with traceable proposed classifications, resolved exceptions, and documented review notes. The exact acceptance standard belongs to the firm and will vary by service.
That distinction changes the pilot's unit of measurement. The useful unit is not a prompt, transaction, document, or generated workpaper. It is a comparable work unit that the firm accepts as complete.
One practical measure is:
Accepted throughput = comparable work units accepted ÷ total human hours used
Total human hours should include setup, prompting, source preparation, normal preparation, review, correction, exception handling, and pilot administration. This is an operational measure, not an accounting standard or a universal benchmark. Its purpose is to stop work from vanishing when it moves from one role or stage to another.
If generated output rises but accepted throughput falls, the pilot has not created usable capacity. If accepted throughput rises while evidence quality and error thresholds remain intact, the firm has a stronger reason to continue.
Build the baseline before the pilot begins
A before-and-after comparison is weak when the "before" side is based on memory. Choose one defined workflow and measure a representative baseline before changing it.
The workflow should be narrow enough that the work units are comparable. "Bookkeeping" is too broad. "Monthly bank reconciliation for accounts that meet the firm's standard intake requirements" is more useful. Record the normal mix of simple and difficult files so the pilot is not tested only on clean examples.
Define four things in writing:
- The start and finish. State when the clock begins and what firm acceptance means.
- The quality standard. List required tie-outs, evidence, notes, formatting, and review criteria.
- The test set. Use a representative group of work items, including known exceptions. Remove or protect sensitive data according to the firm's policies and vendor agreement.
- The failure conditions. Identify errors or control failures that require an immediate stop, escalation, or redesign.
The NIST AI Risk Management Framework is broader than accounting and is voluntary, but its measurement logic is useful here. It calls for task scope, human oversight, test sets, metrics, benchmarks, uncertainty, and documented results to be defined rather than assumed. The AICPA and CIMA AI resource center makes the professional boundary clear: AI can improve efficiency and help validate outputs, but it does not replace professional judgment.
The baseline should use the same acceptance standard as the pilot. Lowering the standard after automation begins creates an impressive comparison and a poor control.
Use a review-adjusted pilot scorecard
A serious pilot scorecard needs to show speed, quality, review burden, exceptions, and economics together. None of the measures below should be read in isolation.
| Measure | What to record | Why it matters |
|---|---|---|
| Gross preparation time | Human time before the output enters review | Shows whether the preparation step became faster |
| Total human touch time | Setup, preparation, prompting, review, rework, and exception time | Reveals whether time was removed or merely moved |
| Cycle time | Elapsed time from accepted intake to firm disposition | Captures waiting and handoff effects that touch time misses |
| First-pass acceptance rate | Share accepted without return for correction | Shows whether the output arrives review-ready |
| Reviewer minutes per accepted unit | Review time divided by accepted work units | Makes senior-capacity consumption visible |
| Rework rate and reason | Returned units and coded causes | Distinguishes formatting, evidence, instruction, source, and judgment failures |
| Exception rate and age | Units requiring non-routine handling and time unresolved | Shows how much work escapes the normal path |
| Evidence completeness | Required sources, tie-outs, notes, and lineage present at review | Tests whether the reviewer can verify the work efficiently |
| Accepted throughput | Accepted units divided by total human hours | Connects productivity to the firm's quality gate |
| All-in pilot cost | License, setup, integration, training, review, rework, and administration | Prevents a narrow subscription-only cost comparison |
The most important metric for many firms will be reviewer minutes per accepted unit. Reviewer capacity is scarce, expensive, and easy to hide. If a tool cuts junior preparation by forty minutes but adds thirty minutes of senior review, the firm did not recover forty minutes of equal value. The workflow changed who carried the work.
Evidence completeness deserves equal attention. A correct-looking answer is not necessarily efficient to review. If the reviewer cannot trace an amount, classification, or exception to its source, confidence has to be rebuilt manually. That work is part of the pilot cost even if the output ultimately proves correct.
A fictional example shows how gross savings can mislead
Consider a clearly illustrative pilot involving twenty comparable reconciliation workpapers. The numbers below are hypothetical and are not a benchmark or a Fintant performance claim.
In the baseline, preparation takes 80 hours and review takes 20 hours. Nineteen of the twenty workpapers are accepted on the first pass, and all twenty are ultimately accepted. Total human touch time is 100 hours.
During the pilot, AI-assisted preparation falls to 35 hours. If the firm stops there, it records 45 hours saved. But review rises to 50 hours because source links and explanations are inconsistent. Corrections and exception handling add another 25 hours. Seventeen workpapers are accepted on the first pass, and all twenty are ultimately accepted after correction. Total human touch time is 110 hours.
The preparation step improved dramatically. The workflow used more human time.
That result does not automatically mean the technology should be abandoned. The problem may be fixable through better source packaging, narrower task scope, different instructions, deterministic checks, or a more consistent workpaper format. The point is that the pilot needs to show the real constraint. "AI failed" and "AI saved 45 hours" are both weaker conclusions than "preparation improved, but missing evidence increased review and rework enough to reduce accepted throughput."
That diagnosis tells the firm what to test next.
Decide whether to stop, narrow, redesign, or scale
A pilot should end with a decision, not a general feeling that the tool has potential. Four outcomes are useful.
Stop
Stop the pilot when it creates an unacceptable privacy or security event, a material error pattern, unauthorized changes, missing audit evidence that cannot be reconstructed reliably, or behavior outside the firm's risk tolerance. Some failures should not be averaged against faster files.
Narrow
Narrow the scope when the system performs well on a stable subset but poorly when judgment or irregular source data increases. The useful result may be a smaller use case, such as document ordering, field extraction, or a first-pass matching task, with exceptions routed to people.
This is consistent with broader research on AI-assisted professional work. In a preregistered experiment with 758 consultants, research published in Organization Science found meaningful speed and quality gains on tasks inside the model's capability boundary. On a task outside that boundary, AI users were 19 percent less likely to reach the correct solution. The work was consulting, not accounting, so the result should not be imported as an accounting benchmark. The lesson is narrower: pilot performance can vary sharply by task, which makes scope design part of measurement.
Redesign
Redesign when the output is potentially useful but the handoff is expensive. Common redesign targets include source citations, required fields, workpaper structure, exception categories, reviewer routing, or deterministic validation before human review. Run a new comparison after the change instead of blending it into the original results.
Scale
Scale only when accepted throughput improves, first-pass acceptance and evidence completeness remain within the firm's thresholds, exception and rework patterns are understood, total cost is supportable, and the result holds across a representative sample. Scaling should preserve the accounting firm's professional judgment, final approval, posting, filing, payments, payroll release, and client communication authority.
Do not dismiss hours saved. Put them in the right place.
Time savings are not a bad measure. They are simply one measure at one stage.
Controlled research has found genuine productivity improvements from generative AI. In a preregistered experiment involving 453 college-educated professionals performing bounded writing tasks, Noy and Zhang reported in Science that ChatGPT reduced average task time by 40 percent and increased assessed output quality by 18 percent. That is strong evidence for those studied tasks. It is not proof that every accounting workflow will improve by the same amount.
Practitioner discussion shows the same range of outcomes in less controlled settings. Some accounting and tax professionals describe tools that organize source documents and make review more systematic. Others report that false positives, missing workpapers, or manual verification consume the expected savings. These accounts are useful for identifying questions and vocabulary, not for estimating prevalence.
The fair conclusion is neither that AI always creates capacity nor that review always consumes the gain. The conclusion is that the firm has to measure the complete path from input to acceptance.
A four-week pilot should produce an operating answer
For one defined workflow, compare a representative baseline with a controlled pilot. Record every work unit, stage, human touch, return, exception, and final disposition. Review the data weekly with preparers and reviewers, and keep the test focused on the workflow rather than individual performance.
At the end, the firm should be able to answer five questions:
- Did accepted throughput improve?
- Did reviewer minutes per accepted unit rise or fall?
- Which rework and exception categories changed?
- Could the reviewer trace every accepted conclusion to adequate evidence?
- Was the recovered capacity actually available for other work after all costs were counted?
If the answers are unclear, the pilot has not failed, but it has not earned a scale decision. The next step may be a better test, a narrower scope, a workflow redesign, or a different intervention entirely.
The most credible AI business case in an accounting firm will not begin with the fastest output from a demonstration. It will begin with work that a reviewer can accept, explain, and approve without hidden labor appearing somewhere else.
Sources
- Thomson Reuters: 2026 AI in Professional Services Report
- Intuit: 2025 QuickBooks Accountant Technology Survey
- NIST: AI Risk Management Framework Core
- AICPA and CIMA: AI Resources for Accounting and Finance
- Organization Science: Navigating the Jagged Technological Frontier
- Science: Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence
This article provides general operational education. It is not accounting, tax, legal, security, employment, or technology advice. Each firm should set its own pilot scope, acceptance criteria, data controls, professional review requirements, and approval authority.
