The bottleneck often moves.
If AI makes one step of a workflow 30% faster, did the business become 30% more productive?
Usually, we don’t know.
That question kept coming up across cases in data, retail, telecom, engineering, and internal agent systems.
The productivity gains were often real. Faster classification. Faster triage. More forecasts. More coverage. More automated activity.
What happened next was more interesting.
Quality became harder to protect. Another team inherited the bottleneck. Aggregate accuracy looked better than the accuracy at the point where someone actually made a decision. Automated agents completed plenty of steps without completing the work.
Across the cases, one measurement rule kept becoming useful:
When AI improves one part of a workflow, measure what happens around it.
That matters as companies move from proving that AI can save time to deciding whether those gains improve operating performance.
Faster can create a quality problem
At Unidata, Head of Data Collection Kirill Meshyk described one of the clearest trade-offs.
AI pre-labeling can reduce the amount of work required from human annotators. That creates an obvious productivity gain.
But the machine suggestion also becomes an anchor.
Once an annotator sees a proposed label, they are no longer evaluating the data from a clean starting point. A wrong suggestion can influence the review itself.
Meshyk’s warning was direct:
Productivity gains “can be misleading if the quality aspect is not taken into account.”
This is particularly important because AI productivity is often measured through throughput.
Items processed per hour. Cases reviewed. Drafts created. Records classified.
Those numbers can improve while the quality of the underlying work deteriorates.
The external evidence points in the same direction. Stanford’s 2026 AI Index found some of the strongest measured productivity gains in structured work where outputs are relatively easy to monitor. Gains become less straightforward as tasks require deeper reasoning or become harder to evaluate.
Speed matters.
But speed without a quality gate can simply make errors cheaper to produce.
The bottleneck doesn’t necessarily disappear
Joel Goldstein saw a different effect at Mr. Checkout.
AI helped reduce the cost and inconsistency of triaging incoming consumer brands before they were considered for distributor referrals.
That worked.
Then another part of the system became constrained.
“The bottleneck moved, it didn’t disappear.”
Once intake became easier, distributor evaluation capacity became relatively more scarce.
That is an operating problem many AI ROI calculations miss.
A workflow is a sequence.
If one stage previously handled 50 cases per day and AI suddenly allows it to process 200, the business only receives the full value if the next stage can absorb the additional 150.
Otherwise, the company has created local efficiency without equivalent system-level capacity.
This changes the question leadership should ask.
Instead of:
How much faster did this task become?
Ask:
What happened to the work immediately afterward?
That question may reveal where the actual investment needs to go next.
Accuracy can look better at the wrong level
Priyank Jain described another measurement problem from forecasting work supporting a retail network of roughly 3,000 stores.
The system could produce approximately 97% accuracy at the network level.
At individual stores, accuracy was closer to 85%.
His response was more useful than either number:
“Nobody decides anything at network level.”
Staffing decisions happen at stores.
Local marketing decisions happen in markets.
Site decisions happen at specific locations.
The aggregate metric can therefore look excellent while masking weaker performance at the level where the organization actually acts.
This is a familiar analytical mistake, but AI makes it easier to reproduce because model performance creates a tempting headline number.
Averages are useful for understanding the system.
They are not necessarily useful for evaluating the decision.
For every AI metric, the operating question should be:
At what level does someone actually use this output to do something?
That is where performance needs to hold.
Activity is not completion
The same problem becomes even more visible with AI agents.
Michael Sebastian of Branded Mayhem shared the results of 16 internal agent runs.
Eight expired.
Six failed.
Two were canceled.
Zero completed.
If the organization had measured approvals, initiated tasks or agent activity, the system might have appeared busy.
It wasn’t accomplishing the intended work.
That distinction is likely to become more important as companies introduce autonomous and semi-autonomous agents into operating workflows.
Sebastian separates:
approval → attempted work → observed completion
Those are three different states.
The system does not get credit because it tried.
That sounds obvious until automation makes activity almost free.
AI can produce thousands of classifications, messages, recommendations or attempted actions. Volume itself becomes less informative precisely because the marginal cost of producing activity has collapsed.
The metric has to move closer to consequence.
Did the work complete?
Did the state change?
Did the customer receive what they needed?
Did the next system accept the output?
Did revenue, cost, quality, or cycle time improve?
Otherwise, automation can generate a tremendous amount of operational theater.
Throughput can reward the wrong behavior
The measurement problem becomes more serious when mistakes have larger consequences.
Rohit Shinde of Atlas Prediction Control works on AI applications in process engineering, including systems designed to identify potential problems in engineering drawings and check them against applicable standards.
Simply counting drawings processed would make the AI look productive.
Shinde rejected that metric.
“Throughput rewards the wrong behavior.”
The more meaningful measure was the number of findings that experienced reviewers agreed were legitimate.
That moves the metric from work produced toward useful work accepted.
And in safety-critical environments, the distinction is not academic.
A false negative can matter considerably more than a faster review.
The same principle applies in less consequential settings.
A marketing system should not be rewarded for producing 100 campaigns.
A research system should not be rewarded for producing 500 insights.
A sales agent should not be rewarded for sending 10,000 messages.
The metric should reflect the result the organization actually wanted from the work.
Four questions for an AI productivity claim
The cases suggest a useful way to assess AI productivity without complicating the evaluation.
| Question | |
|---|---|
| Gain | What became faster, cheaper, more consistent, or more scalable? |
| Trade-off | What could become weaker, riskier, or less reliable? |
| Constraint | Where did the bottleneck move? |
| Outcome | Did the business result actually improve? |
This does not make AI projects harder to justify. It makes the productivity claim more precise.
There is strong evidence that AI can improve productivity. In a 2023 field study of 5,179 customer support agents, Brynjolfsson, Li and Raymond found a 14% average increase in issues resolved per hour, with considerably larger gains among less experienced workers.
There is also evidence that performance can deteriorate when people rely on AI in tasks outside the system’s effective capability range. In research involving BCG consultants, AI substantially improved speed and quality on tasks within the tested frontier, while participants using AI were 19 percentage points less likely to reach the correct answer on a task deliberately placed outside it.
Productivity claims become more useful when the unit of measurement is clear.
Task productivity.
Workflow productivity.
Business outcome.
They are related, but they are not interchangeable.
The operating review should go one step further
For teams implementing AI now, the useful conversation is no longer simply:
Did the automated step improve?
Add four questions:
- What happened immediately before it?
- What happened immediately after it?
- Where did capacity, quality, or risk move?
- Which metric reflects the decision or outcome that actually matters?
That turns AI productivity from a tool metric into an operating question.
And it makes the next investment easier to see.
Sometimes the answer will be more automation.
Sometimes it will be additional capacity downstream.
Sometimes better validation.
Sometimes a different metric.
And sometimes the productivity gain will turn out to matter much less than it initially appeared.
The gain is worth measuring.
So is everything it changes.
If you’re evaluating an AI implementation now, take one workflow and ask where the constraint moved. If the answer isn’t obvious, email us.
Sources
The operator findings and quotations in this article come from Praxable’s August 2026 research into AI inside real operating workflows.
Supporting research:
- Stanford HAI — 2026 AI Index, Economy: productivity, organizational adoption and variation by task. Read the report
- Brynjolfsson, Li & Raymond — Generative AI at Work: field evidence from 5,179 customer-support agents. Read the NBER paper
- Dell’Acqua et al. — Navigating the Jagged Technological Frontier: experimental evidence on AI productivity and performance across different knowledge-work tasks. Read the Organization Science paper

