A year ago there was real urgency around AI and the workforce — a sense that something had to be managed right now or you'd be behind. A year later, most enterprises have settled into something calmer. Not because AI stopped mattering; targeted applications like medical diagnostics and crop yield modeling are producing real results against well-bounded problems. But the broad enterprise transformation hasn't arrived on schedule, and I think that's healthy. It gives us time to ask a better question than "how do we deploy more AI."
The better question is this: nearly all of the productivity gains from AI so far have accrued to individuals, not organizations. And almost nobody is measuring the difference.
If AI makes it faster for me to build a deck, that's a genuine personal benefit. I get an hour back. But the organization only benefits if that hour gets redirected to something it cares about. Most of the time nobody's tracking whether it does. You hit "run" on the agent, you get your hour, and what happens to that hour is entirely between you and your calendar.
The gains are real. They're just in the long tail
Here's the honest shape of it. If I used to get through items 1 through 25 on my backlog, now I might reach 35.
Items 26 through 35 were lower priority to begin with. They weren't going to get properly resourced anyway. So is that valuable? Yes — genuinely. But it's value in the long tail, not on your top priorities, which were always going to get done regardless. AI is helping at the margin.
That's a perfectly good outcome. It's just a very different outcome from the one that got budgeted for. If you approved AI spend expecting your top-line commitments to land faster, you bought the wrong thing. They were already landing. What you actually bought was the bottom of the backlog — and whether that's worth the money depends on what's down there.
Which is a task-selection problem, not a technology problem
Once you accept that AI helps at the margin, the interesting work becomes figuring out which margin. The most useful framing I've found comes from my counterparts in our Automation business unit and sorts work along two axes — how often it happens, and how much it varies.
Work Categorization Framework
Matching AI Deployment Strategy to Volume & Variability
High
Volume
Volume
▼
Deterministic Automation
Have AI generate the deterministic process or code, then hand it to traditional automation. AI wrote the program; the program runs like a program.
Hybrid (Agent + Human-in-the-Loop)
The genuinely hard one, and the only true hybrid. Variability suits agents, but volume increases compounding error risk. Requires the agent and a human in the loop.
AI-Written Script
Have AI write the script once with no agent at runtime. If a rare task is always the same, the honest answer is frequently to leave it alone.
Agent Territory
It doesn't happen often enough to justify building a dedicated system, and it's too variable for a rigid script. Ideal fit for autonomous AI agents.
Sorted this way, AI stops looking like a wholesale replacement for headcount and starts looking like a specialty skill you deploy against specific problems — the way you'd bring in a specialist, not restructure the org around them.
Why the human stays in the loop
The reason that fourth quadrant needs a person isn't that the models are bad. It's that they fail differently than software does.
We're accustomed to computers being verifiable. A script either does what it says or it throws an error, and you can read the code to know which. AI systems don't fail that way. They fail by producing something plausible and wrong, and you often can't tell by inspection — you have to check the output against the world. That's a fundamentally different oversight burden, and most organizations haven't staffed for it or budgeted the time it takes.
It's also why a meaningful share of the time people spend "using AI" is actually spent correcting its assumptions. That correction work is real work, and it requires enough expertise to know when the answer is wrong — which is precisely the expertise that gets cut when a company assumes the tool is a replacement.
There's a live argument about whether this ever fully resolves. The AAAI 2025 Presidential Panel surveyed 475 researchers and found 76% considered it unlikely or very unlikely that scaling current approaches would produce general intelligence. Yann LeCun, a Turing Award winner and until recently Meta's chief AI scientist, has gone further, saying at CES 2025 that autoregressive LLMs of the kind we have today simply will not reach human-level intelligence — he has since left to build a company around a different architecture entirely.
You don't have to settle that debate to plan around it, because the enterprise question is narrower than the AGI question and it's already been answered in public. Klarna went furthest fastest: its AI assistant handled 2.3 million chats in its first month, work the company equated to roughly 700 agents. Within about a year it was recruiting human agents again, with CEO Sebastian Siemiatkowski conceding the company had leaned too hard on efficiency and cost and gotten lower quality for it. The AI is still there, still handling roughly two-thirds of inquiries. What changed is that a human is now on the other side of the hard cases. That's not a failure of AI. It's the fourth quadrant asserting itself.
The management question
So the question in front of most leaders isn't whether to deploy more AI. It's two questions they should actually be able to answer:
-
Which of our work sits in the quadrants where AI genuinely helps? Not which work could use AI — which work has the volume-and-variability shape that makes it worth the oversight cost.
-
When AI frees up capacity, where does it go? This is the one nobody's instrumented. Freed-up hours don't announce themselves. They don't show up in a report. They get absorbed into the day, and the organization has no idea whether it just funded a strategic reallocation or a slightly easier Tuesday for everyone.
You can't manage what you can't see - and right now most organizations can see their AI spend perfectly well and their AI return not at all. That visibility is exactly the kind of question ValueOps AI Tokenomics is built to help answer.
Please contact us to continue the conversation and watch a demo.
Frequently Asked Questions
How does AI impact enterprise productivity versus individual productivity?
AI delivers immediate productivity gains to individual workers by accelerating routine tasks and clearing backlog items, but it often fails to automatically translate into organizational ROI. Individual gains usually accumulate at the margin—reaching lower-priority backlog tasks rather than speeding up top strategic goals. Unless leadership tracks and actively reallocates freed-up time, the organization receives minimal top-line benefit.
How should organizations decide where to deploy AI in their workflows?
Organizations should evaluate work based on two key axes—volume and variability—to select the correct deployment model:
-
Low Volume, High Variability: Deployed best using flexible AI agents.
-
Low Volume, Low Variability: Handled by a single AI-generated script that runs predictably.
-
High Volume, Low Variability: Addressed by using AI to write code, which is then handed off to traditional, deterministic automation.
-
High Volume, High Variability: Requires a hybrid approach combining AI agents with direct human oversight to prevent error compounding.
Why is a human in the loop essential for enterprise AI deployment?
A human in the loop is essential because AI models fail by producing plausible yet incorrect outputs rather than throwing standard software errors. Detecting these subtle flaws requires human domain expertise to verify output quality and catch errors before they compound across high-volume workflows.
What is the primary challenge in measuring enterprise AI ROI?
The primary challenge in measuring AI return on investment is tracking where freed-up capacity actually goes. While AI software expenditure is easily visible on balance sheets, saved employee hours quietly absorb into daily calendars, leaving leaders unable to see whether AI spend funded strategic growth or simply eased daily workloads.