The AI Productivity Paradox: Why Your Team Isn't Faster Yet
I keep watching teams adopt AI and ship at the same speed. Here's why the bottleneck moved, and the few norms you have to rewrite to catch up.

The AI productivity paradox has a simple operational cause. Your team adopted AI and nothing got faster because AI removed the old bottleneck, making things, and the constraint simply moved downstream to checking, deciding, and coordinating. Most companies upgraded the tools and left the operating model untouched, so every new tool now piles up in front of a bottleneck their processes were never designed for. The fix isn't more AI. It's rewriting the handful of team norms that were built around a constraint that no longer exists.
This is the gap behind one of the most-cited findings of the year. McKinsey's State of AI 2025 survey found that while 78% of organizations now use AI, only about 6% are "high performers" capturing 5% or more EBIT impact from it. What sets those few apart is that they redesign their workflows instead of only deploying the tools. Adoption is nearly universal. Value is not. I run my own company on AI employees, and I watch this gap open in client teams every week. If you lead a mid-market team and you can feel it in your own week, this is why.
What is the AI productivity paradox?
The AI productivity paradox is the growing distance between how productive AI obviously makes an individual and how little of that shows up at the level of the business. The individual gains are real: AI super-users report large personal speed-ups. Yet a U.S. National Bureau of Economic Research survey of roughly 6,000 executives across four countries found that over 80% of firms reported zero measurable productivity gains from AI. MIT's NANDA initiative put it more bluntly in its 2025 report The GenAI Divide: about 95% of generative-AI pilots delivered no measurable impact on profit and loss. The cause was how companies adopted AI. The technology itself worked fine.
Economists have seen this shape before. Stanford's Erik Brynjolfsson has long argued that general-purpose technologies follow a "J-curve": output stalls while organizations do the slow work of restructuring around the new capability, and only later does productivity climb. It's the same pattern Robert Solow named in 1987 about computers, visible everywhere except in the productivity statistics. The lesson both times is identical. The tool arrives fast. The reorganization that makes it pay off arrives slowly, and only if you do it on purpose.
I came to this from a factory floor, not a marketing deck. My background is mechanical engineering and Industry-4.0 automation, and there's a cleaner way to say it, borrowed from manufacturing. In the Theory of Constraints, every process has one binding bottleneck: the slowest step everything else waits on. Relieve it, and the bottleneck does not disappear. It moves to the next-slowest step. AI relieved the most expensive step in knowledge work almost overnight. The constraint moved. And almost no one moved their processes with it.
Where did the bottleneck move, from making to checking?
For most knowledge work, the old bottleneck was making the thing: the first draft, the research pull, the first version of the analysis or the proposal. It was slow and it needed a skilled person, so teams built everything around protecting that person's time, with long planning cycles, tight scoping, and careful role boundaries. Then AI made the making cheap and fast. Not perfect. Cheap and fast. And the constraint slid one step to the right, onto the work that AI doesn't do for you: reading the output critically, catching the confident-but-wrong parts, deciding which version is right for this customer, and putting a name behind it.
The clearest measurement of this comes from software, where AI adoption ran ahead of everyone else. Faros AI's 2025 AI Productivity Paradox Report, analyzing more than 10,000 developers, found that AI-assisted teams completed 21% more tasks and merged 98% more pull requests. Yet pull-request review time ballooned by 91%, because human approval had become the new bottleneck. At the company level, Faros found no correlation between AI adoption and better delivery outcomes. More was produced; the queue just moved to review.
That is the entire paradox in one data point. You speed up the step that was already getting faster and pour more volume onto the step that was already the constraint. This is why "everyone uses AI" and "nothing got faster" are both true at once. The tell is simple. If your team produces more first drafts than before but ships at the same speed, your bottleneck has moved to the review-and-decide step, and it is drowning. It gets worse when the output is plausible but wrong, which is exactly when a human has to read every line. I wrote separately about why AI is confidently wrong and what that costs the reviewer.
Which team norms quietly stop working?
Once you accept that the constraint moved, a whole set of everyday norms reveals itself as out of date. None of them were wrong. They were the right answers to the old question, how do we protect our scarce makers?, and they now mis-fire against the new one. The operative word is quietly. Nothing announces itself as broken. It just gets slower until the workaround becomes the cost. Here are the six I see fail most often.
- Planning. Long roadmaps made sense when committing your makers was expensive. Now a first version is cheap, so the fastest way to resolve a question is often to build the thing and look at it. Plan in shorter, just-in-time horizons; keep a direction, stop scheduling the details of month five.
- Decisions. The old reflex was to book a room and argue. Now generating both options is often faster than the meeting to choose between them. Bring two real versions, then decide. Record that the decision holds, or "whoever ships last wins" becomes a quiet failure mode.
- Review. One person checking everything drowns under the new volume. Split the review: automate the mechanical checks that have a right answer (format, completeness, broken numbers), and reserve human judgment for what needs it, like legal and risk sign-off, security, taste, and whether this is right for this customer.
- Ownership. "Who made this?" stops resolving when everything is AI-assisted. Replace it with the real question underneath: who can fix it, who has the context, and automate the answer where you can.
- Roles. Tight specialist lanes blur as non-specialists produce specialist-grade output. Don't fight it. Hire less for raw throughput, which the tools now supply, and more for the two things they don't: deep expertise where the problem is genuinely hard, and the judgment to tell whether fast, cheap output is any good.
- Leadership. This one breaks the quietest and costs the most. A leader who hasn't personally tried to get real work done with these tools is setting adoption, hiring, and process policy on a mental model that's already out of date. I make myself sit in the review seat for our own agents precisely so I don't legislate from a model I haven't tested.
How do I find my new bottleneck, audit one workflow
You don't fix all six norms this quarter, and you shouldn't try. Start by finding where your bottleneck actually moved, and the cheapest way to do that is to audit a single workflow end to end. Skip the biggest one and pick your noisiest: the most expensive, most repeated, or most quietly dreaded. Then walk it through one honest question. Is this still doing the job it was created to do?
Take the weekly status meeting almost every team has. It was created so people stayed aware of each other's work. Over time it grew, with more people, a longer meeting, everyone half-present until it's their turn to talk. Ask what it's actually for, and the answer is shared awareness. The meeting is just one way to get there. The updates are already written down; a short automated digest can deliver the awareness without the hour. So you reach a gate. If the job still matters but the ritual is a clumsy way to do it, automate the job and kill the ritual. If the job no longer matters, just kill it.
This is not hypothetical. In the engineering-leadership talk that prompted this article, the lead behind Anthropic's Claude Code team described cancelling a 50-person weekly review outright the moment someone asked what it was for and got an honest shrug. Separately, the same team replaced a daily standup that had already decayed into a shared spreadsheet with a short script that assembles the same update automatically. Same instinct each time: find the job, then ask whether the ritual is still the best way to do it. We did exactly this internally. Our own standup became a script, and our morning triage became a routine an AI employee runs before anyone logs on.
There is one rule that makes this safe. Don't automate a workflow you haven't questioned. Automating a broken ritual just makes it run faster and cost less to keep, which means it never dies. Map what the process is for first, decide second what stays deterministic versus what AI can assist versus what stays human, and automate third. Mapping before I build is the difference between rewiring an operating model and just buying more software.
How do I measure whether it's working?
You can't manage what you won't measure, but the obvious metric is a trap. Three signals tell you the operating model is catching up: onboarding time falls (new people get useful faster), cycle time falls (work moves start-to-finish quicker), and the share of work that's AI-assisted by default rises. The second one does double duty. Where cycle time won't fall even though everyone produces faster, you've found your new bottleneck lit up in the data.
Then there's the number to ignore, the one every board seems to want: "what percentage of our work is now done by AI?" It's easy to make that figure big and meaningless. Throughput is not the goal. The customer outcome is. This is exactly the mistake behind the headline failure rates. As UC Berkeley and others noted in responding to the MIT study, a six-month "measurable ROI" test misses value that shows up as avoided hiring, faster cycle times, and higher quality. Measure the end goal, meaning value, speed, and reliability, and treat the AI share as an input rather than an outcome. If you want that share to rise safely, the limiting factor is trust. You only delegate work you'd stop checking, which is why I treat getting agents off your supervision as the real productivity lever.
What does this mean for hiring, scale without hiring?
This is the part most leaders miss, and the one I care about most. When the bottleneck was making things, the only way to add capacity was to add makers, then find them, finance them, onboard them, and manage them. That ceiling is where most service firms stall. But once making is cheap and the constraint is review-and-decide, the lever changes. You add capacity by taking whole repeatable tasks off your team's plate, without piling more headcount onto an over-loaded operating model. That's what I mean by scale without hiring: reaching the next level of output by giving real tasks to an AI employee that works inside the tools you already use, while your scarce humans concentrate on judgment.
The distinction that makes this work is the one I build everything on. You don't need another AI subscription, you need AI that does the job. A chatbot waits for your prompt and hands you a draft you still have to check and ship, so it lives in the review bottleneck with you. An AI employee owns the task: it pulls the data, drafts the digest, posts it in Slack or Teams, and only escalates the cases that need a human. That's the difference between buying a faster maker and actually relieving the new constraint, and done GDPR/GoBD-compliant by construction, it's the discipline a regulated service firm can stand behind.
Frequently asked questions
Why isn't AI improving my company's productivity?
Because AI removed the old bottleneck, producing first drafts and first versions, and the constraint moved downstream to reviewing, deciding, and coordinating, which AI doesn't do for you. If your processes, meetings, and roles were built around the old bottleneck, they now sit in front of the new one and slow you down. McKinsey found 78% of organizations use AI but only about 6% capture meaningful EBIT impact, and the differentiator is workflow redesign over the tools themselves.
What is the AI productivity paradox?
It's the gap between large, obvious individual productivity gains from AI and the near-absence of those gains at the company level. A National Bureau of Economic Research survey of about 6,000 executives found over 80% of firms reported zero measurable productivity gains from AI, and MIT's NANDA initiative found 95% of generative-AI pilots delivered no measurable profit impact. Economists explain it with the "J-curve": a general-purpose technology needs organizational restructuring before its benefits show up.
Does giving a team AI tools automatically make it faster?
No. Tools change individual output; they don't change the operating model. Faros AI found AI-assisted software teams merged 98% more pull requests but saw review time rise 91% as human approval became the new bottleneck, with no firm-level correlation between AI adoption and better outcomes. Speed at the team level comes from rewriting the norms around the new constraint, and the tools alone won't get you there.
What should we measure to know if AI is working?
Track onboarding time and cycle time (both should fall) and the share of work that's AI-assisted (should rise). Avoid the vanity metric of "percentage of work done by AI"; it's easy to inflate and says nothing about value. Measure the end goal: customer value, speed, and reliability. Treat AI usage as an input rather than an outcome.
How do we start rewriting our processes for AI?
Audit one workflow first, your most expensive, repeated, or dreaded one. Ask what job it was created to do, find where work now piles up (making versus checking), decide what's mechanical enough to automate versus what needs human judgment, and then either automate the job or retire the ritual. Map before you build; automating an unquestioned process just makes a broken one cheaper to keep.
How does this let me scale without hiring?
Once making is cheap, adding makers is the wrong lever. You add capacity by handing whole repeatable tasks to an AI employee that runs inside your existing tools, while your people focus on the judgment AI can't supply. That lifts output without finding, financing, and onboarding new headcount, giving you capacity at the hiring ceiling, governed and compliant by construction.
Sources and further reading
- McKinsey & Company, The State of AI 2025 (survey of 1,993 respondents, June to July 2025).
- MIT NANDA initiative, The GenAI Divide: State of AI in Business 2025 (reported by Fortune, August 2025).
- Faros AI, AI Productivity Paradox Report 2025 (analysis of 10,000+ developers, June 2025).
- National Bureau of Economic Research, survey of ~6,000 executives across four countries (reported by Fortune, 2025).
- Erik Brynjolfsson (Stanford), the productivity J-curve and the Solow paradox.
- Eliyahu Goldratt, The Goal (Theory of Constraints).
I build this for a living, and I run my own company on it: roughly nine AI employees handle the repeatable work so my team spends its hours on judgment. If you want to find where your bottleneck actually moved and which tasks an AI employee could own next, that's exactly what I do. See how managed AI employees work or tell me about your noisiest workflow.
Christoph Sauerborn is the founder of Brixon AI. He builds AI employees for capacity-constrained service firms, and runs his own agency on them. Mechanical engineer by training (RWTH Aachen), former Industry 4.0 engineer at Bosch. More about how I work.