Before You Sign Off on the AI Budget, Read What the Research Actually Shows
- Johan Steyn

- Jul 10
- 7 min read
Goldman Sachs found that 95 per cent of organisations are getting zero return on AI pilots. PwC found that 74 per cent of AI's economic value goes to 20 per cent of organisations. The NBER documented a productivity paradox where perceived gains are larger than measured gains. Five questions can protect your board from the gap between what vendors claim and what independent researchers find.

Sign up for my Substack daily AI newsletter here.
See my AI Training course portfolio for corporate Business Leaders here.
Follow me on LinkedIn: https://www.linkedin.com/in/johanosteyn/
At some point in the last twelve months, many South African boards approved an AI investment. The justification almost certainly included at least one of the following: McKinsey’s projection that generative AI could add 4.4 trillion dollars annually to the global economy, Goldman Sachs’s estimate that AI could raise global GDP by 7 per cent, or Microsoft’s finding that Copilot users complete tasks 29 per cent faster. These numbers are real. They are cited from real research published by real institutions. They are also, in the context of most board investment decisions, technically true and strategically misleading.
The most important AI governance question South African directors are not yet asking is not whether AI works. It does, in specific contexts, for specific tasks, under specific conditions. The question is whether the research that underpins the investment case presented to your board was measuring the right thing — and whether the gap between what the research measured and what your organisation needs to produce constitutes a basis for the capital allocation being proposed.
CONTEXT AND BACKGROUND
The productivity promise made on behalf of AI is the most important and least scrutinised claim in the AI commercial narrative. It rests on a methodological foundation that most board presentations do not disclose, and most directors have not examined.
The first layer of the foundation is the task level versus the enterprise level. Most AI productivity studies measure speed on narrow, well-defined tasks in controlled conditions. A legal researcher summarising a document faster. A software developer generating code faster. A customer service agent resolving a query faster. These are real gains at the task level. The question that most investment cases do not answer is whether task-level speed translates to enterprise-level outcomes — revenue growth, verified cost reduction, competitive advantage — at the scale and in the timeframe that justifies the investment being proposed.
The DX Longitudinal Study, which tracked software engineering teams across more than 400 companies from November 2024 through February 2026, provides the most precise available answer to that question for one specific domain. During the study period, AI tool usage increased by an average of 65 per cent. Median pull request throughput — the actual measure of what teams shipped — increased by 7.76 per cent. The gap between 65 per cent and 7.76 per cent is not a measurement anomaly. It is the answer to a specific governance question: when individual task speed increases dramatically, does business output increase at the same rate? The study’s explanation for the gap is equally instructive — writing code was never the bottleneck. The bottlenecks were scoping, architectural decisions, cross-team alignment, and the human judgment calls that AI tools do not currently touch.
The second layer of the foundation is the vendor funding question. Microsoft’s 29 per cent faster finding was produced by Microsoft. Anthropic’s estimate that Claude could increase US labour productivity growth by 1.8 per cent over the next decade was produced by Anthropic, who acknowledged in their own paper that the estimates likely overstate current productivity effects. McKinsey’s 4.4 trillion dollar projection was produced by McKinsey — a firm with a direct commercial interest in the consulting engagements that AI adoption generates.
The most rigorous independent assessments reach consistently more modest conclusions. The San Francisco Federal Reserve’s February 2026 economic letter states directly that most macro studies find limited evidence of a significant AI effect and that even firms that say AI is useful find little evidence of transformative gains.
The third layer is what Goldman Sachs’s own research found when it looked past its headline projection. In May 2026 Goldman cited the MIT Labs finding that 95 per cent of organisations are getting zero return on their AI pilots. A 2025 EY survey found that 99 per cent of companies in its sample reported financial losses due to AI-related risks, averaging 4.4 million dollars per company. Goldman’s own AI economist Jim Covello said in a June 2026 interview that he is as or more sceptical on the economics of AI today as he was before — despite acknowledging that the technology itself is incredible.
INSIGHT AND ANALYSIS
The most intellectually honest assessment of the AI productivity landscape available in 2026 comes from a surprising source. Michael Jacobides, Professor of Strategy and Entrepreneurship at London Business School, said at a parliamentary briefing: “If you read LinkedIn and listen to the excitement on the corporate side, you would think that we’re already living on another planet, that the improvements are massive. But if you start looking at productivity, and if you speak behind closed doors at C-suite levels, people will whisper — has anyone really seen the benefits yet? The dirty answer is mostly not. There are very few things we can see as serious end-to-end transformations.”
The NBER working paper published in March 2026, drawing on a survey of nearly 750 corporate executives, documents what it calls the productivity paradox: perceived productivity gains are systematically larger than measured productivity gains, likely reflecting a delay between individual task efficiency and business-level revenue realisation. The executives believe AI is making them more productive. The revenue data has not yet confirmed that belief.
PwC’s April 2026 AI Performance Study of 1,217 senior executives across 25 sectors globally found that 74 per cent of AI’s economic value is captured by just 20 per cent of organisations. The organisations in that 20 per cent are not simply deploying more AI tools. They are using AI as a catalyst for business reinvention — redesigning processes, pursuing new revenue opportunities, and building organisational AI capability that goes well beyond tool deployment. The remaining 80 per cent are stuck in pilot mode.
The distinction PwC identifies between the productive 20 per cent and the pilot-mode 80 per cent is the most practically useful finding in the 2026 AI productivity literature for South African boards. The separating factor is not which AI tools an organisation uses. It is whether the organisation has redesigned its processes, built its data foundations, and invested in the organisational capability that turns task-level speed gains into enterprise-level business outcomes. The DX finding that writing code was never the bottleneck applies more broadly: in most organisations, the tasks that AI makes faster are not the tasks that are limiting business performance. The bottlenecks are the human judgment, process design, and organisational alignment questions that no AI tool currently addresses.
The Goldman Sachs 7 per cent GDP figure deserves specific examination because it appears in more South African board presentations than any other AI productivity claim. It is a ten-year long-run projection dependent on perfect, ubiquitous enterprise adoption across the global economy. It is not a near-term corporate ROI metric. It is not a South African economic projection. And it has since been qualified by its own authors, with Jim Covello saying in June 2026 that the economics remain very much in question and that most enterprises have yet to generate returns from AI spending. The number is real. Its application to a board-level investment case for a South African organisation is a category error.
IMPLICATIONS
The five questions that follow from the 2026 productivity research literature are not a barrier to AI investment. They are the governance discipline that distinguishes AI investment that produces business outcomes from AI investment that produces faster task completion and an ongoing pilot.
The first question is what the baseline is. Every productivity claim requires a comparison point. A 29 per cent improvement over what, measured how, in what conditions, by whom? In the original vendor baseline studies, the speedup was typically calculated for workers executing isolated micro-tasks under controlled conditions, with no compliance requirements, no security review obligations, no legacy software dependencies, and no cross-team coordination burden. The comparison between that controlled baseline and real-world deployment in a South African financial services, legal, or professional services organisation — where compliance, review, and coordination are precisely the tasks that consume the most time — is already suspect before any other methodology question is raised.
The second question is who funded the study. Vendor-funded research consistently produces larger productivity estimates than independent research. The San Francisco Federal Reserve, the NBER, and Goldman Sachs’s own independent research produce more modest findings than the research Microsoft, McKinsey, and Anthropic produce about their own tools. Asking who funded the study is the cheapest and most effective governance tool available to a board evaluating an AI investment case.
The third question is what the organisational throughput metric is — not the individual speed metric. DX found that 65 per cent more AI tool usage produced 7.76 per cent more output. Your investment case should specify in advance how organisational output will be measured, not just how individual task speed will be measured. If the investment case cannot name a measurable organisational outcome, it is a pilot, not an investment.
The fourth question is what process redesign accompanies the tool deployment. PwC’s finding that the productive 20 per cent redesign their processes rather than deploying tools on top of existing ones is the most important governance insight in the 2026 literature. If the investment case does not include a process redesign component, the board is approving a tool deployment that the evidence suggests is unlikely to produce enterprise-level returns.
The fifth question is what the measurement window is and who will verify the outcome. The NBER productivity paradox reflects a delay between perceived and measured gains. If the investment case has no mechanism for independent measurement of business outcomes against the claims in the original case, the board has no basis for evaluating whether the investment was justified — and no accountability structure for the executive team that proposed it.
CLOSING TAKEAWAY
The AI productivity promise is not a lie. The DX study found real gains — just smaller than vendor claims and concentrated in the tasks that were already the easiest to accelerate. The PwC study found real value creation — just concentrated in the 20 per cent of organisations that redesigned their businesses rather than deploying tools. The Goldman Sachs 7 per cent GDP projection may eventually prove correct — just not in the timeframe or on the investment terms most boards are approving today.
The governance response to the productivity paradox is not scepticism about AI. It is the same rigour applied to AI investment decisions that South African boards apply to every other capital allocation — clear baselines, independent evidence, measurable business outcomes, process redesign requirements, and accountability for the results. The research published in the first half of 2026 has provided the questions. The boards that ask them before signing off on the AI budget will be better positioned than those that ask them after.
Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at https://www.aiforbusiness.net



Comments