Almost every conversation about AI for business starts backwards: someone brings in a tool and then goes looking for a place to put it. The order that works is the reverse: which process hurts, which number has to move, and only then which technology applies. That has an uncomfortable consequence: half the fashionable use cases do not survive that question.
88%
of respondents say their organization uses AI regularly in at least one business function 1
39%
of respondents identify some EBIT impact at the enterprise level, and most put it below 5% of EBIT 1
75%
of the potential value estimated by McKinsey would be concentrated in just four functions 3
2.1%
of the Mexican economic units that use digital tools report using AI systems 10
Read them together, with one warning about method: they are not comparable bases. The first three figures are self-reported by executives in global surveys, or estimates of theoretical potential; the Mexican one is a census of economic units, most of them micro businesses, and it cannot be subtracted from the others. Even so, the picture reads clearly: reported usage is everywhere, impact that anyone declares measurable is in very few places, potential value concentrates in a handful of functions, and AI adoption in Mexico is still marginal. None of that means AI does not work: it means the difference between the companies that capture value and the ones that do not is rarely in the use case they picked, but in whether anyone redesigned the process around it.
Where does the value of AI in a company actually concentrate?
In 2023, McKinsey analyzed 63 generative AI use cases across 16 business functions and estimated that, applied broadly, they could add the equivalent of $2.6 trillion to $4.4 trillion a year to the global economy. The useful part is not the amount (it is theoretical potential, not captured value) but the concentration: around 75% of that value would land in four areas, customer operations, marketing and sales, software engineering and R&D 3.
That concentration is your first filter. The second one comes from Stanford HAI's 2026 AI Index: the productivity gains reported by the studies are larger in structured, measurable work, where the output is easy to monitor 4. Translated: the cases that work are the ones where someone can say, without argument, whether an individual output was right or wrong.
Where the work is structured and the result easy to verify, AI delivers. Where judgment is fuzzy, it delivers in the demo and not in the operation.
Key
Sales and marketing: 4 use cases
Marketing and sales is one of the four functions where potential value concentrates 3 and also one of the worst measured: it is easy to prove that more content got produced and hard to prove that more got sold.
Customer service: the 2 cases with the strongest evidence
If I had to pick one function to start with in a midsize company, it would be this one, and not because it is fashionable: it is where the causal data is best. Brynjolfsson, Li and Raymond measured the staggered rollout of a conversational assistant across 5,172 agents at a Fortune 500 business process software company: 15% more issues resolved per hour on average 6. That is not executive self-report. Before extrapolating it, it helps to know the context: a single company, close to 80% of the agents based in the Philippines, and a tool built on GPT-3 6.
Careful
Finance, operations, HR and legal: 7 internal use cases
They never make it into the slide decks because they are not glamorous, and they produce the savings that are easiest to defend in front of a CFO: the process is already instrumented, there is volume, and there is a cost per transaction that someone already calculates.
Finance
Operations and supply chain
Human resources and legal
Software engineering: 2 use cases
This is the function with the most published data and the easiest one to overstate. The most cited randomized experiment, with 95 developers recruited on Upwork, found that those given access to GitHub Copilot completed the task 55.8% faster than the control group 7. Before you take that number to a committee: it was a narrow, synthetic task (an HTTP server in JavaScript), the result is conditional on having completed it, the 95% confidence interval runs from 21% to 89% and it is a preprint by authors from GitHub and Microsoft.
Measurements on real work give more modest figures. The AI Index picks up a study (Cui et al., 2025) with 26% more pull requests completed by developers using Copilot, and the counterpoint: METR found experienced developers were 19% slower with AI assistance on deep reasoning tasks, though it later failed to replicate the result 4. Google's DORA 2025 survey, with nearly 5,000 professionals globally, closes the picture: 90% report using AI in their work and more than 80% believe they gained productivity, but 30% report little or no trust in the generated code 8.
Which functions have measured evidence and which are still a promise
This is the table I use to structure a prioritization conversation. It does not say which case is best for you: it says how much confidence you can put behind the number you are about to promise.
| Function | Best supported case | Evidence | What usually goes wrong |
|---|---|---|---|
| Customer service | Copilot for agents | High. Staggered field study with 5,172 real agents: +15% in issues resolved per hour 6. | Speed gets measured and quality gets ignored; the most skilled agents showed small but statistically significant declines in resolution and satisfaction 6. |
| Software engineering | Code assistant | Medium-high. Randomized experiment on a narrow task, in a preprint by authors from GitHub and Microsoft 7, plus a self-reported global survey of nearly 5,000 professionals 8. | It raises throughput and delivery instability along with it 8; 30% report little or no trust in the generated code 8. |
| Marketing | Creative variants for ads | Medium. A single study cited by the AI Index: 50% more output per worker 4. | Producing more pieces is not selling more; almost nobody connects output to conversion. |
| Operations and supply chain | Demand forecasting | Medium-low. Consulting figures from 2021 on predictive AI, with no published methodology 9. | Dirty master data; the model forecasts well and purchasing keeps doing exactly what it did before. |
| Finance | Document extraction | Medium-low. No public study, but it fits the pattern of structured, verifiable work 4. | The exception rate eats the savings when the review flow is not redesigned. |
| Sales | Lead scoring | Low. No public field study on this case; McKinsey places marketing and sales (the whole function) among the four with the highest theoretical potential value 3. | The CRM does not hold the data the model needs. |
| HR and legal | Clause extraction and internal policies | Low. No comparable public evidence, and the highest regulatory risk on the list. | Decisions about people or contracts with no traceability and no mandatory human review. |
Note
Where to start without burning the budget or the sponsorship
The data point that should structure your planning is not a technology one. Out of 25 organizational attributes evaluated, McKinsey identified workflow redesign as the one with the biggest effect on the ability to see EBIT impact from generative AI; even so, only 21% of the organizations using it say they have fundamentally redesigned at least some workflows 2. Picking the right case is one part of the work; the bigger part is changing how the task gets done around it.
- 1
1. Choose by pain, not by catalog
Take the three processes where the most hours go, or where the customer gets stuck the most. If none of them falls in the functions where potential value concentrates 3, that is fine: pain beats fashion as long as the result is verifiable.
- 2
2. Apply the verifiability filter
Can someone look at an individual output and say whether it was right or wrong? If the answer is *it depends*, the case is not ready. This is the rule that explains why the support copilot works and *AI generated strategy* does not.
- 3
3. Define the number before you build
Issues resolved per hour, minutes per document, escape rate. One number only, with its baseline taken before the system exists. Without a baseline there is no honest way to declare success afterwards.
- 4
4. Prototype against real data in weeks
One case working with your data and real pilot users. If it does not convince there, it stops: that is a valid result, and much cheaper than finding out after the integration.
- 5
5. Redesign the workflow, not just the tool
Who reviews what, which steps disappear, which exception escalates to a human and under what SLA. It is the part almost nobody does and the one most associated with seeing impact in results 2.
One last calibration. In BCG's global study of more than 1,250 companies, only 5% qualify as *future-built* (getting value from AI at scale), 35% are scaling with partial returns and 60% get no material value despite investing 5. That gap did not close by buying licenses. If your plan for the year is one well chosen use case, measured, with the process redesigned around it, you are playing the right game, even if it sounds less ambitious than a full transformation.
References
Every figure quoted in this article comes from the sources listed below. Each one links to the original document so you can check it yourself.
- 01
McKinsey & Company. The state of AI in 2025: Agents, innovation, and transformation (global survey, 1,993 respondents across 105 countries), 2025.
www.mckinsey.com/~/media/mckinsey/business%20functions/quantumblack/our%20insights/the%20state%20of%20ai/november%202025/the-state-of-ai-2025-agents-innovation_cmyk-v1.pdf - 02
McKinsey & Company. The state of AI: How organizations are rewiring to capture value (March 2025), 2025.
www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value - 03
McKinsey & Company. The economic potential of generative AI: The next productivity frontier, 2023.
www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier - 04
Stanford HAI. The 2026 AI Index Report, Economy chapter, 2026.
hai.stanford.edu/assets/files/ai_index_report_2026_chapter_4_economy.pdf - 05
Boston Consulting Group. The Widening AI Value Gap (Build for the Future 2025, n = 1,250 companies), 2025.
media-publications.bcg.com/The-Widening-AI-Value-Gap-Sept-2025.pdf - 06
Brynjolfsson, Li and Raymond (The Quarterly Journal of Economics). Generative AI at Work (NBER Working Paper 31161; QJE 140(2), 2025), 2025.
academic.oup.com/qje/article/140/2/889/7990658 - 07
Microsoft Research / GitHub (Peng, Kalliamvakou, Cihon, Demirer). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (preprint), 2023.
arxiv.org/abs/2302.06590 - 08
Google Cloud / DORA. State of AI-assisted Software Development 2025 (nearly 5,000 professionals globally), 2025.
dora.dev/dora-report-2025/ - 09
McKinsey & Company. Succeeding in the AI supply-chain revolution (April 2021), 2021.
www.mckinsey.com/industries/metals-and-mining/our-insights/succeeding-in-the-ai-supply-chain-revolution - 10
INEGI. Censos Económicos 2024. Resultados definitivos (Comunicado 79/25), 2025.
www.inegi.org.mx/contenidos/saladeprensa/boletines/2025/ce/CE2024_def.pdf