Neuro Ares AI
Use cases14 min read

AI for business: 15 use cases with real impact on the numbers

The list of things AI *could* do in your company is infinite. The list of things whose effect someone has already measured in the field is short and fairly boring. Here are 15 use cases by function, separating the ones with studies behind them from the ones that are still a reasonable promise.

CE
Carlos EspejelSocio · Escalamiento de negocios con IA
Updated August 18, 2026Published August 18, 2026

Almost every conversation about AI for business starts backwards: someone brings in a tool and then goes looking for a place to put it. The order that works is the reverse: which process hurts, which number has to move, and only then which technology applies. That has an uncomfortable consequence: half the fashionable use cases do not survive that question.

88%

of respondents say their organization uses AI regularly in at least one business function 1

39%

of respondents identify some EBIT impact at the enterprise level, and most put it below 5% of EBIT 1

75%

of the potential value estimated by McKinsey would be concentrated in just four functions 3

2.1%

of the Mexican economic units that use digital tools report using AI systems 10

Read them together, with one warning about method: they are not comparable bases. The first three figures are self-reported by executives in global surveys, or estimates of theoretical potential; the Mexican one is a census of economic units, most of them micro businesses, and it cannot be subtracted from the others. Even so, the picture reads clearly: reported usage is everywhere, impact that anyone declares measurable is in very few places, potential value concentrates in a handful of functions, and AI adoption in Mexico is still marginal. None of that means AI does not work: it means the difference between the companies that capture value and the ones that do not is rarely in the use case they picked, but in whether anyone redesigned the process around it.

01

Where does the value of AI in a company actually concentrate?

In 2023, McKinsey analyzed 63 generative AI use cases across 16 business functions and estimated that, applied broadly, they could add the equivalent of $2.6 trillion to $4.4 trillion a year to the global economy. The useful part is not the amount (it is theoretical potential, not captured value) but the concentration: around 75% of that value would land in four areas, customer operations, marketing and sales, software engineering and R&D 3.

That concentration is your first filter. The second one comes from Stanford HAI's 2026 AI Index: the productivity gains reported by the studies are larger in structured, measurable work, where the output is easy to monitor 4. Translated: the cases that work are the ones where someone can say, without argument, whether an individual output was right or wrong.

Where the work is structured and the result easy to verify, AI delivers. Where judgment is fuzzy, it delivers in the demo and not in the operation.
A reading of the Economy chapter of the 2026 AI Index, Stanford HAI 4
02

Sales and marketing: 4 use cases

Marketing and sales is one of the four functions where potential value concentrates 3 and also one of the worst measured: it is easy to prove that more content got produced and hard to prove that more got sold.

01

1. Creative variants for ads · Measured

The best supported marketing case: a study cited by the AI Index (Ju and Aral, 2025) reports 50% more output per worker in teams that used multimodal AI for ad creation 4. It is a narrow context, not a market average, and it measures output, not conversion. It works where you already know which message lands.

02

2. Audience research on your own data · Emerging

Processing tickets, sales call transcripts and reviews to pull out the customer's real objections and vocabulary. The value is not hours saved: it is that you stop writing what marketing *thinks* the customer says. The output is verifiable against the source, which is exactly where it pays off most 4.

03

3. Lead scoring and enrichment · Promise

Classifying inbound, filling in account data and prioritizing the rep's daily list; measured in effective contact rate. We know of no public field study that quantifies it. It almost always fails for the same reason: the CRM does not hold the data the model would need.

04

4. Proposal and quote assembly · Promise

Assembling the proposal from the catalog, pricing rules, account history and approved templates. Measured in cycle time between request and delivery, which in midsize companies is counted in days. If the pricing rules live in the sales director's head, this is a documentation project.

03

Customer service: the 2 cases with the strongest evidence

If I had to pick one function to start with in a midsize company, it would be this one, and not because it is fashionable: it is where the causal data is best. Brynjolfsson, Li and Raymond measured the staggered rollout of a conversational assistant across 5,172 agents at a Fortune 500 business process software company: 15% more issues resolved per hour on average 6. That is not executive self-report. Before extrapolating it, it helps to know the context: a single company, close to 80% of the agents based in the Philippines, and a tool built on GPT-3 6.

01

5. Copilot for live agents · Measured

The system suggests the reply; the agent decides. More important than the +15% average is the distribution: the gains concentrated in novice agents (from 30% to 36% in the lowest skill quintile), while the most experienced ones saw minimal speed gains and even small declines in quality 6. The AI Index picks up this same study in a range of 14% to 15% 4; it is not a second independent measurement.

02

6. Self-service that answers by citing the source · Emerging

Answers the customer only from your own document base, with a link to the policy it came from. The number to move is deflection: queries resolved without an agent. What decides the outcome is not the model but the base: if your policies are out of date, the system will repeat them faster and to more people.

They never make it into the slide decks because they are not glamorous, and they produce the savings that are easiest to defend in front of a CFO: the process is already instrumented, there is volume, and there is a cost per transaction that someone already calculates.

Finance

01

7. Data extraction from invoices and documents · Emerging

Read the document, extract the fields and leave them recorded in the ERP. The archetype of the structured, verifiable work where the AI Index places the biggest gains 4. Measure minutes per document and exception rate: the savings evaporate if one in three ends up in manual review.

02

8. Anomalies in expenses and reconciliations · Emerging

Flag the line that does not add up, the duplicate expense, the supplier with an odd pattern. Classic predictive AI, mature since well before the generative wave, and it audits itself: every alert gets reviewed. The risk is false positives, which make finance stop opening the alerts within three weeks.

Operations and supply chain

01

9. Demand and inventory forecasting · Emerging

The most quoted numerical promise on this list and the most fragile evidence on it. McKinsey reported in 2021 that implementing an AI-enabled supply chain successfully allowed early adopters to improve logistics costs by 15%, inventories by 35% and service levels by 65% against slower competitors 9. It is from 2021, it is about predictive AI, it compares across companies and it publishes no methodology.

02

10. Visual quality inspection · Emerging

Computer vision on the line to catch the defects the human eye lets through at the end of a shift. An indisputable metric (escape rate) and a payback you can calculate per return avoided. It needs labeled images of your own defects: weeks of work before the first line of code.

01

11. Structuring and shortlisting candidates · Promise

Turning heterogeneous CVs into comparable records and ranking them by explicit criteria. It improves hiring cycle time, not hiring quality. It is the case with the most legal risk on this list: the system ranks and explains, a person signs off on the rejection.

02

12. Internal policy and onboarding assistant · Promise

Vacation, benefits, expenses and procedures answered from the official documents, with the citation to the rulebook. Usually the best first internal deployment: low risk, a bounded corpus and everyone notices whether it works. Measured in the repeat tickets HR handles today.

03

13. Clause review and extraction · Promise

Terms, penalties, automatic renewals and deviations from your template, across an archive nobody has read end to end. The value is not replacing the lawyer: it is finally knowing what contract 400 says. Human review is not optional and traceability is part of the deliverable.

05

Software engineering: 2 use cases

This is the function with the most published data and the easiest one to overstate. The most cited randomized experiment, with 95 developers recruited on Upwork, found that those given access to GitHub Copilot completed the task 55.8% faster than the control group 7. Before you take that number to a committee: it was a narrow, synthetic task (an HTTP server in JavaScript), the result is conditional on having completed it, the 95% confidence interval runs from 21% to 89% and it is a preprint by authors from GitHub and Microsoft.

Measurements on real work give more modest figures. The AI Index picks up a study (Cui et al., 2025) with 26% more pull requests completed by developers using Copilot, and the counterpoint: METR found experienced developers were 19% slower with AI assistance on deep reasoning tasks, though it later failed to replicate the result 4. Google's DORA 2025 survey, with nearly 5,000 professionals globally, closes the picture: 90% report using AI in their work and more than 80% believe they gained productivity, but 30% report little or no trust in the generated code 8.

01

14. Code assistant in the developer's flow · Measured

Autocomplete, tests, explaining inherited code. The best studied case, with a nuance DORA documents: AI adoption improves delivery throughput and at the same time increases delivery instability 8. Measure change failure rate and recovery time alongside speed, or you will celebrate a result that operations is paying for.

02

15. Agents for bounded maintenance tasks · Emerging

Repetitive migrations, dependencies, test coverage: tedious work, verifiable by the test suite and low risk if a human reviews the pull request. One of the few places with real usage at scale: the AI Index reports that agent use at scale is still in single digits in almost every function, except in the tech sector, where it reaches 24% in software engineering 4.

06

Which functions have measured evidence and which are still a promise

This is the table I use to structure a prioritization conversation. It does not say which case is best for you: it says how much confidence you can put behind the number you are about to promise.

Maturity of the public evidence by function. The absence of a study does not mean the case does not work: it means you will have to measure it yourself.
FunctionBest supported caseEvidenceWhat usually goes wrong
Customer serviceCopilot for agentsHigh. Staggered field study with 5,172 real agents: +15% in issues resolved per hour 6.Speed gets measured and quality gets ignored; the most skilled agents showed small but statistically significant declines in resolution and satisfaction 6.
Software engineeringCode assistantMedium-high. Randomized experiment on a narrow task, in a preprint by authors from GitHub and Microsoft 7, plus a self-reported global survey of nearly 5,000 professionals 8.It raises throughput and delivery instability along with it 8; 30% report little or no trust in the generated code 8.
MarketingCreative variants for adsMedium. A single study cited by the AI Index: 50% more output per worker 4.Producing more pieces is not selling more; almost nobody connects output to conversion.
Operations and supply chainDemand forecastingMedium-low. Consulting figures from 2021 on predictive AI, with no published methodology 9.Dirty master data; the model forecasts well and purchasing keeps doing exactly what it did before.
FinanceDocument extractionMedium-low. No public study, but it fits the pattern of structured, verifiable work 4.The exception rate eats the savings when the review flow is not redesigned.
SalesLead scoringLow. No public field study on this case; McKinsey places marketing and sales (the whole function) among the four with the highest theoretical potential value 3.The CRM does not hold the data the model needs.
HR and legalClause extraction and internal policiesLow. No comparable public evidence, and the highest regulatory risk on the list.Decisions about people or contracts with no traceability and no mandatory human review.
07

Where to start without burning the budget or the sponsorship

The data point that should structure your planning is not a technology one. Out of 25 organizational attributes evaluated, McKinsey identified workflow redesign as the one with the biggest effect on the ability to see EBIT impact from generative AI; even so, only 21% of the organizations using it say they have fundamentally redesigned at least some workflows 2. Picking the right case is one part of the work; the bigger part is changing how the task gets done around it.

  1. 1

    1. Choose by pain, not by catalog

    Take the three processes where the most hours go, or where the customer gets stuck the most. If none of them falls in the functions where potential value concentrates 3, that is fine: pain beats fashion as long as the result is verifiable.

  2. 2

    2. Apply the verifiability filter

    Can someone look at an individual output and say whether it was right or wrong? If the answer is *it depends*, the case is not ready. This is the rule that explains why the support copilot works and *AI generated strategy* does not.

  3. 3

    3. Define the number before you build

    Issues resolved per hour, minutes per document, escape rate. One number only, with its baseline taken before the system exists. Without a baseline there is no honest way to declare success afterwards.

  4. 4

    4. Prototype against real data in weeks

    One case working with your data and real pilot users. If it does not convince there, it stops: that is a valid result, and much cheaper than finding out after the integration.

  5. 5

    5. Redesign the workflow, not just the tool

    Who reviews what, which steps disappear, which exception escalates to a human and under what SLA. It is the part almost nobody does and the one most associated with seeing impact in results 2.

One last calibration. In BCG's global study of more than 1,250 companies, only 5% qualify as *future-built* (getting value from AI at scale), 35% are scaling with partial returns and 60% get no material value despite investing 5. That gap did not close by buying licenses. If your plan for the year is one well chosen use case, measured, with the process redesigned around it, you are playing the right game, even if it sounds less ambitious than a full transformation.

References

Every figure quoted in this article comes from the sources listed below. Each one links to the original document so you can check it yourself.

  1. 01

    McKinsey & Company. The state of AI in 2025: Agents, innovation, and transformation (global survey, 1,993 respondents across 105 countries), 2025.

    www.mckinsey.com/~/media/mckinsey/business%20functions/quantumblack/our%20insights/the%20state%20of%20ai/november%202025/the-state-of-ai-2025-agents-innovation_cmyk-v1.pdf
  2. 02

    McKinsey & Company. The state of AI: How organizations are rewiring to capture value (March 2025), 2025.

    www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value
  3. 03

    McKinsey & Company. The economic potential of generative AI: The next productivity frontier, 2023.

    www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier
  4. 04

    Stanford HAI. The 2026 AI Index Report, Economy chapter, 2026.

    hai.stanford.edu/assets/files/ai_index_report_2026_chapter_4_economy.pdf
  5. 05

    Boston Consulting Group. The Widening AI Value Gap (Build for the Future 2025, n = 1,250 companies), 2025.

    media-publications.bcg.com/The-Widening-AI-Value-Gap-Sept-2025.pdf
  6. 06

    Brynjolfsson, Li and Raymond (The Quarterly Journal of Economics). Generative AI at Work (NBER Working Paper 31161; QJE 140(2), 2025), 2025.

    academic.oup.com/qje/article/140/2/889/7990658
  7. 07

    Microsoft Research / GitHub (Peng, Kalliamvakou, Cihon, Demirer). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (preprint), 2023.

    arxiv.org/abs/2302.06590
  8. 08

    Google Cloud / DORA. State of AI-assisted Software Development 2025 (nearly 5,000 professionals globally), 2025.

    dora.dev/dora-report-2025/
  9. 09

    McKinsey & Company. Succeeding in the AI supply-chain revolution (April 2021), 2021.

    www.mckinsey.com/industries/metals-and-mining/our-insights/succeeding-in-the-ai-supply-chain-revolution
  10. 10

    INEGI. Censos Económicos 2024. Resultados definitivos (Comunicado 79/25), 2025.

    www.inegi.org.mx/contenidos/saladeprensa/boletines/2025/ce/CE2024_def.pdf