AI Business Strategy: How to Pick Use Cases That Reach the P&L
Most companies use AI but few see profit impact. How to score use cases, decide build vs buy, and measure savings that are real cash, with a worked example.

AI adoption is close to universal. Financial results are not. In McKinsey's State of AI survey published in November 2025, 88% of the 1,993 organizations surveyed used AI regularly in at least one business function. Only 39% could point to any effect on EBIT, and most of those put it below 5%. About one respondent in eighteen said AI accounted for more than 5% of their EBIT.
MIT's GenAI Divide report from mid-2025 reached a harsher headline, that about 95% of enterprise generative AI pilots showed no measurable profit-and-loss impact. That study has fair critics: its sample was small and its definition of success was narrow and short term. But the two sources agree on the pattern. Using AI is easy. Getting it to change the numbers a CFO reports is not.
This guide is about closing that gap: choosing where AI goes in the business, deciding what to build and what to buy, and measuring results in a way finance will accept.
What the companies seeing results do differently
McKinsey's data points to one practice above the others: the organizations with meaningful EBIT impact were much more likely to have redesigned workflows around AI instead of adding a tool to the existing process.
The difference is easy to see in customer support. Giving agents an AI assistant that drafts replies makes each agent somewhat faster. Redesigning the workflow means the assistant resolves routine tickets end to end, routes the rest with a summary already written, and agents spend their time on the hard cases. The second version changes staffing plans. The first mostly changes how the work feels.
Other patterns from the same data: high performers used AI across more functions, had senior leaders who personally owned AI outcomes, and tracked defined KPIs for their AI work. None of that is technical.
Choosing where AI goes
Most companies start with whatever a vendor demoed or an enthusiastic team proposed. A better starting point is a list of every candidate use case, scored on the same criteria.
A scoring method
Score each candidate from 1 to 5 on four questions:
- Value. If this worked, how much would it move revenue, cost, or risk? Estimate in money, even roughly.
- Data readiness. Does the data exist, is it accessible, and is it good enough? Many projects stall here.
- Feasibility. Can current tools do this reliably? Drafting and summarizing are mature. Autonomous multi-step decisions are less so.
- Risk. What happens when it is wrong? Score inversely: 5 means an error is cheap and easy to catch, 1 means an error harms customers or breaks regulations.
Here is an illustration for a mid-sized B2B company:
| Use case | Value | Data | Feasibility | Risk (5 = low) | Total |
|---|---|---|---|---|---|
| Support ticket drafting and routing | 4 | 5 | 5 | 4 | 18 |
| Sales call summaries into CRM | 3 | 4 | 5 | 5 | 17 |
| Demand forecasting for inventory | 5 | 3 | 4 | 3 | 15 |
| Contract review for procurement | 3 | 3 | 4 | 2 | 12 |
| Fully automated credit decisions | 5 | 3 | 3 | 1 | 12 |
The top two are unglamorous and likely to work. The forecasting project is worth more but needs data work first, which is a reason to start that work now rather than skip it. The last one scores high on value but combines weak data with serious regulatory risk; in the EU, AI credit scoring of individuals is a high-risk use under the AI Act from December 2027.
Pick two or three projects from the top of the list. Spreading a small team across ten pilots is the most reliable way to join the majority of pilots that never scale.

Build, buy, or assemble
The choice is rarely "build a model or buy a product." Most real projects assemble pieces.
| Approach | When it fits | Watch out for |
|---|---|---|
| Buy a finished tool | Common needs: meeting notes, writing help, support drafting | Per-seat costs that grow with headcount; data terms |
| Use a model API with your own prompts and retrieval | Your documents or data make the answer useful | Usage costs that scale with volume; quality testing |
| Fine-tune an existing model | A narrow task with lots of examples and a consistent format | Keeping it updated as base models improve |
| Train from scratch | Almost never outside AI companies | Cost, talent, and obsolescence |
The question to ask is where your advantage comes from. If it comes from your proprietary data, such as years of claims history, pricing data, or engineering documents, build the part that uses that data and buy everything else. If a competitor could get the same result by buying the same tool, the tool is not a strategy. It may still be worth buying for the efficiency.
Two costs catch companies out. Usage-based pricing means a successful feature gets more expensive as it spreads, so model cost per transaction belongs in the business case. And switching model providers is easier if prompts, evaluation tests, and retrieval sit in your own code rather than inside a vendor's product.
A worked ROI example
A support team handles 40,000 tickets a month. Average handling time is 12 minutes and the loaded cost of an agent is $45 an hour. A pilot shows that an AI assistant can cut handling time by 25% on the 60% of tickets that are routine. The numbers are illustrative.
- Tickets affected: 40,000 × 60% = 24,000 a month
- Time saved: 24,000 × 12 minutes × 25% = 72,000 minutes, or 1,200 hours a month
- Value of that time: 1,200 × $45 = $54,000 a month
- Assumed running cost (licenses, model usage, and part of a QA reviewer's time): $15,000 a month
- Net capacity value: about $39,000 a month
Here is the part most business cases skip. The 1,200 hours are roughly 7.5 full-time agents' worth of time. That becomes cash only if something changes: the team does not backfill leavers, it absorbs volume growth without new hires, contractor or overtime spend falls, or agents move to work that brings in revenue, such as renewals or upsells. If none of those happen, the company has paid $15,000 a month to make agents less busy.
Agree with finance beforehand which of those outcomes counts as the return, and track it.
Measure before you start
The most common reason AI pilots "fail" is that nobody recorded what things looked like before. Critics of the MIT report made this point: many pilots had no baseline, so no improvement could be shown even when one happened.
Before any pilot, write down:
- Volume and cycle time for the process today
- Error or rework rates
- Cost per unit (per ticket, per invoice, per contract reviewed)
- The target, and what result would mean stopping the project
Then run the pilot long enough to see a stable result, typically six to twelve weeks for operational processes.
Organization and ownership
A named business owner for each use case. Not the AI team. The person who owns the P&L line the use case is supposed to move.
A small central team. It sets standards for model selection, security, evaluation, and vendor contracts, and helps business teams build. It should not be the only place AI gets built.
Governance proportional to risk. An internal meeting summarizer and an automated loan decision need very different controls. The AI governance frameworks guide covers tiered approaches, and generative AI data governance covers what data can go where.
Regulatory dates on the roadmap. EU AI Act transparency duties have applied since 2 August 2026, with high-risk obligations from 2 December 2027. California's automated decision-making rules and Colorado's revised AI law take effect on 1 January 2027. Projects touching hiring, lending, insurance, or housing should plan for them now. Our AI and SaaS data privacy guide has the details.

A first year
- Month 1. List candidate use cases across functions and score them. Pick two or three.
- Months 2 to 4. Record baselines, run pilots, and agree with finance how results will be counted.
- Months 4 to 6. Redesign the workflow around the use cases that work. Stop the ones that do not.
- Months 6 to 12. Scale the winners, start the data work for high-value projects that scored low on readiness, and add the next batch from the list.
For AI in specific areas, see AI for strategic decisions, AI financial forecasting, AI customer experience, and AI in sales and RevOps. For running models reliably once they work, see MLOps best practices.
This article is for informational purposes only and is not professional business, legal, or financial advice. Evaluate AI investments with advisors who know your organization.



