Most pages that promise AI strategy examples describe capabilities. This one names five companies, gives the number each of them actually published, and says what happened afterwards, including the two cases where the company reversed the decision inside a year.
That last part is the whole point. Almost every AI example in circulation is a launch announcement. The launch number is the easiest number to produce and the least useful one to copy, because it is measured at the moment of maximum optimism by the party with the most to gain from it. The examples below are chosen because there is a second data point: a follow-up, a reversal, an apology, or a survey that checked whether the effect held.
Every example that survived contact with a second year kept people in the loop. Every example that had to be reversed had removed them first.
That is not an argument against automation. Three of the five companies below are still automating hard. It is an argument about sequencing: the ones who measured before cutting are still running their systems, and the ones who cut on a projection had to walk it back in public.
The five examples, and what happened next
| Company | The decision | The number they published | What happened next |
|---|---|---|---|
| Klarna | AI assistant took over front-line support | Two-thirds of chats, 2.3 million in month one (Feb 2024) | Rehired human agents in 2025 after the CEO said quality had dropped too far; now runs a hybrid model |
| Commonwealth Bank | Cut 45 contact centre roles after deploying a voice bot | Bot claimed to cut call volumes by 2,000 a week | Volumes actually rose. The bank called the redundancies an error, apologised, reversed them (Aug 2025) |
| Salesforce | Agentforce on its own help service | Support headcount 9,000 to 5,000; 380,000 conversations at 84% resolution over 90 days; support cost down 17% | Still running and still expanding. Every figure is the vendor's own, about its own product |
| Lumen Technologies | Microsoft Copilot for the sales team | Four hours a week back per seller, valued at $50 million a year | Still cited. The dollar figure is an extrapolation from hours saved, not measured revenue |
| JPMorgan Chase | One internal LLM assistant, deployed firm-wide | Around 140,000 employees at announcement, reported near 200,000 since | Still expanding, with a roadmap of production use cases. No firm-wide ROI figure has been published |
Read the fourth column before the third. The third column is what gets quoted in decks. The fourth is what tells you whether the example is one to copy.
What makes an AI strategy example worth copying
An example is worth copying when three things are true: a named company made a specific decision, published a number attached to it, and enough time has passed that you can check whether the number held.
Most of what circulates fails the third test, and a fair amount fails the second. "Company X is using AI for customer service" is a capability, not a strategy. The strategy is the choice underneath it: what was measured first, what was cut, what was kept, and what the organisation did when the first number turned out to be wrong.
Apply that filter and the pool shrinks dramatically, which is the honest reason this page has five examples and not thirty.
Klarna: two-thirds of support automated, then people hired back
Klarna announced in February 2024 that its OpenAI-powered assistant was handling two-thirds of customer service chats, 2.3 million conversations in its first month, and doing the work of roughly 700 full-time agents. It became the single most-cited AI deployment in B2B software, and it was cited almost exclusively in that first form.
In May 2025 the company started recruiting human agents again. CEO Sebastian Siemiatkowski told Bloomberg the cost-focused push had gone too far and had produced lower quality service, and Klarna committed to guaranteeing customers the option of reaching a person. By late 2025 the assistant was handling more volume than ever, alongside the rehired humans, in an explicitly hybrid model.
The useful reading is not that the AI failed. The volume numbers held. What did not hold was the assumption that volume handled is the same as service delivered, and Klarna is unusual mainly in having said so out loud.
What it implies for a revenue team. Deflection rate is not a quality metric. If you are automating any customer-facing step, decide in advance what the escape hatch is and what threshold triggers it, because you will need it and you would rather design it than retrofit it.
Commonwealth Bank: 45 roles cut on a number that was wrong
Australia's largest bank announced in July 2025 that it was cutting 45 contact centre roles after deploying an AI voice bot, citing a reduction of about 2,000 calls a week. Staff and the Finance Sector Union disputed it: call volumes were rising, not falling, the bank was offering overtime, and team leaders were being pulled in to answer phones.
In August 2025 the bank reversed the redundancies, described the decision as an error, apologised to the affected employees, and said it had not adequately considered all the relevant business factors. The 45 people were offered their roles back, redeployment, or exit.
This is the cleanest available example of the failure mode that matters most, and it has nothing to do with model quality. The bot may well have worked as specified. The measurement around it did not, and the headcount decision was made on the measurement rather than on the outcome.
What it implies for a revenue team. Whoever owns the AI deployment should not own the number that grades it. If your routing automation reports its own success rate, you have a reporting problem wearing an AI costume. This is ordinary revenue operations hygiene, and it is the thing AI projects skip most reliably.
Salesforce: 9,000 support staff to 5,000, reported by the company selling the agents
Marc Benioff has said Salesforce cut its customer support organisation from roughly 9,000 people to 5,000 as Agentforce took over more of the load, that Agentforce handled 380,000 conversations at an 84% resolution rate over a 90-day window, and that support costs are down about 17%.
Those numbers are worth taking seriously and worth labelling precisely. They are self-reported, by the CEO, about his own company's use of the product his company sells, in podcast interviews and social posts rather than in audited disclosure. That does not make them wrong. It does mean they carry the weakest evidentiary status of anything on this page, and they should not be used as a benchmark for what the same product would do in a company that is not Salesforce.
What it implies for a revenue team. The strongest version of this example is not the headcount number, it is the shape: a company with unusually clean data about its own product, running agents against its own documentation, in a domain where the correct answer already exists in writing. Support deflection is the easiest agentic use case there is. Nothing about an 84% resolution rate there predicts anything about AI agents in a sales motion, where there is no documented correct answer to retrieve.
Lumen: four hours a week per seller, priced at $50 million
Lumen Technologies reported that Microsoft Copilot gave its sellers back around four hours a week each, mostly by collapsing pre-outreach research from roughly four hours to about fifteen minutes, and that the company valued this at $50 million a year.
The $50 million is the most quoted sales-AI figure in circulation and it deserves a caveat that almost never travels with it. It is an extrapolation: hours saved, multiplied by a value assigned to selling time, published through Microsoft's own customer story programme. It is not measured incremental revenue, and Lumen did not claim it was. Treat it as a credible statement about time and an assumption about what that time is worth.
What it implies for a revenue team. Time returned is real and it is the most reproducible AI benefit available to a sales org today. Whether it converts into pipeline depends entirely on whether the freed hours go into customer conversations or get absorbed. That conversion is a management question, not an AI question, and it is where most of the value in this example is won or lost.
JPMorgan Chase: an assistant for everyone, with no ROI number attached
JPMorgan Chase said in 2024 it was rolling its internal LLM Suite out to around 140,000 employees, and reporting since has put the figure closer to 200,000, with roughly half of eligible employees using it daily and a roadmap toward a large number of production use cases.
What makes this example instructive is the absence. A bank that discloses granular numbers about almost everything has published adoption figures and not financial ones. The most defensible interpretation is the one the bank itself implies: firm-wide assistant deployment is infrastructure, its returns are diffuse and hard to attribute, and the honest way to report it is adoption plus use-case count rather than a manufactured ROI figure.
What it implies for a revenue team. If you cannot attribute the return, say so and measure adoption instead. A made-up attribution number is worse than an honest adoption number, because someone will eventually make a headcount decision on it. See the Commonwealth Bank example.
What the 2026 surveys say these examples have in common
Three pieces of research are doing most of the work in the current debate, and all three were re-checked for this refresh rather than restated from the earlier version of this page.
McKinsey's State of AI survey, fielded in May and June 2026 across 1,719 respondents and published in August 2026, reports that the share of organisations qualifying as AI high performers, meaning they attribute at least 5% of EBIT to AI, has stayed flat at about 6%. Thirty-seven percent attribute at least some EBIT impact, about the same as the year before, while 80% report improved individual productivity. The gap between those two figures is the finding: personal productivity gains are now widespread and enterprise financial impact is not. On agents, 23% report scaling an agentic system in at least one function and another 39% are experimenting, but in any given business function no more than about 10% are scaling.
The distinguishing behaviour is the same one the 2025 edition found. Nearly three-quarters of high performers say they fundamentally redesigned workflows because of AI, against about a quarter of everyone else.
BCG's January 2026 work reaches a similar share by a different route: about 6% of companies qualify as AI leaders, and those leaders outperform peers by around 9 percentage points in industry-adjusted shareholder returns. BCG's older and still widely cited finding, that 74% of companies struggle to show tangible value, sits alongside its 10-20-70 rule: 10% of effort on algorithms, 20% on technology and data, 70% on people and process.
The MIT NANDA report that produced the famous "95% of GenAI pilots fail" line deserves its caveat. The figure comes from The GenAI Divide: State of AI in Business 2025, built on roughly 150 interviews, a 350-person employee survey and an analysis of about 300 public deployments, and the headline has been contested on methodology since publication. The direction it points is consistent with McKinsey's and BCG's numbers. The precision implied by "95%" is not.
How we run AI on our own marketing
The examples above are all large companies with budgets that make the sequencing question easy to get wrong. Here is the small version, which is ours and which we can describe in full because it is our own operation rather than a client's.
The organic marketing motion behind this site runs as a scheduled AI routine against this website's own repository. It reads Search Console data, works out which page is losing ground and why, researches and drafts one piece of work, and stops. The design constraints are the interesting part, because every one of them exists because of a failure mode in the examples above.
It does one unit of work per run, so there is a reviewable diff rather than a batch of output nobody reads. Nothing it writes goes live without a human approving it, and new pages ship set to noindex until that happens. It cannot send anything, post anything, or sign up for anything. Corrections given at the approval step become written rules that every later draft is checked against, so the same mistake cannot be made twice quietly. And the numbers it reports on its own work are pulled from Search Console rather than from its own claims, which is the Commonwealth Bank lesson applied to ourselves.
The operator time this takes is roughly 45 minutes a week, spent almost entirely on approving or correcting, not producing. That is the honest figure for what the system costs to run. We are not going to attach a revenue number to it, for the same reason JPMorgan has not attached one to LLM Suite.
What to copy, in order
Abstract strategy is useless, so here is the sequence the five examples actually support.
The step most teams want to skip is the second one, because it is the only one with no visible output. It is also the one that decides whether the project survives its first executive review.
Frequently asked questions
What is the best example of a company using AI successfully?
Judged on evidence rather than on the size of the claim, Lumen Technologies is the cleanest: a specific, repeatable time saving on a well-defined task, with the company clear about what was measured. Salesforce's numbers are larger but self-reported about its own product, and Klarna's are the most famous but had to be revised. No public example yet pairs a large AI-attributed revenue figure with independent verification.
Why do most AI implementations fail?
The research points at sequencing rather than technology. BCG's 10-20-70 rule puts 70% of the work on people and process and only 10% on algorithms, and McKinsey's 2026 survey found that fundamentally redesigning workflows is the behaviour that separates the roughly 6% of high performers from everyone else. The common failure is deploying a model on top of an undocumented process and then measuring the model instead of the outcome.
Is the MIT statistic that 95% of AI pilots fail accurate?
It comes from MIT NANDA's The GenAI Divide: State of AI in Business 2025, based on around 150 interviews, a 350-person survey and about 300 public deployments, and its methodology has been publicly contested. Other 2026 research points the same direction: McKinsey finds 37% of organisations reporting any EBIT impact and about 6% qualifying as high performers. Treat 95% as directionally consistent rather than precise.
What should a small revenue team automate first?
The task where the correct answer already exists in a document you control: lead routing rules, CRM data hygiene, pre-call research, meeting summarisation into the CRM. These are the small-team version of the support-deflection pattern, and they are the same starting point as what actually works with AI SDRs. The prerequisite is the same in both cases, which is a CRM whose data you trust enough to act on.
Want a second opinion on an AI project before the headcount decision?
Book a free 30-minute call. We will look at what you are automating, what the baseline measurement is, and whether the number you are planning to act on is coming from a system with an incentive to look good. No pitch deck.
Book a call →If the next step is building rather than deciding, our n8n and AI automation work covers the orchestration layer, and Clay for RevOps covers the enrichment and scoring layer most of these workflows need underneath them.
Sources
None of the primary sources below was directly reachable from the environment this page was written in. Figures are therefore credited to the publishers that reported them, and the linked pages are where to verify the current numbers.
- MIT NANDA, The GenAI Divide: State of AI in Business 2025, as reported by Fortune
- McKinsey, The State of AI: Global Survey, 2026 edition, fielded May to June 2026, n=1,719, published August 2026
- BCG, AI Talk Is Cheap. Value Creation Is Rare., January 2026, and BCG's 2024 adoption study for the 74% figure and the 10-20-70 rule
- Klarna's February 2024 announcement of its AI assistant, and the 2025 reversal as reported by CX Dive and Forbes
- Commonwealth Bank's reversal as reported by ABC News and The Register
- Salesforce support headcount and Agentforce figures, from Marc Benioff's public statements as reported by The Register
- Lumen Technologies and Microsoft Copilot, from Microsoft's own customer story
- JPMorgan Chase LLM Suite rollout, as reported by CIO Dive