Somewhere between a board meeting and an e-commerce roadmap, “we need AI” becomes a non-negotiable business requirement.
The pressure is real. In Dataiku’s 2026 survey of 900 enterprise CEOs, 80% said their role could be at risk if their company failed to show measurable AI gains by year-end. These circumstances can move an AI project onto the roadmap before anyone has fully answered a simple question: why this one?
That’s where AI implementation challenges tend to begin. A team can pick a promising use case, run a pilot, and get a good result. The harder part is working out what that result changed for the business and whether it is worth taking further.
So let’s look at what goes wrong in e-commerce, why working pilots stall, how to choose an idea worth funding, and what to measure when it is time to answer the obvious question: did it work?
Why AI adoption challenges often start with prioritization
Data quality, talent, integrations, and infrastructure are usually among the first to blame when an AI project struggles. In many cases, the troubles arise earlier, with a simpler question no answered. And that question is, “Which problem are we solving first?”
E-commerce teams rarely run out of ideas. Smarter search, a support copilot, automated product tagging, demand forecasting, personalized merchandising… They all make perfect sense. The harder call is deciding which one deserves the budget now.
Dataiku found that 65% of enterprise CEOs worry more about over-investing in the wrong AI vendors than under-investing. Vendor selection comes later, but the money is already on the line when a team decides which AI initiative gets funded.
Once that decision is made, the idea meets the store you already have.
Where AI implementation breaks in e-commerce operations
As expected, the usual suspects do show up eventually. In e-commerce, they look pretty ordinary: a pilot with no owner, success criteria that arrive too late, messy product data, or a new feature running into years of platform logic.
Those problems are easier to budget for when the idea has a business case behind it before the work begins. Let’s take a closer look at each.
#1 A pilot with no owner
A pilot can go pretty far before anyone notices this one.
Say the merchandising team tests AI product tagging. The results look promising, the team likes it, and everyone agrees it is worth continuing. Then the test ends and the tool needs a place in the actual workflow.
Someone has to keep an eye on the output and make sure feedback reaches the people who can act on it. A few months later, somebody still needs to care whether the team is using the tool at all.
When nobody owns that part, the pilot can technically succeed and still go nowhere. The company has paid to prove the idea works, then gets very little from it after the test.
#2 Success criteria set too late
A pilot can also look perfectly fine until someone asks what changed because of it.
That is when the hunt for a useful metric can begin. AI search was supposed to improve product discovery, but zero-result searches barely moved. Engagement did, though. Suddenly engagement starts looking very interesting.
You still have numbers. The problem is that the goal has moved after the result was already visible, which makes it hard to tell whether the pilot solved the problem it was funded to solve.
#3 Data that needs more work
E-commerce data usually carries a few years of decisions with it.
Product attributes may have changed as the catalog grew. Customer records can be split across systems that were never designed to agree on every field. Support teams may have changed how they tag tickets more than once.
Most of these inconsistencies can stay fairly quiet during everyday work. An AI feature has a habit of finding them.
Recommendations get worse in certain categories, or the team ends up checking far more output by hand than expected. Part of the project budget is now going into data work nobody had included in the original idea.
#4 Integrations that grow the scope
The store itself may have just as much baggage.
A Salesforce Commerce Cloud store may have custom pricing logic added years ago. BigCommerce might already be connected to ERP and PIM, with several other services sitting around them. An Adobe Commerce setup can have plenty of custom logic behind the storefront too.
Then the AI feature needs something from that setup. Live inventory might be enough for the pilot, while a wider rollout also needs customer data or existing order logic.
The scope grows from there. Work that looked like an AI feature now includes platform changes and considerably more testing than the pilot suggested.
Build the AI business case before implementation
An AI idea becomes easier to judge once you know what the problem is costing you today. A business case turns “this could be useful” into a decision you can actually put money behind.
Take customer service. You probably know how many tickets come in and roughly what they cost to resolve. In merchandising, you can count the hours spent on manual tagging or product enrichment. Checkout already gives you abandonment data.
That number gives the AI business case somewhere to begin. Then ask what would have to change for the investment to make sense and how quickly you could get enough evidence to judge it.
The business case also needs a reality check: can the company run the idea with the data and systems it already has? An AI readiness assessment helps answer that without turning the business case into a technical audit.
If the case survives those checks, the pilot has a reason to exist. A good result moves the conversation to production.
Why pilots stall before they reach production
A pilot gets temporary conditions. It may cover one category, run on limited funding, and receive more attention because everyone knows it is a test. Production has to keep working after the experiment ends. That gap is where some AI pilot projects end up in pilot purgatory. The idea proved enough to continue, while phase two was never treated as its own investment.
Product tagging may work well on one category, for example. Rolling it out across the catalog means paying for the work after the test and making it part of the merchandising operation.
Someone now has to approve a longer-term commitment. Projects have a much easier route forward when a sponsor is already expecting that decision and money for phase two has been discussed before the first budget runs out.
If the original target was met, the sponsor has evidence for the investment. From there, measurement has a different job: finding out what that improvement actually changed for the business.
How to measure AI project ROI
The first measurable gain from AI often appears inside the workflow. The business result usually sits one step beyond that metric.
A customer-care copilot, for example, may first show its impact in shorter handling time. McKinsey reports handling time reductions of 40–60% in some retail deployments, with one sportswear company cutting handling time by more than 40%. That sounds impressive. But on its own, it tells you very little about the business result.
Faster case handling may lower service costs. The same team may also be able to deal with more volume without adding people. Those outcomes put a different value on the same 40–60% figure. For merchandising, for example, saved time only tells part of the picture. Its value becomes clearer when catalog updates move faster or the team can handle more products with the same capacity.
That is the job of AI ROI measurement here: connect the first measurable improvement to the business result that paid for the project. Indirect gains count when you can show what they actually changed.
Where AI has the most leverage in e-commerce
Retail has one thing that makes AI value easier to spot: repetition. Saving a little time on one product or one support request means almost nothing. Repeat the same task across tens of thousands of SKUs, orders, or customer conversations, and the math changes fast.
That is why AI adoption in retail often makes the most sense in high-volume work. Large catalogs create plenty of repetitive product-data and merchandising tasks, while order peaks put the same pressure on support and operations. These are the same kinds of problems that shape broader e-commerce development: the useful AI use case is often hiding in work people already do hundreds or thousands of times.
Before the next AI project starts
A few early decisions make the rest of the project much easier to judge:
- Pick a problem with a number attached. Know what it costs today.
- Give the result an owner before launch. Someone should still care once the test is over.
- Decide what success means while the answer is still unknown. Moving the target later makes every result easier to defend and harder to trust.
- Plan for a good outcome. If the pilot works, there should already be a path to funding what comes next.
So when “we should probably do something with AI” comes up again, this time it comes with a plan.
FAQ
Choosing the wrong problem is often where the trouble starts. After that come the familiar ones: weak ownership, messy data, integrations that turn out bigger than expected, and no clear way to judge the result.
Measure what changed for the business, not just what the AI did better. Faster support matters if it cuts cost or helps the same team handle more work. Better search matters if customers actually find and buy more.
A problem you can put a number on. Know what it costs today, what improvement would make the investment worthwhile, and how soon you can get evidence.
In retail and commerce, AI usually has to work inside a live selling operation. That means dealing with product data, pricing, inventory, orders, customer behavior, and seasonal traffic, often across several connected systems. The closer the use case gets to search, recommendations, checkout, or support, the more its quality depends on how well those parts work together.
Someone close to the result you want to change. If the pilot is meant to improve merchandising, merchandising needs a real stake in it. The same logic applies to support, search, or operations.
Enough to represent the cases the AI will actually see. For product tagging, that could mean a few thousand products across your main categories. For customer support, a few months of tickets covering the most common request types may be enough for a first test. If whole categories, seasonal peaks, or common edge cases are missing from the sample, you probably need more.