The first 90 days of an AI programme in commerce
A 90-day playbook for a commerce team starting its first serious AI programme. One workflow. One number. Honest baselines. No pilot fatigue.
Published 10 June 2026 · Last reviewed 22 July 2026 · BAICA Research
Why 90 days and not 12 months
Commerce AI programmes fail two ways: they stall in a year-long strategy phase, or they launch ten pilots that never make it to production. A 90-day cycle forces the team to pick one workflow, ship it end to end, and either scale it or kill it. Repeat the cycle four times a year and you have a programme.
Days 1 to 14: inventory and pick one workflow
- Run the AI inventory from our EU AI Act checklist. You need a shared picture of what already exists.
- List every candidate workflow: search, recommendations, generative descriptions, chatbot, moderation, forecasting, ad copy.
- For each, note the current owner, the top three pain points, and the metric that would move if AI worked.
- Pick one. The one where the metric already gets weekly attention from leadership, where you already have data, and where you already have a baseline. Not the most impressive, the most measurable.
Days 15 to 30: measure the baseline honestly
Two weeks of clean, uninterrupted measurement of the current workflow. This is the boring, unglamorous phase everyone wants to skip. Skipping it means every future result is arguable.
- Freeze the workflow. No experiments running in parallel.
- Instrument the funnel. If the primary metric is add-to-cart rate from search, measure it by segment and by device, not as one blended number.
- Write the baseline down and get the business owner to sign off on it. This is the number every future comparison is against.
Days 31 to 60: pilot on a small, real slice of traffic
- Build or buy. If a vendor solution covers 80% of the use case, run it as a pilot before you build. See how to evaluate an AI vendor.
- Pick a traffic slice: one country, one category, one device, or a 5% hold-out group. Not the whole site.
- Set a stop rule before you go live: what result kills the pilot early, what result extends it, what result triggers rollout.
- Ship the transparency and oversight controls at the same time as the model, not after. Chatbot notices, generative content labels, human review routes for anything customer-facing.
Days 61 to 80: read the results without flinching
- Compare against the frozen baseline, not against "before we started paying attention".
- Segment the results. A blended number often hides a large positive effect for one group and a small negative effect for everyone else.
- Look for the second-order effects: did average order value move, did returns move, did customer service tickets move, did the team's throughput move.
- Have the debate about what the number means with the business owner in the room. Not in a written report they will skim.
Days 81 to 90: decide, document, and pick the next one
- Three options: roll out, iterate for another cycle, or shut it down. All three are legitimate outcomes.
- Write the case study: goal, baseline, approach, result, cost, effort, what you would do differently. Two pages.
- Contribute an anonymised version to the BAICA Open Case Library. Your peers will do the same and everyone gets smarter.
- Pick the next workflow and start day 1.
The traps to avoid
- Running five pilots at once. Nothing gets across the line, everyone gets tired, leadership loses interest.
- Skipping the baseline. Every result becomes a marketing argument, not a business decision.
- Optimising a metric no one owns. If the business owner does not care about the number today, they will not care about the AI-driven improvement in it tomorrow.
- Waiting for the perfect stack. The first cycle can run on the tools you already have.
- Treating "the AI works" as the outcome. The outcome is the business number moved and the team is ready to run the next cycle.
What good looks like after four cycles
A year in, a team that runs this playbook has four case studies, one or two AI systems in production behind real metrics, a governance record that would survive a regulator's visit, and a shared muscle for shipping AI without drama. That is the goal. Everything else is decoration.
Put the guide to work.
Every guide is free and open-licensed. If a question keeps coming up in your team and we have not covered it, tell us and it goes on the list.