How to evaluate an AI vendor before you sign

Ten questions to ask any AI vendor, the red flags behind each answer, and a pilot structure that keeps you in control of the outcome.

Published 17 June 2026 · Last reviewed 22 July 2026 · BAICA Research

Why vendor evaluation matters more with AI

A traditional SaaS vendor sells you deterministic software. An AI vendor sells you a system whose behaviour depends on data, prompts, and model updates that can change without notice. The switching cost is higher because your workflows adapt to the model. The compliance surface is bigger because the model touches customer data and can make decisions that affect customers. And the ROI is harder to prove because the baseline moves. All of this makes upfront evaluation the highest-leverage moment in the relationship.

The ten questions

  1. What model or models does the system use, and who trained them? A vendor who cannot or will not answer either has thin technology or is protecting a brittle stack. Ask which parts are proprietary and which are third-party.
  2. What data was the model trained on, and does it include our customer data? The answer should be specific. If your data is used to train a shared model that other customers benefit from, you should know before signing.
  3. Where is our data processed and stored, and who has access? Ask for a data flow diagram. If the vendor uses subprocessors (OpenAI, Anthropic, Google), those must be listed with the region of processing.
  4. What happens to our data when we cancel? Deletion within a specific number of days, in writing, or the answer is not good enough.
  5. How does the system behave when it is uncertain? A good answer describes fallbacks: escalation to a human, a safe default, a refusal message. A bad answer talks about accuracy percentages.
  6. How do you evaluate quality, and can we see the evaluation set? Vendors who take quality seriously have a real eval suite. Ask to run your own inputs through it during the pilot.
  7. How often do models change, and what changes with them? Model updates can silently shift tone, safety, and accuracy. Ask for a change log and a notice window before updates.
  8. What is your position under the EU AI Act? Are you a provider or a deployer for our use case? If they do not know, that is the answer. See our EU AI Act checklist.
  9. What does the SLA cover, and what does "uptime" mean when the model itself is degraded? A 99.9% API uptime SLA that ignores model quality is close to worthless for AI.
  10. Give us three reference customers in commerce, at our scale. Talk to at least two of them without the vendor in the room.

Red flags to watch for

  • Case studies without numbers, or numbers without a baseline
  • Refusal to run a paid pilot on your data before the annual contract
  • No named data protection contact and no DPA on request
  • Pricing that scales with input tokens with no cap
  • A demo that only works on the vendor's demo catalogue
  • "We use the latest models" as a substitute for a real technical answer

A pilot structure that protects you

Sign a fixed-scope, fixed-price pilot before any annual contract. Structure it like this:

  • Scope: one workflow, one traffic segment, one success metric. No "let's also try" additions during the pilot.
  • Baseline: measure the current version of the workflow for two weeks before switching anything on. Without a baseline, the pilot will "work" no matter what happens.
  • Duration: four to eight weeks of live traffic, not a demo period.
  • Exit: define in writing what result buys the annual contract and what result ends the pilot with no obligation. Do this before you like the vendor.
  • Data: a separate pilot DPA. Data collected in the pilot is not used to train shared models by default.

What to do after the pilot

Whatever the outcome, write it down. The pilot report is worth as much as the pilot: it becomes the baseline for the next vendor conversation and the evidence the board will ask for. Publish anonymised versions of your outcomes to the BAICA Open Case Library and help the next team avoid the same mistakes.

Put the guide to work.

Every guide is free and open-licensed. If a question keeps coming up in your team and we have not covered it, tell us and it goes on the list.