Buying AI software is harder than buying conventional software, and most procurement processes aren't built for it. The evaluation criteria that work for an ERP or a CRM don't fully translate. Models change, performance degrades in ways that aren't obvious in demos, pricing structures are unfamiliar, and lock-in risks are real but often underestimated. This guide walks through a rigorous vendor evaluation process that accounts for the specific characteristics of AI products.
Before you look at vendors: define the problem
The most common mistake in AI procurement is starting with the solution and working backwards to a problem. "We need an AI tool" is not a procurement brief. "We want to reduce the time our support team spends on first-contact resolution by 40%, using AI to handle common query types automatically" is a procurement brief.
Before approaching any vendor, be specific about: the problem you're trying to solve, the workflow you're trying to change, the metrics you'll use to evaluate success, the data you have available to power the solution, and the constraints you're operating under (budget, timeline, integration requirements, regulatory environment).
This specificity does two important things. First, it lets you evaluate vendors against your actual use case rather than their best demo. Second, it protects you from vendors who are skilled at expanding the problem to fit their solution, a common pattern in enterprise AI sales.
The four evaluation dimensions
1. Technology
The technical evaluation of an AI product should cover more than feature completeness. Key questions:
- What is the underlying model? Is it a proprietary model, a fine-tuned foundation model, or a thin wrapper around a commodity API? The answer matters for both performance and dependency risk.
- How does performance degrade on edge cases? Most AI tools look good in demos constructed to show them favourably. Ask to test on your actual data and your actual edge cases. Measure hallucination rates, error rates, and confidence calibration specifically for your use case.
- What is the model update cadence? When the vendor updates the underlying model, does performance on your use case change? How are you notified? Can you pin to a specific version?
- What are the latency characteristics? For any real-time or customer-facing application, latency matters. Test this under realistic load conditions, not ideal laboratory conditions.
2. Data handling
This is often where deals should be blocked but aren't, because organisations don't ask the right questions early enough. Essential data questions:
- Does the vendor use your data to train or improve their models? Many consumer-tier AI tools do this by default. Enterprise agreements typically allow you to opt out, but you need to ask explicitly and get the answer in writing.
- Where is your data processed and stored? Data residency matters for GDPR compliance, for regulated industries, and for organisations with customers in jurisdictions with data sovereignty requirements.
- What happens to your data if you terminate the contract? How is it deleted? Within what timeframe? What audit trail is provided?
- Who at the vendor has access to your data? Under what circumstances? For what purposes?
If a vendor is unwilling or unable to answer these questions clearly in writing, that tells you something important. A reputable enterprise AI vendor should have comprehensive, accessible documentation on their data practices.
3. Commercial structure
AI pricing models are less standardised than traditional software, and the differences matter enormously for total cost of ownership:
- Per-seat vs consumption-based pricing. Per-seat pricing is predictable but may be wasteful if usage varies widely across users. Consumption-based pricing can scale efficiently but creates cost uncertainty, a surprisingly successful deployment can generate a surprisingly large invoice.
- What counts as a "unit" of consumption? Token-based pricing, query-based pricing, and document-based pricing all have different implications for your expected cost at different usage levels. Model this carefully against your actual anticipated usage.
- Contract flexibility. What happens if you want to scale up significantly? Scale down? Exit entirely? Are there minimum commitments? What are the notice periods?
- Audit rights. For any significant AI deployment, you should have the right to audit usage, costs, and compliance. Make sure this is in the contract.
4. Company stability
The AI vendor landscape in 2026 contains a significant number of companies that will not exist in their current form in 2028. Before making a significant commitment:
- Funding runway. For startups, understand their funding situation. A vendor burning cash rapidly with uncertain prospects for the next funding round is a dependency risk.
- Reference customers. Ask for references from customers in your sector, at your scale, with similar use cases. Then actually call them. Ask specifically about what went wrong, not just what went right.
- Support quality. Test the support process during evaluation. How quickly do they respond? How competent are the responses? What does the escalation path look like?
- Acquisition risk. If the vendor is acquired, what happens to your contract, your data, and your pricing? This should be addressed in the contract.
"The best AI vendor evaluation question is: 'Tell me about a customer deployment that didn't go to plan and what you learned from it.' The answer reveals more than any polished demo."
Structuring the proof of concept
Never skip the PoC. AI products that look compelling in demos regularly underperform against real-world data and real-world workflows. A well-structured PoC typically runs four to six weeks and covers:
- Scope: Define a specific, bounded use case that's representative of your broader use case but small enough to evaluate thoroughly in the timeframe.
- Success metrics: Agree these with the vendor before you start. Don't let them be redefined mid-PoC.
- Data: Use real data if possible, or data that closely mirrors the characteristics of your real data. Synthetic or sanitised data can mask problems that will surface in production.
- End users: Include real end users in the PoC, the people who will actually use the system day-to-day. Their feedback on usability and workflow fit is as important as the technical metrics.
- Failure modes: Deliberately test edge cases and failure modes. What happens when the AI encounters input it hasn't seen before? When it's wrong, how does it fail, gracefully or catastrophically?
The build vs buy question
Before committing to a vendor, genuinely interrogate whether a buy decision is the right one. The build option has become materially more accessible with the maturity of foundation model APIs, you can build surprisingly capable AI applications using off-the-shelf models without training anything from scratch.
Buy is generally right when: the problem is common enough that specialist vendors have deep expertise, the compliance and security work is already done, and your competitive differentiation doesn't depend on the AI capability itself. Build is generally right when: the use case is highly specific to your business, your competitive advantage depends on proprietary data and approaches, or vendor lock-in risk outweighs the development cost.
Red flags in AI vendor evaluation
Experience working on AI vendor selections surfaces consistent patterns that should prompt caution:
- Unwillingness to provide a Data Processing Agreement (DPA) before contract signature. Any legitimate enterprise AI vendor should provide this as a matter of course.
- Lock-in clauses buried in Terms and Conditions — particularly around data portability, integration limitations, and termination provisions. Get legal review of any agreement before signing.
- Demo environments that don't reflect production performance. Ask to see performance data from production deployments at comparable customers, not just curated demo scenarios.
- Pricing structures with no ceiling. Consumption-based pricing without caps or alerts can result in significant unexpected costs. Ensure you have budget controls and notifications in place.
- Resistance to technical due diligence. A vendor who won't let your technical team ask detailed questions is a vendor with something to hide.
Negotiating the contract
Once you've selected a preferred vendor, negotiate on terms that protect your long-term position, not just the initial price. Key negotiating points:
- SLAs with teeth: Define uptime, latency, and accuracy SLAs. Ensure there are meaningful remedies (service credits, termination rights) if they're breached.
- Data deletion: Specify timeframes and methods for data deletion on contract termination, and require a written confirmation that deletion has occurred.
- Exit provisions: What does an orderly exit look like? What data export format is provided? What transition support is available? How much notice is required?
- Price protection: If you're committing to a multi-year contract, negotiate protection against significant price increases. Consumption-based contracts should have a rate guarantee for the contract period.
- Model change notifications: If the vendor updates the underlying model in ways that could affect your use case, what notification are you entitled to, and what rights do you have if performance changes materially?
Need an independent view on an AI vendor selection?
I help organisations evaluate and select technology without vendor bias. let's talk →