AI vendor packs look alike. Appealing use cases, client logos, a compliance pledge. For a CIO, the useful question is drier: what happens the day the tool is wrong, the contract changes, or you want out. If the answer fits in a marketing sentence, risk is underweighted.
Five inspection axes
Five inspection axes help leave the fog. Data location and sub-processors. Ability to log prompts, contexts, and actions, and keep those logs under control. Explicit functional scope, including what the system will never do. Reversibility: export of data, configurations, histories. Real cost at scale, tokens, storage, support, integration, human recovery. Those axes often say more than a model benchmark. They align with our operational criteria for data sovereignty.
Soft lock-in deserves special attention. A knowledge base that will not export, workflows you cannot reproduce, an API that changes without notice. None of that is necessarily illegal. It is simply expensive to discover too late, especially once the business already relies on the tool every day.
Testing the response under pressure
It also helps to test the vendor’s response to a simulated incident. Who can be reached. In what delay. With what technical depth. A well filled security grid does not replace the ability to collaborate under pressure. Teams that have already lived through an incident know this: the contract reads differently on the day something breaks. On running agents, see also governance, logs, and escalation.
A serious evaluation takes time. It saves more time later. Better demanding criteria before signature than exceptions after an incident. The same grid can also be used internally, to judge the bricks you build yourselves. What would not pass the test at a third party should not pass it at home either. The same grid also clarifies the choice between data sovereignty and a well-framed cloud hosting model.