WTWilly Tai
Case study · live demoAssurance · AI Procurement

Every AI vendor says yes. The scorecard only pays for evidence.

Organisations are buying AI with questionnaires built for buying laptops, and vendors have learned that "yes" costs nothing to say. This instrument runs 24 questions across six domains (data rights, security, model risk, performance evidence, resilience and exit, compliance) with one structural idea: a claim without an artifact is capped at a third of its value. Four questions are kill questions, where no score elsewhere compensates for a no. Three synthetic vendors demonstrate why the arithmetic matters. Runs live in your browser.

01 / The pain, and why the naive approach fails

The demo is not the diligence, and the questionnaire that accepts adjectives is theatre.

AI procurement fails in a specific way: the product demo is dazzling, the security questionnaire comes back all-yes, and the contract arrives with the vendor's paper. Nobody asked to see the audit clause, the model-change notice period, or an accuracy number measured on anything resembling the buyer's data. The naive fix is a longer questionnaire, which produces more yeses. The structural fix is to change what a yes is worth: answers score a third of their value until an artifact backs them: a dated SOC 2 report, a clause number in the MSA, a pilot protocol with acceptance criteria. And some questions are not scoring questions at all. A vendor that refuses audit rights, or will not contractually exclude your data from training, or cannot commit to incident notification and exit portability, has answered the engagement question, whatever its total.

02 / One worked example

Three vendors, and the gap that is the finding.

The scorecard's headline is two numbers side by side: what a vendor claims and what it evidences. Veritan Systems, the boring one that shows its work, scores 85 and 85: claims and evidence travel together, verdict proceed. Forma Analytics, young and honest, scores 58 and 58: the gaps are admitted out loud, kill questions all clear, verdict proceed with conditions, and the conditions write themselves from the missing artifacts. Cloudmind AI, the demo that dazzles, claims 77 and evidences 31, and fails a kill question: no customer audit rights, "SOC 2 suffices". For a MAS-regulated buyer that single no ends the evaluation, and the scorecard says so without apology.

VendorClaims / evidenceVerdict
Veritan Systems · shows its work85% / 85%proceed
Forma Analytics · honest about gaps58% / 58%proceed with conditions
Cloudmind AI · dazzles, attaches nothing77% / 31%walk away · kill question failed
The instrument does not call the slick vendor a liar. It calls the gap between 77 and 31 unpriced risk, and prices it. That reframing is what makes the scorecard usable in a real procurement: it never accuses, it only refuses to pay for adjectives, and it hands the buyer an auto-generated list of exactly which artifact to request in the next meeting.
03 / Numbers & honest trade-offs

What it does, and what it does and doesn't prove.

24
questions, 6 domains
The arithmetic printed at the top of the instrument.
4
kill questions
Audit rights · training exclusion · incident SLA · exit portability.
×0.33
the claims cap
What a yes is worth until the artifact arrives.
2
numbers that matter
Claimed vs evidenced; the gap is the finding.

Where this stands, honestly. The vendors are synthetic and deliberately shaped to span the market's three real archetypes. The questions and weighting are a starting instrument, not a standard: a real engagement tunes them to the buyer's sector and reads the artifacts rather than counting them, because a dated SOC 2 report can still describe a scope that excludes the product you are buying. What transfers unchanged is the conviction underneath: diligence that accepts adjectives is theatre, and the buyer's leverage is highest in the one week before signature. The proposition in one line: I sit on the buyer's side of the AI gold rush with an instrument that converts vendor charm into a list of documents.

What this says about how I work.

I have sat on both sides of this table: selling technology programmes into agencies and banks, and training as the auditor who reads what the salesperson left behind. The scorecard is that second training applied to the first market. Its design choices are audit instincts: pre-stated arithmetic, evidence over assertion, red lines that cannot be averaged away, and an output that is a work list rather than a grade. Firms advising clients on AI adoption need exactly this buy-side discipline, and it pairs with the Go-Live Gate: one instrument decides whether to sign, the other decides whether to switch on.

See it run, live.

Score all three vendors, watch the claims-versus-evidence gap tell the story, read which kill question ends the Cloudmind evaluation, and take the auto-generated meeting list from the gaps.

Open the scorecard →Runs client-side over synthetic vendor profiles: no sign-in, nothing to install.