The category

AI assurance, and the one layer that checks the answer

Four different jobs get sold under this phrase. This page says which one this is, which ones it is not, and what your customer's risk function is working from when they ask you for evidence.

The four layers

Four jobs sold under one word, and what each one answers

The market is not empty, it is fragmented - and three of these four are things a vendor should have. They simply answer different questions, and only one of them answers the one a customer's risk function asks second.

AI governance

ISO/IEC 42001, and the compliance platforms built around it

AnswersIs it governed?

Policies, owners, registers and approvals - evidence that the system is permitted and documented. ISO 42001's Annex A governs the management system rather than any output, so a vendor can hold the certificate and still not know whether last month's model change altered an answer. Worth having, and complementary to this; it is usually the thing your customer already accepted before asking the question we answer.

Model risk management

SR 11-7 in the US, PRA SS1/23 in the UK

AnswersWas the model validated?

The discipline a bank runs over its own models: independent validation, documented assumptions, ongoing monitoring. It assumes the institution owns the model and can open it. When the model sits inside a vendor's product, the validator still has to produce evidence and cannot - which is the specific gap this business sits in.

LLM evaluation

Braintrust, Promptfoo, DeepEval

AnswersDid it score well on the test we wrote?

Good instruments, and this is not an argument against them. What they are not is independent: your team writes the cases, holds the suite, decides when it runs, and can edit a scorer after seeing a result. Every one of those is reasonable engineering practice and every one of them is a reason a validator discounts the output.

AI assurance

Where this sits

AnswersIs it still right, and can somebody else check?

Output correctness, measured over time, against a set built from your customer's own material and held by us rather than by you. Re-run on every model change rather than on request, and published as a record your customer's risk function reads directly. It is the only one of the four that produces evidence the vendor cannot quietly revise.

Why you are being asked

The clauses your customer's risk team is working from

None of these regulates you directly if you are the vendor rather than the institution. They reach you through the contract your customer signs, which is why the request arrives from procurement rather than from a regulator.

CPS 230APRA, Australia
Operational risk for material service providers: audit rights, subcontractor disclosure, and APRA's own access. In force with no transitional relief since 1 July 2026, which is why Australian vendors started being asked this year rather than next.
DORA Article 30(3)European Union
Contracts with ICT providers supporting critical or important functions must grant the financial entity - and its regulator - rights of access, inspection and audit. The obligation lands on your customer, and reaches you through what they sign with you.
Insurance Circular Letter No. 7NYDFS, New York
Where a third party's AI affects underwriting or pricing, an insurer cannot rely on that vendor's self-attestation alone. It is the sharpest version of the question in any jurisdiction: it makes your word about your own system insufficient as a matter of supervision rather than of taste.
SS1/23PRA, United Kingdom
Model risk management expectations, including for models a firm did not build. Narrow in scope - a handful of IRB banks rather than the sector - but those are the customers whose validators ask the hardest questions and whose sign-off takes the longest.
SR 11-7Federal Reserve and OCC, United States
The supervisory guidance most US model risk functions are still organised around: independent validation, documented limitations, and monitoring that continues after go-live. Written for models a bank owns, which is exactly why a vendor-supplied one strands the validator.
EU AI ActEuropean Union
Obligations for high-risk systems reach the deployer as well as the provider, and a deployer cannot document a system it did not build without the provider's material. The compliance deadlines are the reason this moves from a procurement question to a contractual one.

This is a description of what makes the question arrive, not legal advice. The wording that binds you is in your customer's contract, and it is usually stricter and more specific than the regime it derives from.

Asked before a demo

The questions that arrive first

Answered here rather than in a call, because a question a reader has to book a meeting to get answered is a question we would rather they took to somebody else.

What is AI assurance?
AI assurance is independent evidence that a deployed AI system still produces correct output - not that it is well governed, and not that it scored well on a test its own vendor wrote. In practice it means a fixed set of cases with known-good answers, held by somebody other than the vendor, re-run when the system changes, and reported to whoever carries the risk.
Is AI assurance the same as AI governance?
No, and the distinction is the whole category boundary. AI governance answers whether a system is permitted, documented and owned - policies, registers, approvals. AI assurance answers whether its output is still right. A vendor can be fully governed and have no idea whether last month's model change altered an answer.
Does ISO 42001 certification show that our AI answers correctly?
It does not. ISO/IEC 42001 certifies an AI management system - that the processes, roles and controls exist - and its Annex A governs that system rather than the output any model produces. It is worth holding, and it is not an answer to a customer asking whether your system still returns what it returned last quarter.
Who is responsible when a vendor's AI gives a customer the wrong answer?
The organisation that put it in front of the customer. Moffatt v Air Canada settled the point plainly: a company owns what its chatbot said, and could not treat the chatbot as a separate party. That is why a bank or insurer's risk function asks its vendors for evidence rather than assurances - the liability does not travel with the software.
Can we do this with our own eval tooling?
You can measure with it, and eval tools are good instruments. What you cannot do with them is answer the independence question, because your team writes the cases, holds the suite, chooses when it runs and can change a scorer after seeing a result. A customer's validator discounts evidence the supplier controls end to end, whatever the numbers say.
What does AI assurance cost?
Presoja is US$1,500 per month for Verification and from US$4,500 per month for Assurance, flat, with no metering on cases, runs, deployments or readers. The comparable is a model-validation engagement rather than an eval tool subscription, which is typically a five- or six-figure piece of consulting per review.

If the fourth layer is the one you are being asked for, the mechanism is on how it works, the record your customer reads is Assurance, and the limits we state before any result are in the same place as the claims.