The model changed. Prove the answers didn't.
Independent verification for AI vendors selling into regulated finance. We hold a test set built from your customer's data, re-run it on every model change, and publish a record their risk team reads directly.
- 23 Oct
- OpenAI retires roughly twelve model families on one day.
- 60 days
- Anthropic's floor for deprecation notice. That is the whole warning.
- No notice
- What an unpinned alias gets when the provider repoints it underneath you.
The problem
A question you cannot answer yet
You migrate because a provider made you, and you will do it again, on their calendar, forever. Each time, your customer's risk team asks whether your system still works.
Today you answer with a dashboard you control, from a test suite you can edit, with no record of what you ran last quarter.
You are marking your own homework.
Why it persists
Nothing you already pay for answers it
Complements, not substitutes. The seam between them is why this exists.
- Governance platforms
- Record that the model version changed. Never whether the answers survived it.
- Eval tooling
- Measures a run your own team starts, with a scorer you write and can change afterwards.
- Assurance consultancies
- Independent about your process, once a quarter. Nobody re-runs the test.
A record that never measures the output. A ruler you hold. Independence pointed at your paperwork.
What we do about it
An operated service, not a tool you have to run
We build the test set from your data, and we hold it
Mined from your documents, tickets and past decisions, confirmed once by your expert in under four hours, and versioned from that day on.
A case you failed cannot quietly leave the set.
We run it to a fixed schedule, not on request
The full set weekly, a smoke set daily, and a confirm run within the hour of any change we detect. Nobody starts a process, and no port opens on your side.
A test that only runs when you start it is a test whose timing you control.
We produce a record built to be read by your customer
Stamped with the fingerprint of the model that actually answered, hash-chained, timestamped, and verifiable in the reader's own browser - limitations stated before results.
Nothing in it rests on trusting us.
Your customer reads it directly
A read-only portal for your customer's risk team - about twenty minutes a quarter, without a call with you first.
Absence is visible: no run in 92 days is itself a finding.
How a case actually reaches your system, and what comes back
Product · Verification
One tile does the work of the whole pitch
Not a score and not a percentage. A state, resolved by comparing the fingerprint of what is in production against the fingerprint recorded on the last sealed run.
Verified against current model
Current
The system answering your customers right now is the one the last sealed run was made against.
Verified against current model
Stale
Production has moved since the last sealed run. Nothing is known to be wrong, and nothing is known to be right.
Verified against current model
Never
Nothing has been verified against what is running. Said plainly rather than shown as a blank.
Only a party holding the set across the change can say the thing in production has been verified.
Explore Verification - the set, the runs, and the regression screen
Product · Assurance
A record read by somebody who does not work for you
A read-only portal for your customer's risk team, held for the contract term under standing disclosure - where the cadence is stated, the runs are counted, and a gap is a finding rather than a blank.
Explore Assurance - the portal, the evidence pack, and what we don't claim
Security
Built to pass your customer's security review
- No inbound path - every connection is yours, outbound, on 443.
- Your production traffic never reaches us.
- PII removed before a case enters the set - measured, with published numbers.
- Tenant isolation enforced twice - in the application, and again inside Postgres.
The full review, answered - and the certifications we do not claim yet
This is for you if
You sell AI into regulated finance
Your customers are banks, insurers or brokers, and their risk function has a say in whether you get bought.
Somebody is accountable if an answer is wrong
Not a disclaimer - a named person on your customer's side who has to answer for what your system said.
Your models have dates on them
A published retirement date, or an alias with none, which a provider can repoint underneath you.
You would rather not build this yourself
A quarter of an engineering team on eval infrastructure your customer would still not accept as independent.
Start here
Begin with what doing nothing costs you
Paste the model ids out of your configuration. You get each one's published retirement date, the days remaining, and which of them are aliases a provider can repoint without telling you. No account, no integration, nothing stored about you.
If you already know the answer
The offer, with the number on it
Two plans, with the numbers on the page rather than behind a call - and one question back: whether they are roughly right.