Independent AI model verification belongs in your RFP

•Riz Pabani

Most AI vendor due diligence still runs through the standard supplier questionnaire: ISO 27001, SOC 2, data residency, and whether your data will be used for training. Those questions were written for ordinary software. None of them asks for independent AI model verification, meaning someone outside the vendor who has tested the model itself.

The companies that make the models are now giving buyers a reason to ask. On 15 September 2026, OpenAI’s global policy chief Chris Lehane told reporters that OpenAI, Anthropic and Google DeepMind had been working together on AI safety for weeks. He also said OpenAI supports the provision in the US FRONTIER Act that would force the largest labs to let licensed independent verification organisations in. Anthropic and OpenAI have both committed to embedding outside evaluators. Google hasn’t gone that far, but on 24 September The Information reported that all three labs aim to launch a standards body by the end of 2026 or early 2027. It would set qualifications for independent auditors of models and labs.

When the model makers say an outsider should check the model, “who has independently evaluated this?” is no longer an awkward question to put in an RFP. Here’s what’s been proposed, what it doesn’t cover, and six questions I’d add to your next AI procurement.

How we got here

DateWhat happened
14 July 2026Demis Hassabis calls for an independent, US-led standards body for frontier AI
23 July 2026Reps Jay Obernolte and Lori Trahan introduce the bipartisan FRONTIER Act (H.R. 9925)
12 September 2026Dario Amodei publishes “We Must Pace the Frontier” and commits Anthropic to embedding outside evaluators
12 September 2026Sam Altman says OpenAI will do the same. Elon Musk posts “Dario is right”
14 September 2026The Information reports the three labs are discussing an industry standards body
15 September 2026Lehane confirms the joint safety work and OpenAI’s support for the verification provision
24 September 2026The Information reports the labs aim to launch the standards body by the end of 2026 or early 2027, covering auditor qualifications and incident reporting
29 September 2026AI and tech leaders sign a voluntary accord at the White House

The Anthropic commitment is more concrete than most safety pledges. According to TechCrunch’s write-up of the essay, evaluators from organisations such as METR get company badges, desks and laptops. Their access is “mostly comparable to what internal risk assessment teams have”, with exceptions when required by law or contracts. Note that carve-out. It’s exactly the part a buyer would want to understand.

Amodei had suggested a narrow government waiver so the labs could coordinate without breaking competition law. Lehane said they don’t need one.

Not everyone is applauding. Cohere’s CEO Aidan Gomez called the arrangement “a cartel by another name”. FTC Chair Andrew Ferguson said that when companies ask Washington for regulation and an antitrust exemption at the same time, “all of my alarm bells go off”.

President Trump said on 14 September that “AI taking over the World, destroying Humanity, and all other things bad, is a HOAX”. David Sacks, who co-chairs his science and technology advisory council, said “it is on them to make their products safe”. The Attorney General, Todd Blanche, said he wouldn’t prejudge whether the labs should get an antitrust exemption. Then, on 29 September, leaders including Dario Amodei and Sundar Pichai signed a White House accord that Trump called “morally” binding. It is voluntary, and House Speaker Mike Johnson described it as “a statement of principles”.

Both things can be true. The coordination may suit the incumbents commercially, and independent verification is still something buyers should want.

What the FRONTIER Act would require

The verification clause applies to “very large frontier developers”. The bill’s section-by-section summary sets that at more than $5bn in gross revenue plus at least $10bn of AI development spending over 36 months. In practice, a handful of companies.

Each of them would have to retain a licensed independent verification organisation. The verifier assesses whether the lab’s safety framework, governance, risk monitoring and mitigations achieve acceptable levels of catastrophic risk mitigation. It gets access to unredacted materials, records and assessments at any time. If it finds a model poses an imminent catastrophic risk, it must refer it to the government within 72 hours.

A new Under Secretary of Commerce for AI Security would license the verifiers. The Government Accountability Office would report every year on whether they are staying independent of the industry. That last provision tells you Congress already expects the obvious problem: verifiers depending on the companies they check.

The bill is still in committee. Senate Majority Leader John Thune has said getting anything done “in the near term is going to be challenging”. Treat it as a signal of where US rules are heading. It won’t be on anyone’s compliance calendar this year.

If you buy from London or Frankfurt, some of this already exists. The EU’s General-Purpose AI Code of Practice expects signatories to appoint independent external evaluators for models with systemic risk, with limited exceptions. OpenAI, Anthropic and Google all signed it in 2025. The evidence may be sitting there. Ask for it.

What independent verification won’t tell you

The FRONTIER Act verifiers are looking for catastrophic risk, which in frontier AI policy usually means things like weapons uplift and large-scale cyber attacks. That matters. It tells you nothing about whether the model drafts an accurate credit memo, or whether it quietly drops a client name into a summary it shouldn’t.

That second kind of check is use-case evaluation, and it belongs to the buyer. UK banks already know the principle. The PRA’s SS1/23 on model risk management expects independent validation of models, including models bought from third parties. A foundation model inside a lending workflow is a model. The lab’s verifier covers the bottom layer. Your validation covers what you’ve built on top.

There’s a version problem too. An evaluation applies to a specific model snapshot. Your contract probably names a product tier. A clean assessment of the model a vendor shipped in March says very little about what’s answering your staff in October.

Six questions to add to your next AI procurement

1. Who evaluated this model, and with what access?

Ask for the organisation’s name, the date and the model version. Then ask about access. An evaluator testing through the public API sees far less than one sitting inside the lab with the kind of access Anthropic has described. Ask who paid for the work. There’s no agreed standard yet for who counts as a qualified evaluator. Setting one is part of the labs’ proposed standards body, so until it exists, ask what qualifies this one.

2. Can we see the findings?

A full report under NDA is best. A written summary is the minimum. You want what was found, what was fixed and what was accepted as residual risk. “We’ve been independently evaluated” with nothing behind it gives your risk committee nothing to review.

3. Which model version does our contract cover?

Get the model identifier into the order form. Ask for written notice before the model is replaced or materially changed, and for evaluation evidence on the replacement. Where the vendor offers it, ask for the right to stay on the previous version for a fixed period.

4. How will you tell us about a safety incident?

The FRONTIER Act would make labs report safety incidents to government. The labs’ proposed standards body would also set out how developers should report safety and security incidents. Neither gives you, the customer, a right to be told. Ask for a customer notification commitment with a timeframe, a definition of what triggers it and a named contact.

5. What sits between the lab’s model and our users?

Most enterprises don’t buy directly from a lab. They buy a copilot, a CRM add-on or an analytics tool built on one. The lab’s evaluation doesn’t cover that product’s system prompts, retrieval, tool access or data connections. Ask the product vendor which lab model and version is underneath, and what testing they’ve done on their own layer. Watch that relationship too, because model providers are increasingly building the applications themselves.

6. Who validates our use case?

Decide this before you sign. It could be your model risk function, an independent third party or a specialist partner. It shouldn’t be the same people who built the workflow.

A clause you can adapt

This is a starting point for your legal team to rework. It isn’t legal advice.

For each AI model used to deliver the Services, the Supplier shall disclose: (a) the model identifier and version; (b) any independent evaluation of that model version completed in the previous 12 months, including the evaluating organisation, the level of access provided, and a summary of findings and remediations; and (c) the Supplier’s process and timeframe for notifying the Customer of material safety incidents. The Supplier shall give the Customer no less than [30] days’ written notice before replacing or materially changing a model, together with equivalent evaluation evidence for the replacement.

Where Exponential Partners fits

We build software on top of these models for clients, so we sit on the use-case side of that line. Before anything we build goes live, we test its outputs against the client’s real work.

On an invoicing and payroll system we’re building for a staffing firm, that means parallel runs. Each week the system produces the invoices, the finance lead produces hers the old way, and we compare the two. The first run came back with zero exceptions, and that worried us more than it reassured us. It was too early for there to be no mistakes. The next runs found them. An export from the scheduling system arrived with its column headers in English instead of French, and the import rejected it. One invoice came out with the wrong subtotal, which she spotted because she knew what the pre-tax total should be. A night-premium rule was counting one hour where it should have counted two.

None of that would show up in a model evaluation. The model was fine. The problems sat in the data, the rules and the edge cases only the client knew about. That’s the layer a lab’s verifier will never see, and the one you have to test yourself.

If you’re putting together an AI procurement pack or RFP and want someone to pressure-test the evaluation section before it goes out, we’ll tell you plainly whether what you’ve got is already enough. Schedule a Conversation.

Riz Pabani

Execution, Exponential Partners

Riz helps executives and their teams figure out where AI actually creates value — then builds the capability to capture it. Former Goldman Sachs, Nomura, and Bank of England; led partnerships at the Cardano Foundation. MIT-certified in AI products.

Related Insights