In regulated AI, the model never grades its own work. Proof has to come from outside it
Dual Run, verification tools and calibrated reviewers share one rule: evidence that an AI output is right must come from something independent of the model. A test for regulated teams.
Listen to this article · 5 min
AI-generated narration of the full article.

Two announcements this week look unrelated. Google’s mainframe modernization tools use Gemini to read legacy code, then rely on Dual Run, a deterministic replay of live transactions, to prove the new system behaves the same. OpenAI’s text watermark comes with a frank list of what detection cannot prove. Put them next to what we have covered over the past week and one rule emerges for regulated organizations: evidence that an AI output is correct has to come from something independent of the model that produced it.
This is a point of view, and we will say where it rests on fact and where it is our reading.
The evidence behind the rule
- Modernization. In Google’s Dual Run design, Gemini’s reading of legacy code is a set of hypotheses; the proof is what both systems do with the same real transactions. Intesa Sanpaolo describes Dual Run as the evidence it builds for leadership, internal control and regulators.
- Provenance. OpenAI’s own numbers show watermark detection falling from about 92% to 66% when 10% of words are swapped. A probabilistic signal cannot carry a decision about a person by itself.
- Engineering. In GPT-Synopsys, the model operates chip design tools, but verification tools have the last word.
- Finance operations. Snowflake’s accounts payable system keeps matching and calculations deterministic, and Chatham Financial redesigned trade validation around evidence and exceptions.
- People. IBM’s study of HR leaders shows that a reviewer who cannot challenge the AI is not a control.
Four sources of proof that count, three that do not
- Comparison with a trusted system
- Domain verifier or rule engine
- Expert-curated reference set
- Qualified reviewer who can reject
- The model checking itself
- Uncalibrated LLM judge
- Vendor benchmarks
Counts as evidence
What counts. A deterministic comparison against a system you already trust (Dual Run, reconciliations, parallel runs). A domain verifier that does not share the model’s failure modes: simulators, rule engines, calculation engines, sign-off checks. A reference set built by experts for your own cases, held out from anything the model was tuned on. A qualified reviewer with the time, information and authority to reject the output.
What does not count on its own. The model explaining why it is right. Another LLM judging the first, unless you have measured how often it agrees with your experts. Benchmarks published by the vendor, which test the vendor’s tasks, not yours.
Regulators already speak this language
None of this is new to regulated industries; AI just makes it urgent.
In U.S. banking, the model risk guidance the Federal Reserve, OCC and FDIC reissued in April 2026 (SR 26-2, which replaced SR 11-7) keeps the idea of effective challenge: critical analysis by objective experts with the expertise, the independence and the organizational standing to force changes. It also states that generative and agentic AI models are outside its scope, leaving banks to govern them through their broader risk practices. That gap is where the principle should be carried over by choice. A model that validates itself is the opposite of effective challenge, whatever the guidance’s scope says.
In medical devices, the FDA’s Computer Software Assurance guidance, reissued in February 2026, asks manufacturers for a risk-based approach to establishing confidence in production and quality system software, with more rigor where risk is higher. The higher the risk of an AI-assisted step, the more the evidence should come from independent testing rather than from the tool’s own reports.
What follows for architecture
Our reading, and it is a view rather than a fact: the regulated AI systems that reach production first will be designed as two systems, one that produces and one that proves. The proving system is often old (a reconciliation, a validation protocol, a review step) and deliberately not AI. Budget for it from the start; it is usually where most of the effort goes, and it is what the auditor will read.
What to do now
- For every AI step in a regulated process, name the independent evidence that shows its output is right.
- If the only evidence is the model or a similar model, do not move the step past “recommend.”
- Build expert reference sets from your own cases and keep them out of tuning.
- Calibrate any LLM judge against qualified reviewers before relying on it.
- Give reviewers authority and time to reject, and track how often they do.
- Package the evidence for the auditor as you build, not after go-live.
The bottom line
AI can read, draft, translate and transform faster than any team. In a regulated organization, it still cannot be its own witness. Design every production use so that the proof comes from outside the model, and the conversation with risk, quality and the regulator gets much shorter.