← All insights

Case analysisLens: United States4 min read

Don't automate the old workflow. Chatham Financial redesigned trade validation around evidence and exceptions

Chatham Financial cut an early trade validation review from about 30 minutes to under 4 by asking what evidence each decision needs, not where to put an agent. Its model routing follows suit.

Listen to this article · 6 min

AI-generated narration of the full article.

A magnifying glass resting on a pale surface, in La Madre duotone, beside the words Evidence first
Photo: rawpixel (CC0)

Most enterprise AI projects start with a question that sounds practical: where can we put an agent? Chatham Financial, a capital markets advisory firm, starts somewhere else. Its CEO, Matt Henry, describes the approach as starting “with the outcome we want” and identifying “where judgment matters” before deciding how to use AI.

OpenAI published the case on October 2. All figures below are reported by Chatham and OpenAI.

What Chatham built

Chatham uses OpenAI’s Codex to build internal and client-facing software and GPT-5.6 to power AI features across three areas:

  • Process Zero, its workflow reengineering service. For each workflow, Chatham identifies the minimum inputs and evidence required, determines where human judgment is essential, and decides how AI and AI-built tools should support the work.
  • Trade validation, an early Process Zero result. Chatham’s Controls and Data Integrity team checks that each system record matches what the client authorized and what was executed. The new application gathers supporting transaction evidence, compares key terms and flags discrepancies for review. In early measurement, review time fell from roughly 30 minutes to under 4.
  • Chatham Onyx, its capital markets operating system, where clients and advisors work from connected, governed data and use AI “without losing traceability to the underlying source.”

Two details deserve as much attention as the time saving. Chatham says it is comparing the application’s results with experienced reviewers on real transactions before expanding automation. And the application does not decide whether a trade is correct. It assembles the evidence and points a person at the exceptions.

Evidence first, then automation

This is a different design target from “an agent that validates trades.” It breaks the work into what can be checked and what must be judged:

  1. Define the outcome: every record matches authorization and execution.
  2. List the minimum evidence: confirmations, client instructions, execution data.
  3. Automate collection and comparison: pulling documents and matching terms is where software is fast and consistent.
  4. Route exceptions to experts: a discrepancy is a judgment call, so it goes to a person with the evidence attached.
  5. Measure against the old way before widening scope.
Evidence-first redesign, as Chatham describes itDESIGNAUTOMATEJUDGE AND PROVE01Outcome toprotect02Minimumevidenceneeded03AI gathersandcompares04Discrepanciesflagged05Expertdecidesexceptions06ComparedwithreviewersbeforescalingDefined by the businessDone by software
  1. Outcome to protect
  2. Minimum evidence needed
  3. AI gathers and compares
  4. Discrepancies flagged
  5. Expert decides exceptions
  6. Compared with reviewers before scaling
  • Design: Outcome to protect · Minimum evidence needed
  • Automate: AI gathers and compares · Discrepancies flagged
  • Judge and prove: Expert decides exceptions · Compared with reviewers before scaling

Defined by the businessDone by software

The application does the legwork and surfaces exceptions. People keep the judgment, and the automation earns its scope against them.

The pattern fits what we called agent-ready work in our framework for finding where agents fit: an existing process, a clear record of right and wrong, and a review step people already own.

Routing is part of the same design

The case also shows model routing in a regulated firm. Onyx uses GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.4 and GPT-4.1: simpler analysis and non-production testing go to cheaper models, and GPT-5.6 is reserved for complex tasks where accuracy carries more value. On Chatham Vibes, the internal platform where employees build their own applications, AI features run on GPT-5.6 Terra by default, with Sol as an optional upgrade configured per application.

That default-plus-upgrade rule is a governance decision as much as a cost one. It keeps citizen-built apps on a known, cheaper model unless someone deliberately asks for more, and it makes the upgrade visible in configuration. It is the practical form of the argument in our analysis of routing by task in OpenAI’s GPT-6 guide.

Why this matters for regulated firms

For a U.S. bank or a firm that serves banks, the comparison with experienced reviewers is not just good practice. The Federal Reserve’s SR 11-7 guidance on model risk management expects models to be validated before use and monitored after, with effective challenge from people who understand them. Running a new validation tool side by side with expert reviewers, on real transactions, is the evidence that guidance asks for. It is also what lets the firm expand automation later without guessing.

Traceability is the other half. A system that compares terms and points to the underlying document leaves an audit trail by construction. A system that answers “this trade is fine” in prose does not.

What to do now

  1. Pick one control process where people already check records against sources.
  2. Write the outcome and the minimum evidence before choosing a model or a tool.
  3. Automate collection and comparison first; leave judgment on exceptions with people.
  4. Run side by side with experienced reviewers on real cases, and record agreement and disagreement.
  5. Set a default model and an explicit upgrade path for internal builders, and log upgrades.
  6. Expand scope only on measured results, one trade type or document type at a time.

The bottom line

Chatham’s 30-to-4-minute figure will get the headlines, and it is an early, self-reported number. The more durable lesson is the order of operations: define the outcome, name the evidence, automate the legwork, keep judgment where it belongs, and let measured agreement with experts decide how far automation goes.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case