
How to tell a rubber stamp from a real oversight chain

Every AI governance document ever written contains a sentence promising that “a human remains ultimately responsible.” The sentence is free. What costs effort is the machinery that makes it true: an oversight chain with named links, write-down rules and the ability to act within minutes instead of meetings.
This piece walks that chain link by link. When you next read a claim about AI-driven investing, check which links the firm describes in concrete terms — and which ones it waves at generically.
Link one: defined authority
Somewhere in the chain there must be a person, or a role with a name in the job description, who can stop the model. Not “monitor” it. Stop it: freeze its outputs, block its orders, pull its recommendations out of the workflow until someone has looked.
The test question is disarmingly simple: who, right now, has the authority to switch this off? In a well-run operation you get an answer within a sentence — a portfolio manager, a chief risk officer, a defined duty officer — because that authority was assigned on paper long before you asked.
Authority also has a geography worth checking. Some firms concentrate halt authority in one duty officer; others distribute a limited version of it to every portfolio manager whose book is touched. Both designs can work. What cannot work is an authority that exists in the governance manual but is wired to no actual switch in the production system — a stop permission nobody has ever tested is a rumour, not a control.
Link two: escalation paths with time bounds
Oversight fails at the speed of delay. A model that flagged its own drift yesterday is harmless; a model whose alert sits in a queue until next Tuesday’s meeting is not. The documents to look for are the service rules: escalation thresholds, who gets paged, and the maximum time allowed before a human must respond — with consequences defined for what happens if the deadline passes unanswered.
The phrase “duty officer” should appear somewhere. If escalation is illustrated by a flow chart ending in a committee, ask what the committee does on a Friday night — markets do most of their embarrassing work outside office hours, and a response process that exists only in the organigram responds to nothing.
Well-crafted escalation ladders also say what happens to unacknowledged alarms. Does the drift threshold page a named person, then their deputy, then a supervisor? Does a third-level escalation de-risk the model automatically after so many unanswered minutes? An escalation rule with no default action is a staircase that ends in mid-air.
Link three: logged interventions, four eyes on changes
Good oversight leaves a sediment of records. Curbing an output, overriding a recommendation, retraining a model: each leaves a log line with a timestamp, a name and ideally a second approver — the four-eyes principle. Where override records exist, ask what they show in total: a model overridden three hundred times a quarter has quietly revealed that the deployment is mostly human improvisation wearing a machine’s badge. A model with zero overrides in five years of live markets has revealed something equally uncomfortable — that no one is actually at the desk.
The four-eyes principle earns its keep on changes, not just on halts. Retraining decisions, threshold adjustments and data-feed substitutions are the changes most likely to import silent regime shifts, which is why change approvals need a second signature from someone who loses sleep over a different aspect of the business than the changer.
Link four: scheduled challenge
Oversight is not only reactive. Statistical supervision and challenger arrangements exist precisely so that challenge is institutional rather than a personal mood: alternatives that might do the same job with less fragility, regular reviews of whether the model still earns its risk tier, and standing invitations to a red team to try to break the deployment before the markets volunteer to do it.
The vocabulary matters less than the calendar. Challenge that happens on a date is challenge that survives busy months; challenge that happens “when warranted” is challenge that happens never. Ask for the cadence of the committee that owns the model, the proportion of meetings with the model on the agenda, and the last occasion on which the challenge broke something — a healthy oversight system should be able to name its own repairs.
Link five: accountability that lands
The final link closes the loop to a person who cannot disappear into the organisational chart: a named senior owner who signs off the model’s status for the board, carries the consequences when oversight fails, and — the part that separates accountability from scapegoating — has the standing to refuse deployment in the first place. When regulators criticise firms after model-driven losses, the recurring finding is not that no documents existed but that no one person carried the duty. Diffusion, not silence, is the standard governance failure.
Case study: forty-five minutes in 2012
The standard cautionary tale of automation without fast human control is Knight Capital, a market-making firm, on 1 August 2012. A deployment error reactivated obsolete code on a newly connected set of stocks. The result: within about forty-five minutes the firm had accidentally built a loss of roughly US$440 million — enough to bankrupt a business that had been profitable the day before.
Two details matter from this desk’s viewpoint. The error was discovered during the incident — the humans were on the case within minutes. The loss still grew to fatal size because every surrounding step (understanding the position, telling counterparties, halting the flow) moved at meeting speed while the code moved at market speed. The second detail is what came after: across the industry, serious firms began rehearsing the unlucky day in advance — practised shutdowns, assigned halt authority, drills for the corner case. Oversight, it turned out, is a thing you rehearse, not a paragraph you publish.
Oversight theatre: the warning signs
Set against that day, these familiar phrases read as theatre rather than control:
- “Experienced team” naming no roles, no structure.
- “Monitored around the clock” producing no logs anyone outside the firm may examine.
- A risk committee that is never described as having blocked, slowed, or reordered anything.
- No incident ever having occurred — which almost always means none has ever been recorded, not none has happened.
- “Human-in-the-loop” used as a paragraph heading with no description of which humans, doing what, within what time.
The five-link checklist
For any firm running a model near your money, ask:
- Who can stop the model today, in one sentence?
- What is the longest time between an alert and a required human response?
- Where are override decisions recorded, and what does their frequency suggest?
- When is the model formally re-challenged on the calendar?
- Which named person’s signature carries the model for the board?
A firm that answers all five concretely is showing you its oversight chain. A firm that answers with mood — cutting-edge, battle-tested, robust — is showing you an unlit dial, and the closing article of this reading list compresses the whole tour into five questions you can ask in a single meeting.