
How to read an investment model's paperwork

When a model starts touching allocation decisions, the paperwork is the only part of it an outsider can actually inspect. The code stays sealed; the marketing slides exaggerate; but a governance file — when it exists — was written for a different audience, usually supervisors or internal risk committees, and that makes it the most honest artefact a firm produces.
This piece is a map of that file. Not a statistics lesson: a reading order, plus the red flags that suggest no one is really driving.
The stack of papers a governed model produces
Investment firms that run models under documented governance maintain a small, recognisable shelf of artefacts. You will rarely see all of them from outside, but a firm that has them can usually say so plainly — and one that has never heard of them has told you something too.
- A model inventory entry. A register line stating what the model is, who owns it, what decisions it influences, and how risky the firm itself judges it to be. Multi-tier systems rank models by materiality; the highest tier gets the heaviest review.
- A model card or fact sheet. A short technical summary: intended purpose, the data it consumes, the population it was built for, known limitations. The format migrated from machine-learning research into finance; versions of it now appear in supervisory expectations worldwide.
- A validation report. The independent review: what was tested, on what data, out-of-sample or not, and who signed. Crucially it is written by someone other than the builders — that separation is the whole point.
- Version and change logs. Every retraining, every parameter change, every deployment, with dates and named approvers.
- A monitoring plan. Which metrics are tracked, at what cadence, against which thresholds, and what happens when a threshold trips.
- Incident and override records. What the humans did when the model misbehaved, and how long that took.
The classic template, and why it spread
The clearest articulation of this discipline remains the US Federal Reserve’s 2011 guidance on model risk management, known everywhere in the industry as SR 11-7. Its logic is simple to state: models are simplifications by definition, therefore every model carries error, therefore every model needs both independent validation and ongoing monitoring, with accountability assigned to named people.
Regulators elsewhere have converged on the same expectation. The EU’s AI Act grades systems by risk tier and attaches explicit obligations to the high tier — logging, documentation, and human-oversight duties among them. Supervisory bodies in other markets, Taiwan included, increasingly describe AI in finance in the same vocabulary of transparency and accountable persons rather than as a free technological instrument.
So when a firm says its model is “validated”, the honest question is: validated against what, by whom, and when?
What a validation report must contain to count
A validation that deserves the name answers five things on paper:
- Independent authorship. Written by a reviewer who does not report to the model’s builders, ideally not even to the same profit line.
- Out-of-sample evidence. The model’s engine was tuned on one stretch of history; the test that matters is on data it never saw. Walk-forward tests, which roll the training window forward step by step, are the more honest variant.
- Stress behaviour. Not just “it works in normal markets” but explicit scenarios: rate shocks, spread blowouts, a liquidity drought, a sudden correlation collapse.
- Benchmark honesty. What the model is measured against, and whether a naive alternative — an index, a rule of thumb — already does most of the work.
- Limitations, stated by the validator. The section that marketing would never write: where the model should not be used, and with what caveats. A validation with no limitations section is a brochure.
What version history reveals
Among the artefacts above, the version log is the most underrated and the hardest to fake. A model that has genuinely been in production carries a fingerprint of small accidents: a data feed that changed format, a retraining that was postponed twice, a threshold that was quietly loosened after one noisy week and equally quietly restored. Read enough version logs and you can tell a live system from a laboratory trophy — the trophy has no scars.
Three patterns in a change log deserve particular attention:
- Clustered retraining. A model re-retrained eleven times in two months is telling you the data shifted and nobody knew why. Stable learning environments produce boring, evenly spaced log entries.
- Changes that precede incidents. The most instructive line in any version history is the one just before the first excursion — what was adjusted, by whom, and whether the adjustment was validated or simply shipped.
- Gaps. A change log that jumps six quiet quarters between material events often means the log is incomplete rather than the model immaculate. Silence is a data point, but rarely the one firms intend it to be.
From the outside, asking for “the date the current version entered production, and the number of versions since the first” is a reasonable, non-confidential question. The answer is a number: factories of governance produce small regular numbers, marketing departments produce metaphors.
The red-flag vocabulary
Certain words, repeated in disclosures, mark paperwork that is thinner than it sounds:
- “Proprietary” used to deflect a governance question. Trade secrets legitimately exist; governance summaries don’t have to reveal the recipe to prove the kitchen exists.
- “Fully automated” with no mention of who can halt it. Automation is an operating choice; the absence of a named stop authority is a governance gap.
- Backtest curves as the only evidence. A backtest is an anecdote about a friendly past. Without out-of-sample or live-tracking records, it is decoration — and survivorship bias decorates it further (the failed variants never get published).
- Undated (or eternal) validation. “Independently validated” with no date, reviewer, or re-test schedule is a photo of a car without an inspection sticker.
- Monitoring described as continuous but unverifiable. “Monitored 24/7” is frequently a dashboard nobody owns. The question “what was the last escalation?” will outperform any number of the word robust.
- Tiering that never tiers. Firms that describe every model as critical to investors often describe none as critical internally — materiality ranking is what makes heavy review affordable, and its absence suggests nobody priced the risk.
A worked example of the pattern: a disclosure that a “cutting-edge machine-learning engine” is “subject to rigorous ongoing validation” answers none of the five validation questions above — no reviewer, no window, no scenarios, no benchmark, no limitations. Every noun is an adjective away from meaning anything. The way to read such text is not to mistrust it but to flatten it: strip the adjectives, and the number of concrete claims it makes is the number you actually learned. Frequently that number is zero, and knowing it without irritation is a skill.
The ten-line checklist
Before trusting a firm’s account of its own AI, ask for — or look for — these ten items:
- The model inventory entry, including its internal risk tier
- Who owns the model, by role if not by name
- Who validated it, and when that validation is due again
- The out-of-sample window used in the test
- Which stress scenarios the model was shown
- The naive benchmark it must beat, and by how much
- The current version number and the change log history
- The drift metrics monitored, with thresholds
- The named authority who can halt it
- The last logged intervention — date and reason
What to do with missing lines: treat each absent item as a follow-up question rather than a verdict. A firm that answers eight of ten concretely and explains the two gaps candidly has shown you a governance process that knows where its own edges are. A firm that answers ten of ten fluently with zero dates is the one to watch — fluency without anchors is what unrehearsed marketing sounds like.
What the paperwork cannot tell you
A complete shelf of documents still says nothing about the culture that reads it. Reports can pile up unread; committees can meet and minute and move on. The residue of every governance review is therefore behavioural: does the firm change what it does when the monitor blinks? That question is answered incident by incident, and it is exactly where the next piece in this series — the oversight chain — picks up.