Risk: From Market Risk to Expected Credit Loss, Part 8 — the final post in this series. Previously: IFRS 9 Expected Credit Loss for Vietnamese Banks.
Model validation is the discipline that decides whether a model is allowed to be used. A model can be statistically sound and still fail validation, because validation asks a question the modeller never had to answer: can this withstand independent challenge from someone whose job is to find what is wrong with it?

What validation actually covers
Newcomers assume validation means re-running the statistics. That is one component of three, and usually not the one that generates findings.
| Pillar | Question | Typical finding |
|---|---|---|
| Conceptual soundness | Is the approach right for this purpose? | TTC model used for IFRS 9 provisioning |
| Ongoing monitoring | Is it still working? | No defined performance thresholds |
| Outcomes analysis | Do predictions match reality? | Backtesting never performed |
Most severe findings come from the first pillar. A technically competent model built for the wrong purpose is a worse problem than a crude model used correctly, because the error is invisible in the output.
The tests that get run
Discriminatory power asks whether the model ranks borrowers correctly — do higher-risk grades actually default more? Measured by the Gini coefficient or AUC. A Gini below roughly 0.4 on a corporate rating model invites questions; the threshold varies by portfolio and should be set in advance rather than after seeing the result.
Calibration asks whether predicted levels match observed outcomes. A model can rank perfectly while being systematically too optimistic. The binomial test and Hosmer-Lemeshow test are standard, with the caveat from Part 6 that thin data makes most calibration tests underpowered.
Stability asks whether the population has shifted since development. The population stability index is the usual measure; a reading above 0.25 generally indicates a shift material enough to warrant redevelopment.
Benchmarking compares output against an alternative model or external source. Often the only meaningful test available when internal data is too sparse for statistical power.
What regulators actually ask
Supervisory examination has a recognisable shape, and the questions are rarely the ones a modeller prepares for.
- Who owns this model? A named individual accountable for performance, not a department.
- What data was it built on, and can you reproduce it? Lineage from source system to model input. This question fails more reviews than any statistical test.
- Why these variables? Economic rationale, not just statistical significance. A variable that works without an explanation is a finding.
- What are the limitations? A model owner who cannot list them has not understood the model.
- What happens when it breaks? Defined performance thresholds and a pre-agreed action when they are crossed.
- Who overrides it, and how often? Override rates above roughly 10% suggest the model is not being used, whatever the documentation says.
Questions 2 and 5 generate the most findings in practice. Data lineage is tedious to document and nobody enjoys doing it, so it is frequently incomplete. And performance thresholds are often absent entirely — the model is monitored, but nothing specific has been agreed to happen when monitoring shows deterioration.
The documentation standard
The working test: could a competent quantitative analyst who has never seen this model rebuild it from the documentation alone? If not, the documentation is incomplete regardless of its length.
What a validator looks for, in order of how often it is missing:
- Data lineage and any exclusions applied, with reasons
- Rejected alternatives and why they were rejected
- Limitations stated by the owner rather than discovered by the reviewer
- Assumptions listed explicitly, including implicit ones
- Version history with rationale for each change
The second item is undervalued. A document showing three approaches tested and two discarded for stated reasons reads as considered work. One showing a single approach reads as the first thing that worked.
Independence
Validation must be performed by someone who did not build the model and does not report to whoever did. This is a structural requirement, and the common failure is subtle: validation formally independent but practically captured, where the validator lacks the seniority or technical depth to challenge the developer.
The symptom is a validation report with no material findings. Every model has limitations. A clean report usually means the review was not searching, not that the model was flawless.
In smaller institutions genuine independence is hard to staff. The workable answer is external validation for the highest-impact models and documented internal challenge for the rest — not pretending that a colleague two desks away constitutes independence.
Where this series ends
Eight posts from Value at Risk to here, and the thread running through them is narrower than it looks. Every measure covered — VaR, expected shortfall, stress scenarios, PD, LGD, EAD, ECL — is an estimate produced by a model, and every one of them is wrong in some specific way.
What distinguishes competent risk management is not better models. It is knowing precisely how each one is wrong, documenting that honestly, and deciding in advance what to do when the wrongness starts to matter.
That is also the connection back to the foundations series, which ended on the same point from the pricing side: models are structured arguments about uncertainty, not descriptions of truth.
Frequently asked questions
What is model validation?
It is independent review of whether a model is conceptually sound, still performing, and producing outcomes that match reality — covering judgement and process, not only statistics.
What is a good Gini coefficient?
It depends on portfolio and model type. Below roughly 0.4 on a corporate rating model usually invites challenge, but the threshold should be set before seeing results rather than after.
What do regulators ask first?
Ownership and data lineage. The ability to reproduce model inputs from source systems fails more reviews than any statistical test.
Why do high override rates matter?
Because they indicate the model is not actually driving decisions. Rates above roughly 10% suggest users do not trust the output, whatever the documentation claims.
Try it yourself
Run a self-validation before someone else does, in this order:
- Pick a model your team owns and name the accountable individual.
- Trace one model input back to its source system and document the path.
- List the model’s three most material limitations in writing.
- State the performance threshold that would trigger redevelopment.
- Check the override rate over the last twelve months.
Whichever step takes longest is where your next validation finding is going to come from.