Skip to content
Request Demo

Model Risk Scoring Framework Whitepaper

Resource • Advanced • 4 min read

Audience
CFOs • Investment Committees • Advisory Firms
Last Reviewed
July 2026
Updated
Version 1.0

Executive Summary

A model risk score is only useful if it means the same thing every time it is produced, comparable across models, stable across auditors, and traceable back to specific findings. This whitepaper sets out a methodology for scoring model risk systematically, built on two independent dimensions, the severity of structural findings and the materiality of the model to the decision it supports, and explains why collapsing those two dimensions into a single undifferentiated score produces a less useful, less defensible result.

Key Takeaways

  • A defensible model risk score is built from two independent dimensions, finding severity and model materiality, scored separately before being combined.
  • Severity should be assessed against the model's own structural characteristics, independent of who relies on the model or for what.
  • Materiality should be assessed against the decision the model supports, independent of how well or badly the model itself is built.
  • Collapsing severity and materiality into a single undifferentiated score obscures which lever, fixing the model or scoping the decision differently, actually reduces risk.
  • A risk scoring framework is only useful if applied consistently across models and over time, not re-derived ad hoc for each engagement.

Overview

The Model Risk pillar page introduces the concept of a model risk score as a quantified, structural assessment used to prioritise which models warrant independent audit or closer scrutiny. This whitepaper develops the methodology behind that score in more depth: what it should measure, how its components should be combined, and why the specific design choice of scoring severity and materiality separately produces a more useful and more defensible result than a single blended number.

Framework / Methodology

The framework rests on two independent dimensions.

Severity measures how structurally reliable the model itself is, assessed against a defined set of structural risk categories: formula and logic consistency, hardcoded values overriding live formulas, circular references and whether they resolve stably, hidden or broken content, and documentation quality sufficient to support independent review. Severity is a property of the model, not of who is using it or for what.

Materiality measures how much a real decision actually depends on the model's output. A model directly determining a large capital allocation is highly material. A model used for internal reference only, with no direct decision attached, is low materiality, regardless of how well or poorly it is constructed. See Model Materiality.

The combined score is a function of both dimensions, not a simple average. A structurally severe finding in a low materiality model represents contained risk. The same finding in a highly material model represents concentrated risk that should be prioritised well above it. A scoring methodology that only sums or averages the two dimensions loses this distinction; a methodology that treats them as a matrix, severity against materiality, preserves it and produces a more actionable output.

Key Findings

Scoring severity and materiality independently, rather than as a single blended figure, surfaces a distinction that matters operationally: two models can carry an identical overall risk score for entirely different reasons, one because it is structurally weak but low stakes, the other because it is structurally sound but extremely high stakes. Treating those two cases identically, which a blended score does by construction, leads to the wrong remediation response in both directions: over-investing in fixing a low stakes model's structural issues, and under-investing in independently verifying a high stakes model that happens to look clean.

A second finding is that materiality tends to be the more volatile of the two dimensions over the life of a model. A model's structural severity, once assessed, changes only when the model itself is modified. Its materiality can change abruptly, when a deal is upsized, when the model is repurposed for a new decision, when it moves from an internal planning tool to a lender facing document, which argues for materiality being reassessed more frequently than severity in an ongoing governance framework.

Practical Implications

For organisations building a governance framework, this suggests tiering criteria, described on the Financial Model Governance pillar page, should be driven primarily by the materiality dimension, since tiering is fundamentally a resource allocation decision based on stakes, while the audit methodology applied within a given tier should be driven by comprehensively assessing severity.

For investment committees and boards, a risk score presented without its severity and materiality components separated out should be treated as incomplete. A committee should be able to ask which dimension is driving a given score, and receive a specific answer, not just a single number.

For advisory firms and audit engine providers, this framework implies that a risk scoring output should always expose both components, not just the blended figure, so downstream consumers of the score, committees, boards, lenders, can apply their own judgement about which dimension matters most for their specific decision.

References & Further Reading

The following sources have been verified against their primary publisher and are listed in full, with links, in the References section below. - Federal Reserve, OCC & FDIC — Supervisory Guidance on Model Risk Management (2026, supersedes SR 11-7) - GARP — Model Risk Management (GARP Risk Institute)

Continue Reading

How OXXON tests thisRun a free structural check with FMAE

Frequently Asked Questions

What is a model risk score?

A quantified assessment of how much structural and materiality driven risk a specific model carries, intended to be comparable across models and stable across the auditors or reviewers producing it.

Why should severity and materiality be scored separately?

Because they answer different questions. Severity asks how structurally reliable the model is. Materiality asks how much the decision depends on the model's output. A model can be low severity but high materiality, or the reverse, and collapsing the two obscures which one actually drives the organisation's exposure.

How is severity typically assessed?

Against the model's structural characteristics, formula consistency, circularity, hardcoded values, documentation quality, independent of who is relying on the model or for what purpose.

How is materiality typically assessed?

Against the decision the model supports, how much capital, and how directly, is determined by the model's output, independent of how well or badly the model is actually built.

What is the practical benefit of combining the two dimensions rather than scoring severity alone?

It allows limited audit and governance resource to be directed at the models that carry the most real world exposure, rather than only the models with the most visible structural issues.

Can a well-built model still carry a high overall risk score?

Yes, if it is highly material, a large, direct capital decision, even a structurally clean model carries meaningful residual risk simply because so much depends on it.

Does a high risk score mean the underlying investment is bad?

No. It reflects the model's structural reliability and the decision's dependence on it, not the merits of the underlying opportunity, a distinction developed further on the Model Risk pillar page.

How does this framework relate to model tiering?

Model tiering, used in governance policy to decide review depth, is typically driven primarily by the materiality dimension of this framework, since tiering is fundamentally a resource allocation decision based on stakes.

Can this scoring methodology be applied consistently by a deterministic audit engine?

Yes. A deterministic engine's rule based findings map naturally onto a severity score; materiality is typically supplied as a structured input describing the decision the model supports.

Related Articles

Request Demo