Enterprise Risk Management
You Cannot Prove ERM ROI With a Loss That Never Happened
The academic literature has spent twenty years trying to measure what enterprise risk management is worth, and it still struggles to cleanly separate the program from the company that built it. Here is what a CFO can evidence instead.
By Eric Kennedy · Tue Sep 01 2026 · 13 min read
TL;DR
You cannot prove the ROI of enterprise risk management by assigning a dollar value to losses that never happened. For a single company looking back at one realized year, the stronger evidence is what can be documented and remeasured: exposure reduced after mitigation, material decisions changed because risk information existed, and cost-of-risk trends over time. That is not ROI in the textbook finance sense. It is a defensible record of value a CFO and a board can inspect.
Academic research points the same direction. Reviewing the field in 2025, the authors of one of its foundational value studies note that adopting ERM is an endogenous decision, so estimates of its effect are biased whenever the same factors that drive adoption also drive performance. The confounder researchers name is managerial quality.
The Question Every ERM Program Eventually Fails
Somewhere between year one and year two, a CFO asks the risk owner what the program has delivered. The answer is almost always some version of the same thing: we identified exposures, we improved controls, and we probably avoided losses we will never be able to point to.
Perhaps the program prevented the loss. Perhaps it did not. For a single company looking backward at one realized year, the alternative outcome is not directly observable, and CFOs recognize a claim that cannot be checked.
The problem gets worse when the number is specific. Claiming a program prevented a $20 million loss requires knowing what would have happened in a world where the program did not exist. That world is not available for inspection. A finance executive who accepts that reasoning about risk spending would not accept it about a capital project.
None of which means causal effects are unknowable in principle. Researchers estimate them with control groups, matched samples, and quasi-experiments. What a single mid-market company cannot do, looking back at one year that actually happened, is run itself twice.
So the honest starting position is that the headline ROI question, framed as a comparison against an imagined alternative history, cannot be settled from inside one company. Recognizing that is not a concession. It is what makes the rest of the measurement possible, because it forces attention onto things that leave a record.
The Research Has the Same Problem
This is the part worth sitting with, because it changes how the question should be asked.
Academic finance has been trying to measure the value of ERM since the early 2000s. In a 2025 review chapter in the Handbook of Insurance, Robert Hoyt and Andre Liebenberg summarize where that literature has landed. They are not neutral observers of it. Their own 2011 study in the Journal of Risk and Insurance became one of the foundational empirical papers on ERM value.
Their summary is that most studies find firms engaging in ERM, and those with higher quality programs, tend to have greater firm value. Then the footnotes explain why that finding is harder to use than it looks.
The adoption problem. Because the decision to adopt ERM, or to invest in a high quality program, is an endogenous decision made by the firm, estimates of ERM's impact can be biased whenever the factors correlated with that decision are also correlated with observed differences in value. In plain terms, companies choose to build ERM. The kind of company that chooses to build it may already be the kind of company that performs well.
Researchers have tried to correct for this with treatment effect models, two-stage least squares, and instrumental variables. Writing in the Journal of Risk and Insurance, Lundqvist and Vilhelmsson note that a treatment effects model cannot solve omitted variable bias when the omitted variable is unobservable, and name the one that matters in corporate finance research: managerial quality.
That is the whole difficulty in a phrase. Well-run companies build risk programs. Well-run companies also perform better. Separating the two requires observing management quality directly, and nobody can.
The measurement problem. The review notes that the absence of common reporting requirements for whether a firm engages in ERM has been a barrier both to identifying which firms use it and to measuring program maturity. Studies disagree partly because they cannot agree on what counts as having ERM. Early work used the hiring of a chief risk officer as the proxy. Later work used text-based indicators, which capture adoption across many firms but say nothing about quality, or S&P ERM ratings, which assess quality but cover few firms, or direct surveys, which are costly and subject to selection bias.
The sample problem. Most of this work measures insurers and other financial institutions, because that is where the data and market characteristics allow it. The review explicitly identifies research on nonfinancial firms as an open opportunity. There is essentially no evidence base for privately held mid-market companies, which is the honest thing to say rather than borrowing an insurance-industry coefficient and hoping.
The chapter closes by calling for further research using refined measures, methods, and samples larger and more representative of the aggregate economy.
Read that as a CFO. Twenty years of panel data, third-party ratings, and econometric correction, conducted mostly in the one industry where the data is best, and the field's own authors are still asking for better measures. If that apparatus cannot cleanly isolate ERM's contribution to value, a risk owner is not going to isolate it in a single company with one slide.
That is not an argument for giving up. It is an argument for measuring something else.
Three Ledgers a CFO Can Actually Check
Everything below shares one property: it can be documented, remeasured, and challenged without invoking a version of the year that did not happen. I think of these as the ERM Value Ledger, three columns that each hold evidenced and repeatable entries.
A note on what "evidenced" means here, because it matters. A quantified exposure is a modeled estimate, not an observed fact. It rests on assumptions about likelihood, impact, and scenario. What makes it usable is not that it is true in some absolute sense but that it can be documented, calculated the same way twice, remeasured, challenged, and replicated by someone else. That is a much lower bar than proof and a much higher bar than an avoided-loss story.
Exposure reduced. Not "a bad thing did not happen," but a stated starting position, a spend, and a stated ending position. Plausible exposure on the top five risks was quantified at one figure at the start of the year. Management funded specific mitigations. Plausible exposure on the same five, measured the same way, is now a different figure.
This only works if exposure was quantified in dollars in the first place, which is a prerequisite rather than a part of this exercise. If your risk reporting still runs on colors, that has to be fixed before any of this is measurable, and it is a separate piece of work covered in why your risk heat map is failing the board.
The entry in the ledger is not a claim about outcomes. It is a claim about a management decision that was made, priced, and executed. Whether the event occurs is a separate matter.
Decisions changed. This is the most valuable output of a risk program and the least often recorded.
Keep a decision log with three columns: the decision, the risk information that informed it, and what was different as a result. Acquisition terms adjusted after diligence surfaced a customer concentration. A capital project resequenced because a dependency was identified. Contract language changed on a supplier agreement. Inventory policy adjusted on a single-sourced component. Coverage restructured after an exposure was quantified.
The test for an entry requires no hypothetical: what changed after the risk information entered the decision? Record the change in terms, timing, funding, sequencing, limits, ownership, or scope. If you cannot point to a documented change, it does not go in the log.
Do not be surprised if the honest count is small. A handful of material decisions that changed the shape of a year is worth more than forty risks that were catalogued.
Cost of risk. The line items where risk work should eventually show up in the financials: insurance premiums and retentions, claims frequency and severity, control failures and their remediation cost, audit findings, unplanned downtime, contract penalties, expedited freight, inventory write-offs, customer concessions.
Two honest constraints. Most of these move for reasons that have nothing to do with ERM, so attributing a change to the program requires the same care being applied to everything else here. And a single period's movement proves very little. This ledger needs a trend across multiple periods before it carries weight.
There is at least one measurable effect of this kind in the literature. Bailey, Collins and Abbott found that insurers and reinsurers with high quality ERM programs had lower audit fees, shorter audit delays, and were less likely to file financial statements late. Insurance industry, and subject to the same endogeneity caveat as everything else in that literature, but it is an observable outcome rather than an avoided catastrophe.
What a Real Example Looks Like
The literature's own practical illustration of ERM value is instructive, because it is not a prevented loss.
Hoyt and Liebenberg cite an interview with the ERM leader at the ACE Group, who described the function raising awareness of roughly $15 billion in reinsurance recoverables and implementing a plan to reduce both the amount and the risk of those recoverables. The result was improved quality of reinsurance counterparties and less reinsurance ceded, which he connected to increased investor trust and a higher market valuation.
Notice the structure. Nobody claims a disaster was averted. The claim is that the risk function surfaced an exposure nobody had aggregated, management made a decision about it, and the balance sheet position changed. That is a decision-log entry and an exposure-reduced entry, in a single story.
It is also a very large insurer, and the specific number is theirs rather than a benchmark. What transfers is the shape of the argument, not the magnitude.
When the Honest Answer Is No
An ERM advisory writing about how to justify ERM spending should say plainly where the answer is that you should not spend it.
When the real problem is one function, not the enterprise. If the exposures that keep leadership awake all sit inside a single process, a targeted fix is cheaper and faster than a program. Buying a program to solve a departmental problem is the most common way this money gets wasted.
When nobody will sustain the cadence. A program without a recurring rhythm and a named owner decays into a document within about two quarters. If there is no one to run it after the build, the spend buys a deliverable rather than a capability, and the deliverable ages badly.
When the company is too simple for it to pay. Below a certain threshold of complexity, few products, one facility, concentrated decision rights, a competent leadership team already holds the risk picture in a weekly meeting. Formalizing it adds process without adding information.
When it is being bought for a document. A program built to satisfy a lender, an insurer, or a diligence checklist will produce a binder that satisfies them. It will not change decisions, and it should not be evaluated as though it might.
A useful sanity check before committing: if the program worked exactly as intended, name the first decision it would change. If nobody in the room can name one, the timing is wrong.
What This Means for the Business Case
The practical consequence is that the business case for a risk program should be written the way a capital request is written, with a baseline, a defined deliverable, and a measurement plan agreed in advance.
Agree with the CFO, before the work starts, on which of the three ledgers will be used, what the opening balance is, and when it will be reviewed. A program that defines its own measurement after the fact will always be able to tell a good story, which is exactly why nobody believes it.
And be honest that the three measures run on different clocks. Decision changes can be documented as they happen. Exposure reduction requires a comparable remeasurement after the mitigation lands, which may be a quarter or may be a year depending on what was funded. Cost-of-risk trends usually need several periods before an attribution claim is credible. A business case assuming all three will show clean improvement inside twelve months is probably overpromising.
Where to Start
Open a decision log this week. One page, three columns: the decision, the risk information that informed it, and the documented change in terms, timing, funding, sequencing, limits, ownership, or scope. Backfill it for the last twelve months from memory and from meeting notes.
That backfill is the fastest diagnostic available. A program that has been running for a year and produces zero honest entries has told you something important, and it did not cost anything to find out.
If the log is thin and the question is whether the program is structured to produce those decisions at all, the ERM Program Diagnostic is a one-to-two-week, fixed-fee review built for mid-market organizations, and the fee is credited toward a larger engagement if you move forward.
Take the ERM Scorecard Explore the ERM Diagnostic
Frequently Asked Questions
Can you calculate the ROI of enterprise risk management?
Not as a conventional return calculation, because the central claim depends on knowing what would have happened without the program, and that comparison is unobservable. Losses that did not occur cannot be evidenced. What can be measured are three things that do not require proving a negative: exposure reduced against a stated baseline, decisions that changed because risk information existed, and cost-of-risk line items such as insurance retentions, claims, downtime, contract penalties, and remediation costs. A business case built on those three is defensible in a way that an avoided-loss figure is not.
Does research show that enterprise risk management increases firm value?
Most studies find an association between ERM engagement and higher firm value, but the causal claim is weaker than the association suggests. In a 2025 review of the literature, Hoyt and Liebenberg note that adopting ERM is an endogenous decision by the firm, so estimates can be biased when the same factors that drive adoption also drive value. Researchers have named managerial quality as the unobserved variable that statistical corrections cannot fully address. The evidence base is also concentrated in insurers and financial institutions, and the review identifies nonfinancial firms as an area needing further research. Treat the literature as directional rather than as a promised return.
How do you measure the effectiveness of a risk management program?
By what it changes rather than what it produces. A risk register, a heat map, and an annual assessment are outputs, not effects. The practical measures are the number of material decisions that were made differently because risk information existed, the movement in quantified exposure across a full cycle against a stated baseline, and cost-of-risk line items over multiple years. A decision log with three columns, the decision, the risk information, and the documented change in terms, timing, funding, sequencing, limits, ownership, or scope, is the single most useful record and takes minutes a month to maintain.
How do you justify the cost of an ERM program to a CFO?
Write it like a capital request. State the baseline before the work starts, define the deliverable, and agree the measurement plan in advance rather than after. Be honest that the measures run on different clocks: decision changes can be documented as they happen, exposure reduction needs a comparable remeasurement once mitigation lands, and cost-of-risk trends need several periods before attribution is credible. A program that defines its own success criteria after the fact will always produce a favorable story, which is precisely why finance discounts it.
When is an ERM program not worth the investment?
When the exposures all sit inside one function, in which case a targeted fix is cheaper. When no one will own the recurring cadence, in which case the spend buys a document rather than a capability. When the company is simple enough that leadership already holds the risk picture in a weekly meeting. And when the program is being bought to satisfy a lender, insurer, or diligence checklist, in which case it will produce a binder and should not be judged on decisions it was never meant to change. A useful test before committing: name the first decision the program would change. If nobody can, the timing is wrong.