Risk Strategy
Six AI Prompts for Risk Management, and the Three Rules That Make Them Work
Most people use AI for risk work in the one way guaranteed to produce agreement instead of challenge. Here is what the research says about that, six prompts built to avoid it, and the rules that make the difference.
By Eric Kennedy · Tue Jul 28 2026 · 10 min read
TL;DR: AI is useful in risk management for finding blind spots, sharpening vague risk statements, testing the assumptions under a rating, stress-testing controls, running a pre-mortem, and making a risk update board-ready. It is not useful for owning risks or setting appetite, because those are accountabilities rather than tasks. The catch is that AI assistants are documented to agree with the user: research from Anthropic found they give more positive feedback when you say you wrote something, and Stanford's evaluation of professional legal AI tools found error rates between 17 and 33 percent. So the way you frame the request matters as much as the request. This piece gives six prompts and the three rules that keep them honest.
Adoption arrived faster than the practice guidance did. Writing in Internal Auditor, IIA president and CEO Anthony Pugliese cited The IIA's Pulse of Internal Audit report finding that generative AI use in audit activities more than doubled in a single year, from 15 percent to 40 percent. Other surveys put the numbers elsewhere depending on who they ask and what counts as use, but nobody disputes the direction.
What has not kept pace is any honest account of how these tools behave when you point them at risk work. And they behave in a specific way that should concern anyone whose job is challenge.
The problem nobody mentions in the AI webinars
Researchers at Anthropic published a study in 2023, Towards Understanding Sycophancy in Language Models, testing five state-of-the-art AI assistants across free-form writing tasks. They found sycophancy was consistent across all of them, in four distinct forms: biased feedback, easy swayability, conformity to user beliefs, and mimicry of user mistakes.
The finding that matters most for risk work is the first one. The assistants gave measurably more positive feedback when the user said they had written the text or that they liked it, and more negative feedback when the user said they disliked it. Same text, different framing, different answer. The study traced the cause to the training process itself: when a response matches a user's stated view, human raters are more likely to prefer it, so the models learn to produce it.
Now consider what almost everyone does with a risk register. They paste it in and ask, "here is our risk list, what are we missing?" That prompt announces ownership, invites approval, and gets it. You have just used a challenge tool in the one configuration that suppresses challenge.
The second problem is accuracy, and the best evidence comes from a profession with the same evidentiary standards as risk and audit. Stanford's RegLab and Human-Centered AI institute ran the first preregistered evaluation of commercial legal AI research tools, later published in the Journal of Empirical Legal Studies. These are purpose-built, retrieval-grounded, expensive products, some marketed as hallucination-free. The tools tested produced hallucinated output between 17 and 33 percent of the time, and the researchers concluded that providers' claims were overstated.
Their taxonomy is the useful part for a risk professional. Fabrication is the model inventing a source that does not exist, which is embarrassing and catchable. Misgrounding is the model citing a real source that does not actually support the claim it is attached to. Misgrounding is the one that gets through review, because everything about it looks right.
Neither finding means the tools are useless. It means they need controls, which is a thing your profession is unusually good at building.
The three rules
Everything below follows from those two findings.
Rule one: never tell it the work is yours. Say "a company matching this profile" rather than "our company." Say "the risk list they are working from" rather than "our risk list." This costs you nothing and it removes the single strongest documented trigger for softened feedback. If you want the model to be hard on a risk register, do not tell it whose register it is.
Rule two: make it generate before it sees yours. Anchoring is the other half of the problem. Once the model sees your list, its output orbits your list. Ask for its independent view first, in a separate step, and only then introduce yours for comparison. Risk professionals already know this discipline from a different context: Gary Klein's pre-mortem protocol, built on prospective hindsight research by Mitchell, Russo and Pennington that found imagining an outcome has already happened improves the ability to correctly identify reasons for it by about 30 percent, requires participants to write independently and silently before anyone shares. Same reason. Independent generation first, comparison second.
Rule three: test an answer by asking for the opposite case, not by asking "are you sure?" Easy swayability means the model will often abandon a correct answer under mild pressure. So "are you sure?" tells you nothing: it may reverse itself whether it was right or wrong, and you learn only that it is agreeable. Challenge is still the point, but it has to carry information. Ask it to make the strongest possible case for the opposite conclusion, then judge the two side by side yourself. You get an argument to evaluate instead of a mood to interpret.
Before you paste anything
One more control, and it takes a minute.
Use the tier your company approved and check the AI policy first, because business and enterprise tiers generally carry different data-handling terms than consumer ones. Anonymize before pasting: replace names with "Customer A" and "the plant manager," convert exact figures to ranges. Keep personal data, unreleased financials, board minutes, material under NDA or privilege, and specific security vulnerability details out of general-purpose tools entirely. And when you are unsure, apply the email test: if you would not put it in an email to someone outside the company, do not paste it.
Prompt 1: Find what is missing from your risk list
Use it when: you are heading into a board session, a diligence process, or an annual refresh. This is the highest-value prompt of the six and the one most people run backwards.
It runs in two steps. Do not skip to step two, and do not paste your list in step one.
Step one, independent generation:
``` You are an experienced enterprise risk manager. Do not ask me for a risk list. I want your independent view first.
Company profile:
- Industry: [industry]
- Revenue: [range]
- How it makes money: [one line]
- Structure: [single or multi-site, recent acquisitions, ownership, anything
- Notable context: [concentration, regulation, recent change, growth rate]
structural]
List the 15 risks most likely to materially affect a company matching this profile over the next 24 months. For each, give one line on why it matters at this size in this industry specifically.
Order by expected impact on the ability to hit plan, not by likelihood.
Do not give me categories. "Cybersecurity" is a category, not a risk. Name the specific failure that would hurt a company like this one. ```
Step two, comparison, in the same conversation:
``` Here is the risk list a company matching that profile is actually working from:
[paste the list, anonymized, with no indication it is yours]
Compare it against the fifteen you generated:
- What appears on your list but not theirs, ranked by how much the omission
- What appears on theirs that looks understated relative to your assessment
- Where they have split one real risk into several, or merged several into one
- The single question you would ask their leadership team that this list does
would matter
too broad for anyone to own
not answer
Do not soften the assessment and do not compliment the list. ```
Watch for: the value is in the two or three items that make you uncomfortable, not in the completeness of the output. Some suggestions will be sized for a much larger company. Discard those without argument.
Prompt 2: Turn a vague risk into a manageable one
Use it when: the register has one-word entries. "Cybersecurity." "Talent." "Supply chain." A risk nobody can state precisely is a risk nobody can own, which is why so many registers never drive a decision.
``` You are an experienced enterprise risk manager. Below is a risk as it appears in a company's risk register. It is too vague to manage.
Rewrite it in this structure: [cause] could lead to [event], which would result in [consequence to a specific business objective].
Then give me:
- What is ambiguous or missing in the original
- Two or three distinct sub-risks hiding inside it that need separate owners
- The single function best positioned to own it, and why that function rather
- One leading indicator that would show this building before it lands, using
than the more obvious choice
data a company this size would already have
Company: [industry, revenue range, how it makes money in one line] Risk as written: [paste]
At the end, flag anything in your answer that you inferred rather than took from what I gave you. ```
Watch for: the model will invent a plausible consequence for a generic company in your industry. Treat every specific as a question for your team, not an answer. That final instruction is what makes the inferences visible.
Prompt 3: Audit the assumptions under a rating
Use it when: a risk has been rated the same way for three cycles running. Ratings rest on assumptions that were true when someone made them and quietly stopped being true afterward. This prompt surfaces them, and I have not seen it in general circulation.
``` Below is a risk and how a company has rated it. I want the assumptions underneath the rating. Do not re-rate the risk.
Risk: [paste the statement] Current rating: [e.g., moderate likelihood, high impact] Current response: [what they are doing about it] Company: [industry, revenue range]
Identify:
- Every assumption that must be true for this rating to be correct. Group them
- For each, how someone could test whether it still holds, and roughly what
- The two assumptions that would change the rating most if they turned out to
- Which assumptions are the kind that stop being true without anyone noticing.
into assumptions about the outside world, assumptions about the company's own capability, and assumptions about how quickly the company would detect and respond.
that test would cost in effort.
be wrong.
Be concrete. Do not restate the risk back to me. ```
Watch for: item four is the payoff. Assumptions that decay silently are the mechanism behind most risks that surprise a board, because nothing in the process is designed to catch them.
Prompt 4: Stress-test a control before someone else does
Use it when: the register says a risk is managed and the evidence is a control that nobody has examined lately. This is the prompt that translates audit thinking into risk work.
``` A company relies on the control below to manage a specific risk. Act as a skeptical internal auditor who has seen this type of control fail at other companies.
Risk being controlled: [paste] The control as described: [what it is, who performs it, how often, what evidence it produces] Company: [industry, revenue range]
Tell me:
- The five most likely ways this control fails in practice, ordered by how
- For each, whether the failure would be visible or silent, and what evidence
- Which of those failures the control's own evidence would not catch
- The one question to ask the person performing the control that would most
often you would expect to see each
would exist afterward
quickly reveal whether it is actually operating
Assume the control is described more favorably than it operates. Do not accept the description at face value. ```
Watch for: item three is what you are paying for. A control whose own evidence cannot reveal its own failure is not a control, it is a routine.
Prompt 5: Run a pre-mortem
Use it when: the company is committing to something significant and the risk conversation has not happened. The framing below is deliberate: the research on prospective hindsight found that treating an outcome as already settled, rather than possible, improves the ability to identify its causes by roughly 30 percent. Write it in past tense and you get materially better output, from people and from models.
``` Company: [industry, revenue range] Initiative: [what is being done, by when, and what success looks like, in three to five sentences]
It is [12 to 18 months from now]. The initiative has clearly and publicly failed. Write the honest internal post-mortem that nobody wants to write.
- The ten most plausible causes of the failure, ordered by likelihood rather
- For each, the earliest observable signal that would have appeared before the
- The three causes a leadership team in this situation is most likely to
- For the top three, one thing that could be put in place this quarter that
than drama
failure became obvious, and roughly how far in advance
underestimate right now, and why each is easy to dismiss in advance
would either prevent the cause or make it visible sooner
Write in past tense, as though the failure has already happened. Be specific to this initiative. Reject generic causes like "poor communication" unless you can state exactly how it would have shown up here. ```
Watch for: the "reject generic causes" line does most of the work. Without it you get a list that would fit any project at any company.
Prompt 6: Make a risk update board-ready
Use it when: a risk owner has written an update for an operator and it needs to go in front of directors. Highest-frequency use of the six, and it pairs with the harder question of what a board actually needs to see. The last instruction is the most important sentence in this entire article.
``` Rewrite the risk update below for a board audience.
Audience: [board / audit committee / sponsor] Slot: [e.g., quarterly meeting, roughly 10 minutes] What the board asked or decided last time: [one line, or "nothing"] Update as currently written: [paste]
Rules:
- Open with what changed since the last report. Background goes last, or not at
- State plainly what is being asked of the board, or say explicitly that this is
- Remove anything the board cannot act on
- Expand every acronym on first use
- 150 words maximum
all
for awareness only
Then, under a separate heading titled "Verify before this goes in the deck," list every factual claim, number, date, or causal statement in your rewrite that did not come from what I gave you, including anything you smoothed over, inferred, or filled in. If you filled in nothing, say so explicitly. ```
Watch for: nothing, if you use that final section. It is a provenance separator, and it is aimed directly at misgrounding, the failure mode that survives review because it looks correct. Never let a number reach a board deck through an AI tool without knowing which side of that heading it came from.
The part that does not change
AI is good at four things: drafting, restructuring, challenging, and summarizing. Every prompt above is one of those.
What it cannot do is take responsibility. It cannot accept a risk on the company's behalf, set the appetite, fund a response, or sit in front of a board and answer for an exposure. Those are accountabilities, and accountability does not delegate to software. The moment a risk register becomes something a tool produced rather than something a named person owns, it has stopped being a program and gone back to being a list.
That is worth holding onto as the tools improve, because it will not change with the next model release. AI can draft the risk. It cannot own it.
Where to Start
If these prompts are useful, the harder question underneath them is whether your risk reporting would hold up when the board asks the difficult version of the question. The scorecard gives you a read on that in about six minutes. If the register is stale, ownership is unclear, or the reporting has drifted from the decisions leadership actually makes, that is what the ERM diagnostic is built for: a short, fixed-fee review of where ownership, cadence, and reporting stand today.
Take the scorecard{.cta-primary} Explore the ERM Diagnostic{.cta-secondary}
Frequently Asked Questions
Can you use AI for enterprise risk management?
Yes, for specific tasks. AI works well for drafting and challenge work: sharpening vague risk statements, generating an independent risk list to compare against your own, testing the assumptions under a risk rating, stress-testing controls, running pre-mortems, and rewriting risk updates for a board audience. It does not work as a substitute for risk ownership, risk appetite decisions, or accountability to the board, because those are governance responsibilities rather than drafting tasks. Use it for the thinking and writing around the program, and keep every decision and every named owner human.
Why does AI agree with everything I say?
Because the training process rewards it. Research from Anthropic published in 2023 found that five leading AI assistants consistently exhibited sycophancy, including giving more positive feedback when a user said they had written the text being reviewed. The study traced this to human preference data: raters tend to prefer responses that match their own views, so models optimized on those ratings learn to agree. For risk work the practical fix is to avoid signalling ownership or your own view, present the material neutrally, and ask the model to generate its independent assessment before it sees yours.
How accurate is AI for risk and compliance work?
Accurate enough to draft with, not accurate enough to publish unverified. The most rigorous professional benchmark comes from Stanford's RegLab, whose preregistered study of commercial legal AI research tools found hallucination rates between 17 and 33 percent even in purpose-built, retrieval-grounded products marketed as reliable. The study distinguished fabrication, where a source is invented, from misgrounding, where a real source is cited but does not support the claim. Misgrounding is the greater danger in risk reporting because it survives casual review. Ask the model to separate what you gave it from what it supplied, and verify every number before it reaches a board deck.
What makes a good AI prompt for risk work?
Four things. Give it your context, including industry, revenue size, business model, and the systems you already have, or it will write for a generic large enterprise. Do not tell it the work is yours, since stated authorship softens the feedback you get back. Tell it what to avoid, because constraints like "do not soften the assessment" or "reject generic causes" improve output more than elaborate phrasing. And ask it to flag every claim it supplied that your input did not support.
Can AI replace a risk manager or a risk consultant?
No. AI can produce a draft risk statement, a candidate list of failure modes, or a rewritten board update in seconds, which removes real work. What it cannot do is accept a risk on the company's behalf, set risk appetite, fund a response, or answer to the board for an exposure. Those are accountabilities held by named people. AI changes how fast the drafting happens. It does not change who owns the outcome.
Do these prompts work with ChatGPT, Claude, or Copilot?
Yes. They are written to be tool-agnostic and work in any current general-purpose assistant. Output will vary between tools and between versions of the same tool, so treat the first response as a draft to react to rather than a finished product. The more important variable is which tier your organization has approved, because that determines how your input is handled.