Using AI to Build Realistic Breach Scenarios for Board-Level Drills
Using AI to build realistic breach scenarios for board-level drills means letting a generative model draft the attacker narrative, business impact, and decision inflection points of a tabletop exercise — a practice drill of your incident response plan — so directors rehearse the judgment calls they would actually face during a cyber crisis. The value is speed and specificity: instead of a security lead spending significant time stitching together a plausible ransomware or third-party breach storyline from old reports, an AI assistant produces a scenario grounded in your sector, your regulatory perimeter, and your current threat exposure in a single working session. That shifts the bottleneck from authoring the drill to running it, which is where board-level readiness is actually built.
The catch, and the reason this article exists, is that a convincing narrative on a slide is not the same as a drill the board can execute. A realistic scenario has to force real decisions — disclose or wait, pay or refuse, notify the regulator now or later — and it has to be delivered through a channel that still works when the corporate network, email, and Teams are the very things under attack.
How does AI generate realistic breach scenarios for board drills?
AI can help produce realistic breach scenarios for board-level drills by combining threat-intelligence patterns with an organisation's own context, producing narratives that feel plausible to directors rather than generic textbook exercises. The realism comes from tailoring — AI-generated guidance blends known attacker techniques, industry-specific pressures, and the company's actual systems and reporting obligations into a scenario the board would recognise as one that could happen to them next quarter.
What inputs shape a tailored scenario?
A useful generator draws on a defined set of attributes. Each attribute constrains the output so the drill lands where the board's decisions actually matter.
| Attribute | Allowed values / range | Why it matters |
|---|---|---|
| Threat vector | Ransomware, business email compromise, third-party breach, insider, supply-chain compromise | Anchors the scenario to a technique the board has likely read about |
| Industry context | Banking, insurance, healthcare, manufacturing, public sector | Drives realistic regulator, customer, and media dynamics |
| Regulatory trigger | Sector disclosure and incident-reporting obligations such as DORA, NIS2, SOC 2, or ISO 27001 | Forces disclosure-clock decisions the board must own |
| Business impact surface | Customer-facing outage, data exfiltration, financial fraud, safety event | Aligns the drill to the risks on the enterprise risk register |
| Escalation pressure | Media inquiry, regulator call, ransom deadline, customer SLA breach | Tests governance under time pressure, not just technical response |
| Injects cadence | Timed updates paced to keep a short-form board session moving | Keeps directors engaged and reveals decision bottlenecks |
How does the AI assemble the narrative?
The model composes a timeline from these attributes: an initial detection, a plausible dwell period, the pivot that escalates the incident to board attention, and the injects that force decisions — whether to notify a regulator, engage counsel, pay a ransom, or take a customer-facing service offline. Because the scaffolding is structured, the same engine can produce a variant with a different vector or regulator when the board reconvenes, so directors practise judgement rather than memorising one storyline.
Why are traditional board tabletop exercises falling short?
Traditional board tabletop drills built from static, human-written scenarios are falling short because the threat landscape moves faster than any document author can revise a Word file. When directors gather once or twice a year to walk through a breach scenario written months earlier — often recycled from last year's binder — they rehearse a version of reality that no longer resembles what their security team is actually seeing.
What do we mean by "board tabletop" — and what's actually broken?
The phrase gets used two different ways, and disambiguating them matters:
- The governance walk-through. A facilitator narrates a breach story to directors, who discuss disclosure, legal exposure, and public messaging. Useful for oversight literacy, but rarely tests decisions under pressure.
- The executive decision drill. Directors are put in the seat: ransom pay-or-don't-pay, regulator notification timing, customer communications. This is where readiness is actually forged — and where static scripts fail hardest.
Most boards get the first and call it the second.
Where do static, human-written scenarios break down?
- They age on contact. A scenario written months earlier rarely reflects the attacker techniques your security team is seeing this quarter.
- They're generic. Off-the-shelf scripts rarely model your crown-jewel systems, your third-party dependencies, or your regulatory reporting clocks — a DORA incident-notification window or an NIS2 significant-incident threshold, say.
- They're expensive to produce. A tailored scenario can take a lean in-house security team significant time to build — so it gets built once and reused until it's stale.
- They don't branch. A paper narrative marches forward regardless of the board's decisions, which strips out the very thing tabletops are supposed to teach: consequence.
- They live in a binder. When the drill ends, findings get typed into a report nobody opens again until next year's exercise.
The net effect: directors leave the room feeling rehearsed, while the organization's actual capacity to respond has barely moved.
Which AI techniques and data sources power scenario realism?
The AI techniques and data sources behind realistic board-level breach drills combine large language models with curated threat data and adversary-emulation frameworks — a narrow specification of scenario design, not general "AI for security." The goal is a narrative a board will recognise as plausible for your industry, geography, and tech stack, not a generic ransomware story pulled from a template.
Here are the key building blocks and what each contributes:
- Large language models (LLMs)
- Role: Draft the narrative arc, inject-by-inject wording, media questions, and regulator notifications.
- Why it matters: Turns raw indicators into a story executives can follow without a SOC background.
-
Watch-out: LLMs must be grounded in your organisation's context (sector, systems, obligations) or the scenario feels generic.
-
Threat intelligence feeds
- Role: Supply current adversary tradecraft, targeted sectors, and recent breach patterns.
- Sources: Commercial CTI providers, ISAC/ISAO sharing communities, and open-source feeds.
-
Why it matters: Anchors the drill in what is actually happening now rather than in years-old playbooks.
-
Attack simulation and adversary-emulation frameworks
- Role: Provide the technical spine — plausible kill-chain sequences, lateral movement, and dwell-time behaviour.
-
Why it matters: Ensures the "how" of the breach is technically coherent, not hand-waved.
-
Your own environment data
- Role: Names of real business units, critical applications, third-party dependencies, and applicable regulatory triggers such as DORA, NIS2, or NYDFS.
- Why it matters: This is what makes a directors-and-officers-level conversation feel real — the scenario names your payment processor, your claims platform, and the specific reporting clock (a DORA or NIS2 deadline) that applies.
One underappreciated angle: the value is less in the LLM itself and more in the retrieval layer feeding it. A drill built on generic prompts produces a generic tabletop; a drill built on adversary techniques mapped to your crown-jewel systems, seasoned with sector-specific intelligence, produces a boardroom conversation people remember.
How do AI-driven drills compare to consultant-led tabletop exercises?
To fairly compare AI-driven drills with consultant-led tabletop exercises, teams should first agree on the criteria that matter — because the two approaches optimize for very different things.
Which criteria matter before you compare?
Before weighing options, define what "good" looks like for your program. We suggest five criteria, ordered by weight for most in-house security teams:
- Cadence — how often you can realistically run a drill. A plan practiced once a year decays fast; muscle memory needs repetition.
- Realism — whether the scenario reflects your actual stack, regulators, and threat profile, not a generic ransomware storyline.
- Cost per exercise — fully loaded, including facilitator fees, prep hours, and executive time.
- Executability under pressure — whether the drill produces artifacts (roles, decisions, comms trees) you can actually use during a live incident.
- Auditability — whether outputs map cleanly to the tested-plan evidence a DORA, NIS2, SOC 2, or ISO 27001 auditor asks for.
How do the two approaches stack up?
| Criterion | AI-driven drills (platform-based) | Consultant-led tabletops |
|---|---|---|
| Cadence | On-demand; quarterly or monthly is practical | Typically annual or semi-annual |
| Realism | Scenarios tailored to your environment, sector, and prior incidents | High for the day, dependent on the consultant's sector depth |
| Cost per exercise | Low marginal cost after platform subscription | Significant per-engagement professional-services fees |
| Executability | Same platform runs the drill and the real response, out-of-band | Output is usually a report; execution reverts to the paper plan |
| Auditability | Timestamped logs, decisions, and participant actions captured automatically — the evidence DORA, NIS2, SOC 2, and ISO 27001 auditors ask for | Written after-action report, manually assembled |
What's the verdict?
Consultant-led sessions still shine for set-piece board exercises, novel regulatory scenarios, or when you want an outside expert to challenge assumptions. AI-driven drills win on cadence, cost, and — the underappreciated angle — continuity between practice and response, because the same guided workflows you rehearse are the ones you execute when systems are down. Most regulated mid-market teams should treat the two as complementary: consultants once a year, AI-driven drills every quarter.
What risks and governance issues should CISOs plan for?
The risks CISOs must weigh when adopting AI-assisted breach scenario design fall into four governance issues: sensitive-data exposure, model hallucination, reputational blowback, and unclear ownership of the exercise itself. Each risk has a direct action-and-mitigation pairing, and treating them as a governance discipline — not a tooling choice — is what separates a defensible drill from a liability.
What actions should you take, and what should you watch for?
| Do this | But watch out for | Mitigation |
|---|---|---|
| Feed the AI enough context to generate realistic scenarios | Leaking crown-jewel architecture, control gaps, or M&A detail into a public model | Use a scoped, tenant-isolated environment; redact identifiers; prefer platforms that keep prompts and outputs off training pipelines |
| Let AI draft injects and adversary narratives | Hallucinated threat-actor TTPs, fabricated regulatory citations, or invented CVEs presented to the board | Require a human CSIRT (Computer Security Incident Response Team) reviewer to validate every inject against a known framework before the drill |
| Use AI to accelerate tabletop exercise prep | Board members mistaking a synthetic scenario for a real-world benchmark | Label AI-generated content explicitly; anchor scenarios to plausible sector incidents rather than named third parties |
| Run drills out-of-band from primary systems | Introducing a new SaaS dependency that itself becomes an incident surface | Confirm the platform is genuinely out-of-band — not hosted on the same identity provider or network as production |
Why this follows from the AI shift
If AI is helping shape scenarios that inform how executives will decide under pressure, then the outputs are governance artifacts, not just training material. It follows that they need the same review, retention, and audit treatment as the underlying incident response plan — the same tested-plan evidence a SOC 2 or ISO 27001 auditor expects to see. The underappreciated risk is not hallucination; it is complacency — a slick AI-assisted deck that convinces the board readiness is handled when the plan has never actually been executed out-of-band.
Frequently Asked Questions
How is an AI-assisted breach scenario different from a traditional tabletop script?
A traditional tabletop script is a static document a facilitator reads aloud, with pre-written injects and a fixed narrative. An AI-assisted scenario is more dynamic: guidance can shape the initial breach premise and suggest new injects, adversary moves, and stakeholder pressures in response to how the board actually decides during the drill. That makes the exercise feel less like a rehearsal and more like a real incident, where the next problem depends on your last decision.
Can we run a board-level drill without exposing sensitive incident data to a public AI model?
Treat it as a procurement question. When evaluating any platform, ask specifically where prompts and outputs are processed, whether your inputs train external models, and how the environment stays reachable out-of-band — independent of your primary network — during a real incident. Exigence, for example, provides AI-generated tabletop scenarios and guidance inside an out-of-band incident-response platform, so the same platform you rehearse on is the one you would use to respond when primary systems are down.
How often should the board practice a breach scenario?
For regulated mid-market and lower-enterprise organizations, a board-level drill roughly once or twice a year, supplemented by shorter executive-level exercises, is a commonly cited cadence. Firms in heavily regulated sectors such as financial services and insurance — where DORA and NIS2 press for demonstrable testing in 2026 — typically need more frequent evidence of practice. The bigger point: cadence matters less than whether each drill actually stresses new decisions rather than repeating last year's script.
Who should be in the room for an AI-driven board breach drill?
At minimum: board members or the risk committee, the CISO, the CIO or head of IT operations, general counsel, communications, and the BCDR (business continuity and disaster recovery) lead. For scenarios involving regulated data, add compliance and the privacy officer. AI-assisted injects are most useful when they force cross-functional decisions — legal disclosure timing, customer communications, regulator notification — rather than staying purely technical.
What evidence should we capture from the drill for auditors?
Auditors generally want to see three things: the scenario itself (including how it was tailored to your environment), a timeline of decisions and actions the board took, and a post-exercise review with assigned follow-ups. This is the documented, tested-plan evidence that DORA, NIS2, SOC 2, and ISO 27001 assessments look for. A platform-based approach captures it automatically as an execution log; a paper-plan approach usually leaves you reconstructing it from memory and email threads days later.
Does AI replace the human facilitator in a tabletop exercise?
No. AI accelerates the parts humans are slow at — drafting a fresh, plausible scenario, suggesting injects on the fly, tailoring the narrative to your sector — but a skilled facilitator is still essential to read the room, probe assumptions, and push the board on uncomfortable decisions. The right framing is AI as scenario engine, human as director; teams that try to fully automate the exercise tend to produce drills that feel technically rich but strategically shallow.