Skip to contentNew: Does ChatGPT recommend your brand? Free 60-second AI visibility check →
By The DDH Team · Digital Dashboard Hub

AI Incident Response Playbook 2026: NIST AI RMF, MITRE ATLAS, OWASP LLM Top 10, EU AI Act Article 73, and ISO 42001 Mapped to Real LLM Disasters

Six frameworks, one job: get you through the first 60 minutes when your chatbot lies, leaks, or libels in production. NIST AI RMF gives the lifecycle. MITRE ATLAS gives the threat taxonomy. OWASP LLM Top 10 gives the failure modes. The AI Incident Database gives the cautionary tales. EU AI Act Article 73 sets the regulator clock. ISO 42001 ties it into your management system. Sources cited inline, June 2026.

By DDH Research Team at Digital Dashboard HubUpdated

Air Canada lost a small-claims tribunal in February 2024 because its chatbot invented a bereavement fare policy, then argued — unsuccessfully — that the chatbot was a separate legal entity (https://www.bbc.com/travel/article/20240222-air-canada-chatbot-misinformation-what-travellers-should-know). In April 2025, Cursor's support bot fabricated a 'one device per subscription' policy and forced a founder apology on Hacker News. Google's Gemini image generator was pulled in February 2024 after producing historically inaccurate images. Meta withdrew Galactica three days after launch in 2022. NYC's MyCity small-business chatbot told users to break the law. Microsoft's Tay survived 16 hours in 2016. None of these were technical surprises — every one was foreseeable under the OWASP LLM Top 10 and the NIST AI RMF GenAI profile. The companies that owned the mistake fastest paid the smallest price. Before the rest of this playbook helps you, run your stack through the LLM jailbreak prevention guide so you catch failures before they become incidents.

This playbook is built around six reference frameworks that overlap deliberately. **NIST AI RMF 1.0** plus the **AI 600-1 Generative AI Profile** (https://www.nist.gov/itl/ai-risk-management-framework) is the US lifecycle model: Govern, Map, Measure, Manage. **MITRE ATLAS** (https://atlas.mitre.org/) is the adversarial-tactics knowledge base — ATT&CK for ML systems, covering prompt injection, data poisoning, and model evasion. The **OWASP Top 10 for LLM Applications** (https://owasp.org/www-project-top-10-for-large-language-model-applications/), updated in 2025, is the developer-focused failure-mode catalog. The **AI Incident Database** from the Partnership on AI (https://incidentdatabase.ai/) is the public corpus of 3,000-plus documented AI failures you should be reading every Friday. **EU AI Act Article 73** sets the legally binding incident-reporting window. **ISO/IEC 42001** is the management-system standard that ties it into an auditable program. All references sourced from the framework owners' pages as of June 2026.

The rest of this guide walks the incident life cycle: detection signals to wire up today, the first-60-minute triage call, containment (rollback, disable, geo-fence), comms to regulators, users, and press, root-cause analysis under MITRE ATLAS, post-mortem culture, and the retro learnings that feed back into your NIST AI RMF Govern function. We map the EU AI Act Article 73 timeline (15, 10, 2 days) against the reality of learning about a chatbot lie via Twitter at 11pm on a Friday. Pair this with responsible AI platforms for enterprise and the EU AI Act compliance checklist.

Digital Dashboard Hub

Safety work depends on prompts you can audit. DDH's Saved Prompt Library versions every prompt — diff what changed, branch for tests, export for review. The audit surface starts with knowing what you actually shipped.

Start free 14-day trial — AICHAT30 = 30% off Pro for 3 months.

NIST AI RMF, MITRE ATLAS, OWASP LLM Top 10, AI Incident Database, EU AI Act Art. 73, ISO 42001 — framework overview, June 2026

Feature
NIST AI RMF
MITRE ATLAS
OWASP LLM Top 10
AI Incident Database
EU AI Act Art. 73
ISO 42001
ScopeFull AI lifecycle risk management (Govern/Map/Measure/Manage) plus GenAI Profile 600-1Adversarial tactics, techniques, procedures targeting ML and AI systemsTop 10 most critical security risks for LLM-powered applicationsPublic registry of real-world AI failures and harms with structured taxonomyMandatory serious-incident reporting for high-risk and GPAI systems in the EUAI management system standard — governance, risk, controls, audit
AudienceAI program owners, risk officers, executivesSecurity engineers, red teams, threat modelersApplication developers, AppSec engineers, ML engineersResearchers, journalists, policy, anyone doing post-mortem analysisProviders and deployers of high-risk AI systems and GPAI models in the EUAI program leads, internal auditors, certification bodies
Prescriptive vs descriptiveDescriptive (voluntary framework, outcome-based)Descriptive (threat knowledge base, no controls mandate)Descriptive with recommended mitigations per riskDescriptive (incident corpus, no controls)Prescriptive (legally binding reporting obligations and timelines)Prescriptive (auditable management-system requirements)
Free vs paidFree (https://www.nist.gov/itl/ai-risk-management-framework)Free (https://atlas.mitre.org/)Free (https://owasp.org/www-project-top-10-for-large-language-model-applications/)Free (https://incidentdatabase.ai/)Free regulation; compliance costs varyPaid standard (CHF 173 from ISO); certification fees separate
Reporting cadence requiredNone mandated; voluntary self-assessmentNone mandated; updated by MITRE communityNone mandated; refreshed roughly every 12-18 monthsVoluntary public submission; editorial reviewMandatory: 2 days widespread harm, 10 days serious incident, 15 days serious malfunctionContinuous via management-system reviews; annual audit cycle
Sector-specific guidanceGenAI Profile (AI 600-1) plus sector overlays in developmentTactic mappings include healthcare, finance, defense case studiesGeneral-purpose; OWASP also publishes API and ML Top 10sTagged by sector — finance, healthcare, government, social media etc.Sector tied to Annex III high-risk categories (employment, education, law enforcement)Cross-sector; pairs with ISO 27001 and ISO 27701
Certifiable by third partyNo formal certification (NIST is voluntary)No certificationNo certificationNo certification (database, not standard)EU notified bodies for high-risk conformity assessmentYes — accredited bodies issue ISO 42001 certificates
Year published / current versionAI RMF 1.0 (Jan 2023); GenAI Profile AI 600-1 (Jul 2024)Launched 2021; living taxonomy as of June 2026v1.0 (Aug 2023); v2025 update (Nov 2024) — verify at owasp.orgLaunched 2020 by Partnership on AI; 3,000+ entries as of June 2026AI Act in force Aug 2024; Art. 73 effective Aug 2026 high-risk, Aug 2025 GPAIISO/IEC 42001:2023 (published Dec 2023)
Used byUS federal agencies, Fortune 500 AI programs, FedRAMP-adjacent buyersMicrosoft, Google, IBM, Anthropic security teams, MLSecOps communityMost enterprise AppSec programs touching LLM appsResearchers, journalists, regulators, AI safety teamsMandatory for providers and deployers placing AI on the EU marketEnterprises pursuing AI governance certification; vendors signaling trust
Best fitSetting your overall AI governance baselineThreat modeling and red-team scopingSecuring the specific LLM application surfaceLearning from documented failures; tabletop exercisesOperationalizing legally mandated reporting for EU exposureBuilding an auditable, certifiable AI management system
Cost to adopt (rough)Internal time only; 2-6 months to operationalizeInternal time only; 1-2 weeks into threat modelsInternal time only; 2-4 weeks to map to AppSecInternal time only; ongoing reading cadenceLegal review plus reporting infrastructure; $50K-$500K+$100K-$500K+ for ISO 42001 certification depending on scope
Update frequencyMajor updates every 12-24 monthsLiving taxonomy; updated continuouslyRoughly every 12-18 monthsNew incidents added daily by community submissionRegulation amended via delegated acts; guidance ongoingISO standards typically reviewed every 5 years

Sources as of June 2026 — verify at the framework owners' pages before relying on any specific obligation: https://www.nist.gov/itl/ai-risk-management-framework, https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook, https://atlas.mitre.org/, https://owasp.org/www-project-top-10-for-large-language-model-applications/, https://incidentdatabase.ai/, https://artificialintelligenceact.eu/, https://www.iso.org/standard/81230.html. Regulatory dates and reporting windows change — confirm with counsel before any incident-comms commitment.

What each framework actually does (and the marketing copy you should ignore)

**NIST AI RMF 1.0** is a voluntary framework, not a regulation, and the common misread is treating it like a checklist. The four functions — Govern, Map, Measure, Manage — are outcomes you must show evidence for. The companion **AI 600-1 Generative AI Profile** published July 2024 (https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook) is where the operational meat lives: 12 risks specific to generative systems, including confabulation, harmful bias, data privacy, intellectual property, and value-chain risk. Read the GenAI Profile before the base RMF.

**MITRE ATLAS** (Adversarial Threat Landscape for AI Systems) is the threat knowledge base at https://atlas.mitre.org/. It uses the same tactic-and-technique structure as ATT&CK, so your existing SOC can map it onto familiar workflows. Tactics include Reconnaissance, ML Model Access, Execution, Exfiltration, and Impact, with techniques like prompt injection, data poisoning, and model evasion. The case studies — including Tay and the Tesla model-evasion experiments — are the best starting point for red-team scoping.

**OWASP Top 10 for LLM Applications** is what your AppSec team will actually adopt because it speaks their language. The 2025 update at https://owasp.org/www-project-top-10-for-large-language-model-applications/ promoted system prompt leakage, vector and embedding weaknesses, and unbounded consumption to top-tier risks. LLM01 (Prompt Injection) and LLM02 (Sensitive Information Disclosure) remain the categories you will see in real incident write-ups.

**The AI Incident Database** at https://incidentdatabase.ai/ is run by the Partnership on AI with editorial review by the Responsible AI Collaborative. It is a structured corpus, not a news feed — incidents tagged by harm type, affected parties, technology stack, and contributing factors. Use it for tabletop exercises: pick an incident from your sector, walk your team through how your stack would have responded, and use the gaps to update your runbook.

**EU AI Act Article 73** is the only item on this list that is law. Once high-risk obligations come into force August 2026 (with GPAI obligations live from August 2025), providers and deployers must report serious incidents to national market surveillance authorities. The clock is 15 days from awareness for a serious malfunction, 10 days for harm to a person, and 2 days for widespread infringement or harm to critical infrastructure. The full timeline is at https://artificialintelligenceact.eu/article/73/. Verify with counsel — implementing regulations may tighten these windows.

**ISO/IEC 42001:2023** is the management-system standard published December 2023 (https://www.iso.org/standard/81230.html). If you know ISO 27001, the structure rhymes: context, leadership, planning, support, operation, performance evaluation, improvement. Annex A lists 38 controls covering AI policy, risk treatment, lifecycle, third-party, and continual improvement. It is what you certify against to make a third-party-audited claim that your AI governance is real.


Real incidents and what each framework would have caught

**Air Canada chatbot, February 2024** — A bereavement-fare hallucination cost the airline a tribunal loss and a viral PR cycle (https://www.bbc.com/travel/article/20240222-air-canada-chatbot-misinformation-what-travellers-should-know). Under OWASP LLM09 (Misinformation) and the NIST GenAI confabulation risk, this was category-one foreseeable. The mitigation pattern — retrieval-grounded responses, refusal-to-answer for policy questions, clear human fallback — is exactly the kind of control ISO 42001 Annex A.7.2 would have flagged as missing in an audit.

**Google Gemini image generation, February 2024** — Image generation produced historically inaccurate outputs after over-correction for representational diversity. Google paused the feature within days. Under MITRE ATLAS this maps to model-behavior shifts during fine-tuning; under NIST AI RMF the failure is in Measure — insufficient evaluation across edge cases before promotion. The lesson is not that diverse outputs are bad; it is that the eval suite did not cover historical-figure prompts, and the rollback path worked.

**Microsoft Tay, March 2016 and Meta Galactica, November 2022** — Tay, the grandparent incident, was prompt-injected and data-poisoned by Twitter users in under 16 hours. Microsoft pulled it the same day — the founding case study for MITRE ATLAS adversarial inputs and the cleanest example of fast containment. Galactica, Meta's scientific-paper LLM, survived three days before withdrawal after researchers documented fabricated citations. Under NIST Govern, the open question was who had authority to call the launch; under ISO 42001, whether a documented risk treatment plan existed. A fast pull is not the same as a controlled rollback if the launch itself was ungoverned.

**NYC MyCity chatbot, March 2024** — The city's small-business chatbot was reported by The Markup as giving advice that would have caused users to violate housing and labor law. The city kept the chatbot online with a disclaimer rather than pulling it — a defensible and contested risk decision under the NIST Manage function. Under EU AI Act terms, a comparable system serving EU residents would have been a high-risk public-services use case requiring conformity assessment and incident reporting.

**Cursor.com support bot, April 2025** — A support bot fabricated a 'one device per subscription' policy, users hit Hacker News, and the founder issued a public apology with refund and process update. The post-mortem cited insufficient grounding and missing human-review escalation. Under OWASP LLM05 (Improper Output Handling) and the NIST GenAI confabulation risk, textbook. The recovery — fast acknowledgment, named owner, concrete fix, refund — is why brand damage was contained.


Detection: the signals you should be wiring up today

Detection is the part of the playbook most teams underinvest in, then regret. The bare-minimum detection layer has four channels. First, **output safety filters** on every generation — Llama Guard, OpenAI's moderation endpoint (https://platform.openai.com/docs/guides/moderation), or a hosted equivalent like Lakera catches obvious policy violations before they ship. Second, **structured logging** of every prompt, system message, retrieved document, model output, and user feedback signal, with retention long enough to support 30-day root-cause analysis.

Third, **user feedback signals** — thumbs-down, abuse flag, support ticket. These are your highest-precision low-recall signal, and OWASP LLM09 calls out feedback loops as a control. The mistake most teams make is collecting feedback and not routing it to a real human queue within 24 hours. By the time you read the negative feedback batch on Monday, the Twitter screenshot is already in tech journalism.

Fourth, **external monitoring** — branded keyword search across X, Reddit, Hacker News, LinkedIn, and the AI Incident Database. Mention, Brand24, or a $30/month Google Alerts plus Slack-webhook setup is the floor. The Cursor incident broke on Hacker News before it broke in internal Slack — assume yours will too. Wire alerts into your on-call channel, not a marketing inbox no one reads on weekends.

Above the basics, instrument **evaluation drift detection**. Run a golden set of prompts against production every hour and alert on regression. NIST AI RMF Measure requires this kind of continuous evaluation for GenAI systems (https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook). The drift detector caught the Gemini issue inside Google before it caught it externally — the question was whether the launch gate respected the drift signal, which is a Govern question.

Finally, **adversarial monitoring** — a small portion of eval traffic should be MITRE ATLAS techniques (https://atlas.mitre.org/techniques) on a schedule: prompt-injection probes, data-exfiltration probes, jailbreak suite runs. JailbreakBench (https://jailbreakbench.github.io/) provides a maintained corpus. The point is not to score perfectly; the point is to detect when your score drops because a model update or system-prompt change opened a new attack surface. If any channel is missing, fix it before you tune your playbook. For deeper coverage see LLM jailbreak prevention and AI audit trail requirements 2026.


The first 60 minutes: triage, containment, and the bridge call

The first hour determines whether you are managing a controllable problem or a public-relations event. Open an incident channel and a video bridge. Page the named incident commander, LLM application owner, security on-call, legal and comms on-call, and the executive sponsor. ISO 42001 Annex A.6.2 expects this rotation to be documented and rehearsed; if you have not run a tabletop in six months, your bridge will be slower than it should be.

Triage in the first 10 minutes. Assign severity using a documented rubric — most teams use 1-to-4 where Sev 1 is widespread harm or regulatory exposure, Sev 2 is significant user impact, Sev 3 is contained, Sev 4 is internal. The EU AI Act Article 73 definitions at https://artificialintelligenceact.eu/article/73/ are useful inputs but they are legal thresholds, not operational ones — your operational severity is usually one notch higher because you want to over-respond.

Containment in the next 20 minutes. The decision tree is short: can you **rollback** the model or prompt version that introduced the issue? If yes, do it now and confirm reversion. Can you **disable** the specific feature? If yes, do it. Can you **geo-fence** (turn off in EU, turn off for unauthenticated users)? If yes, do it. The Air Canada lesson is that the chatbot kept answering during the dispute — every additional answer was incremental liability. Microsoft killed Tay the same day for the same reason.

Comms in the next 20 minutes. Three artifacts: an internal status update, a public acknowledgment for users (tweet, status-page entry, or banner), and a holding statement for press. The Cursor incident showed that named acknowledgment with a real human and a concrete next step beats a corporate-comms statement. Do not lie. Do not blame the model. Do not call it 'hallucination' to non-technical audiences — call it 'incorrect information' and explain what you are doing about it.

Document everything in the last 10 minutes. Timestamped log of decisions, who made them, what was disabled, what was preserved for forensics, what was communicated to whom. This is source material for the EU AI Act Article 73 notification (if it applies) and the internal post-mortem. NIST Manage expects this trail; ISO 42001 auditors will ask for it.

If you do nothing else from this playbook, run a one-hour tabletop this quarter against the Cursor scenario. A team that has rehearsed once is materially better than one that has read the runbook three times. Reps help more than frameworks.


Communication: regulators, users, press, and your own team

External communication is where most incidents are won or lost in 2026. Order matters: internal team first (so they hear it from you, not Slack rumors), then affected users, then regulators (per legal timelines), then press (after a holding statement is ready). Galactica demonstrated this in reverse — researchers reported failures publicly while Meta was deciding what to say, and outsiders set the narrative.

**Regulator notifications under EU AI Act Article 73** are the most time-bound. For serious malfunction in a high-risk system: 15 days; for harm to a person's health: 10 days; for widespread infringement or critical-infrastructure harm: 2 days. Notification goes to the national market surveillance authority in the EU member state where the incident occurred. The European AI Office (https://digital-strategy.ec.europa.eu/en/policies/ai-office) coordinates cross-border cases. Verify with EU counsel.

**User communication** should be specific, named, and actionable. The Cursor post-incident note worked because it said: this is what we did wrong, this is what we changed, here is a refund, here is the human you can email. The Air Canada response failed in part because the airline argued the chatbot was a separate entity rather than apologizing — which made the news cycle worse than the original fare dispute. Users tolerate mistakes; they do not tolerate deflection.

**Press communication** requires a single spokesperson, a holding statement within 60 minutes, and a fuller statement within 24 hours. The holding statement should acknowledge the incident in factual terms, name a containment action already taken, and commit to follow-up. The fuller statement should include root-cause framing, named accountability, and a concrete remediation timeline. NIST GenAI Profile risk-treatment language is useful as a starting point but too technical for press — translate it.

**Internal communication and AI Incident Database submission.** Employees will be asked by friends, family, and recruiters; arming them with a clear internal FAQ within 24 hours reduces both bad external commentary and internal anxiety. Tay's internal Microsoft post-mortem was reportedly fast and frank — part of why Microsoft's AI safety practice in 2026 is one of the strongest in the industry. After containment and legal review, consider submitting the case to https://incidentdatabase.ai/. For incidents that became public anyway, a structured account is a public-interest act and a credibility signal. Microsoft, IBM, and Salesforce submit their own — the corpus is more valuable when it is not only journalists reporting failures.


Root-cause analysis under MITRE ATLAS and NIST

Root-cause analysis for LLM incidents is harder than traditional software because the failure mode is often statistical, not deterministic. MITRE ATLAS (https://atlas.mitre.org/) gives you the threat-actor lens: adversarial input (prompt injection, jailbreak), upstream issue (poisoned training data, compromised supply chain), or deployment-time issue (system prompt regression, model swap, retrieval bug)? Starting with ATLAS tactic categories prevents the 'must be a bad prompt' tunnel vision that wastes the first 48 hours.

The NIST GenAI Profile risk taxonomy (AI 600-1) is the complementary lens. The 12 generative-specific risks include confabulation, dangerous recommendations, data privacy, harmful bias, human-AI configuration, information integrity, intellectual property, and value-chain integration. Tagging the incident gives you both the post-mortem vocabulary and the mitigation library to consult.

Use a structured 5-whys or fishbone analysis, extended for LLMs. The traditional 5-whys assumes deterministic causation; for an LLM you need additional questions: was the failure reproducible (run the prompt 100 times), model-version-specific, system-prompt-specific, or context-dependent (which retrieved documents were in scope). Cursor was reproducible; Gemini was prompt-dependent. Different root causes, different fixes.

Distinguish proximate cause from contributing factors. Proximate cause for Air Canada was the chatbot's confabulation about bereavement fares. Contributing factors: no retrieval grounding, no 'I don't know' refusal path for policy questions, no output safety filter, deployment for a use case the bot was not evaluated against. Fixing only the proximate cause leaves you exposed to the next incident with a different proximate cause.

Document the analysis in a format your team will actually read. The two-page incident report with timeline, contributing factors, immediate actions, and follow-ups is the format that gets read; the 30-page post-mortem is the format that gets filed. NIST AI RMF Playbook (https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook) suggests outcome-based reporting; ISO 42001 Annex A.8.4 wants evidence of corrective action.

Loop the analysis back into your evaluation suite. Every confirmed root cause should generate at least one new evaluation prompt that would have caught this incident pre-deployment. Over time the eval suite becomes a regression test for past incidents — the discipline that separates teams whose incidents repeat from teams whose incidents teach. JailbreakBench (https://jailbreakbench.github.io/) is the public example.


Post-mortem culture and closing the loop on governance

The 'blameless post-mortem' tradition from Google SRE and Etsy is the right starting point, but it needs adaptation. Blameless does not mean accountability-free; the incident review focuses on systemic factors rather than individual blame, while leadership separately addresses governance failures with the people responsible. ISO 42001 Annex A.6.1 on roles exists so accountability is clear before the incident, not assigned after it.

Run the post-mortem within 5 to 10 business days of containment. Sooner is too raw; later loses detail. Invite the on-call team, the application owner, security on-call, a legal observer, and an executive sponsor. The sponsor's job is to authorize follow-up actions on the spot, not grade the response. Structure around five questions: what happened, what was the impact, what went well, what went poorly, what will we change. Time-box each.

Follow-up actions need owners and dates, tracked in your normal project tooling. The most common failure mode is producing a thoughtful action list no one is on the hook to ship. ISO 42001 Annex A.8.4 (corrective action) and the NIST Manage function both expect this loop to close in evidence; auditors will ask to see the action items, owners, and completion dates 6 months later.

An incident response that does not change your governance posture is wasted. Update your **AI risk register** with the new failure mode, contributing factors, and residual risk. Update your **launch-gate criteria** to add the evaluation that would have caught this incident. Over 12 months this discipline produces an eval suite materially better than any off-the-shelf benchmark — because it reflects your actual failure surface.

Update your **third-party assessment template**. If the incident involved a vendor model, retrieval system, or guardrail product, the new vendor questionnaire should include the question that would have surfaced this exposure during procurement. OWASP LLM03 (Supply Chain) covers this surface — most enterprises have a 200-question security questionnaire and a 20-question AI questionnaire, and the AI one is where the real exposure is hiding.

Finally, share what you can externally. Vendor trust programs — Microsoft's Responsible AI Transparency Report, Anthropic's Responsible Scaling Policy, Google's AI Principles report — set the bar in 2026. A quarterly summary of what you fixed and how governance changed is a credibility signal that compounds. Companies public about failures are trusted with bigger contracts in 2027.

How to stand up an AI incident response capability in your organization

  1. 1

    Step 1: Name the incident commander and document the bridge

    Before you write a runbook, decide who owns an incident at 11pm on a Friday. Document the on-call rotation for incident commander, LLM application owner, security on-call, legal/comms on-call, and executive sponsor — primary and secondary for each. Publish the bridge URL, Slack channel name, and paging mechanism in a single page every named person has bookmarked. ISO 42001 Annex A.6.2 requires this documentation, and EU AI Act Article 73 timelines presume you can convene the bridge inside an hour. Run a 30-minute dry-run page exercise this month — page everyone, time how long it takes to populate the bridge, fix the slow path.

  2. 2

    Step 2: Wire the detection layer end-to-end

    Inventory your detection signals against the four-channel baseline: output safety filters, structured logging, user feedback routing, external monitoring. For each channel that does not exist or does not route to on-call, file the work. Add the fifth and sixth channels — eval drift detection on a golden set, and an adversarial-monitoring suite from MITRE ATLAS techniques (https://atlas.mitre.org/techniques) plus JailbreakBench. Aim for end-to-end detection inside 5 minutes for high-severity signals; user feedback should reach a human queue within 24 hours. Validate by injecting a known failure into staging — does the alert reach on-call, and do they know what to do?

  3. 3

    Step 3: Tabletop the Cursor scenario this quarter

    Pick a published incident from https://incidentdatabase.ai/ that resembles your stack — Cursor for SaaS support, Air Canada for customer service, NYC MyCity for public-sector, Gemini for media generation. Convene the named incident commander, app owner, security, legal/comms, and executive sponsor for a 90-minute exercise. Walk through detection, triage, containment, comms (regulator clock, user statement, press holding line, internal FAQ), and post-mortem (who runs it, what artifact, what follow-ups). Capture every gap. The gaps become your priority backlog for 30 days. Run a second tabletop in 90 days against a different incident type to validate the fixes.

  4. 4

    Step 4: Map your obligations under EU AI Act Article 73

    If you place AI on the EU market — as provider, deployer, importer, or distributor — get counsel-led clarity on which incidents you must report, to which national market surveillance authority, and within which window (15, 10, or 2 days per https://artificialintelligenceact.eu/article/73/). Build a one-page decision tree your incident commander can run during the bridge call: is this a serious incident, is it widespread, does it involve critical infrastructure or rights, what is the clock. Identify the named recipient in each EU member state where you operate. Pre-draft the notification template with counsel so it can be filed within hours, not days. For non-EU teams, do the equivalent mapping for any sector regulators (SEC, FINRA, HHS, FTC, state AGs).

  5. 5

    Step 5: Adopt ISO 42001 as your management-system frame, even if you do not certify

    Even without plans to pursue ISO 42001 certification, adopt its structure as your internal management-system frame. The Annex A controls map cleanly to NIST AI RMF outcomes and give you an auditable artifact set: AI policy, role definitions, risk register, lifecycle controls, third-party assessment, incident response, corrective action, continual improvement. The standard itself is paid (CHF 173 at https://www.iso.org/standard/81230.html) but control summaries are widely published. Adopting the structure now means when a future customer, investor, or regulator asks for evidence of an AI governance program, you have an organized binder rather than a scramble. Pair with the NIST AI RMF Playbook (https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook) for outcome guidance.

Continue your research on adjacent topics — calculators, rate limits, head-to-head comparisons, and guides.

Frequently Asked Questions

What counts as a 'serious incident' under EU AI Act Article 73?

Per Article 73 at https://artificialintelligenceact.eu/article/73/, a serious incident is any malfunction of a high-risk AI system that directly or indirectly leads to death or serious harm to a person's health, serious disruption of critical infrastructure, infringement of fundamental rights under EU law, or serious harm to property or environment. Windows: 2 days for widespread infringement, 10 days for harm, 15 days for serious malfunction. The clock starts at provider awareness, not user complaint. As of June 2026, high-risk obligations come into force August 2026; GPAI obligations have been in force since August 2025. Verify with EU counsel.

Should I report my incident to the AI Incident Database?

If the incident became public, yes — submitting a structured report to https://incidentdatabase.ai/ is a public-interest act that improves the corpus everyone else learns from. The database accepts researcher-reported and vendor-self-reported entries; contributing your own gives you control of the framing. If the incident did not become public, talk to counsel first — a voluntary public report creates a discoverable record. Microsoft, IBM, and Salesforce submit their own incidents as part of responsible-AI programs. The credibility benefit of doing this consistently outweighs the discomfort of doing it once.

How is an LLM post-mortem different from a normal software post-mortem?

Three meaningful differences. First, the failure is often probabilistic — the same prompt may produce the bad output 5 percent of the time, not 100 percent — so reproducibility analysis needs 50-100 trials, not 1. Second, contributing factors are layered across model version, system prompt, retrieved context, and user input, so the 5-whys must extend to cover each layer. Third, the fix is usually not a code change — it is an eval addition, prompt refinement, guardrail update, or launch-gate change. The traditional SRE template works as a skeleton; add an 'eval added' section and a 'model/prompt versions tested' section. The NIST AI RMF Playbook (https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook) has examples.

What is the difference between NIST AI RMF and ISO 42001?

NIST AI RMF 1.0 (https://www.nist.gov/itl/ai-risk-management-framework) is voluntary and outcome-based — it tells you what to achieve (Govern, Map, Measure, Manage) but not exactly how. ISO/IEC 42001:2023 (https://www.iso.org/standard/81230.html) is a paid, certifiable management-system standard — it tells you what controls to put in place and gives accredited bodies the right to audit and certify against it. The two are complementary: NIST is your strategic frame, ISO is your operational structure. Annex B of ISO 42001 maps controls back to NIST functions. Most mature programs adopt both — NIST for executive communication and risk language, ISO for the management system you show to auditors and procurement teams.

Can I use OpenAI's or Anthropic's incident response capability instead of building my own?

No — but you can lean on it. Both vendors run mature incident response programs and publish Trust pages with shared-responsibility models (https://openai.com/trust/ and https://www.anthropic.com/trust). If the failure is in the underlying model, the vendor will roll forward — but if the failure is in your application's use of the model, system prompt, retrieval data, or downstream user impact, that is your incident to own. The Air Canada chatbot ran on a third-party platform — the airline still lost. Vendor IR covers vendor incidents; your app's incidents are yours. Plan for both surfaces and include vendor escalation paths in your bridge documentation.

What is MITRE ATLAS and how is it different from OWASP LLM Top 10?

**MITRE ATLAS** at https://atlas.mitre.org/ is a threat-actor knowledge base — it describes how adversaries attack ML and AI systems using the same tactic-and-technique structure as ATT&CK. Right tool for red-team scoping, threat modeling, and SOC analyst training. **OWASP LLM Top 10** at https://owasp.org/www-project-top-10-for-large-language-model-applications/ is a developer-focused catalog of critical security risks specific to LLM applications, with recommended mitigations. Right tool for AppSec reviews, secure-coding training, and launch-gate criteria. They overlap (both cover prompt injection) but serve different audiences. Mature programs use ATLAS for offensive thinking and OWASP for defensive engineering.

How fast do I really need to publish a public statement, and how do I run a tabletop?

Cadence: holding statement within 60 minutes, substantive statement within 24 hours, fuller post-incident note within 7 days. The Cursor founder note hit Hacker News inside 24 hours and contained brand damage. The Air Canada statement came after the tribunal ruling and made the cycle worse. Galactica's withdrawal was fast but the explanation was slow — outsiders set the narrative. Speed beats polish in hour one; substance beats speed in day one. For tabletops, pick a published incident from https://incidentdatabase.ai/ that resembles your stack, brief participants on the scenario but not the resolution, convene the bridge for 60-90 minutes, and walk through detection, triage, containment, comms, and post-mortem. Run one per quarter against different incident types.

What should be in our public AI incident response page?

At minimum: a named contact for vulnerability disclosure (a security@ or aisecurity@ alias), your incident-classification rubric in plain English, a commitment to acknowledge reports within X hours, a commitment to update affected users when feasible, your stance on coordinated disclosure, and a link to recent public post-mortems. Microsoft's Responsible AI Transparency Report, Anthropic's Responsible Scaling Policy updates, and OpenAI's safety pages set the current bar. You do not need to publish at that scale on day one. A one-page security.txt-style page beats nothing and signals to researchers, journalists, and regulators that you have thought about this. See the EU AI Act compliance checklist.

You now know how to respond when your LLM makes a public mistake. Now make every prompt your AI tools run actually hit.

AI Prompt Generator builds production-ready system prompts that work across ChatGPT, Claude, Gemini, and every safety, eval, and governance tool in this article — so your incident response, red-team reviews, and post-mortems get sharper data, not generic AI fluff. Stop tweaking prompts by hand and start shipping prompts that drive measurable lift. 14-day free trial, no credit card required.

Browse all prompt tools →