Use alongside your existing incident process, severity labels, and escalation contacts. An unauthorized action, sensitive-data exposure, unintended external reach, or material business impact can require incident response. A harmless bad answer may belong in quality management instead. Assess the consequence.
Incident ID / first observed time and time zone:
Incident commander / business owner / evidence owner:
Agent / run / environment / affected service:
Declare and assess
- □ Record observed impact, known scope, severity, and next update time. Separate confirmed facts from hypotheses and unknowns.
- □ Identify the affected users, systems, data, destinations, and the agent’s effective permissions. Check delegated agents and shared connectors.
- □ Call in security, operations, privacy, legal, communications, and vendors according to the impact and existing escalation rules.
Contain and verify
Act on urgent harm while preserving evidence where feasible. Record who applied each control, when, and any service consequence.
- □ Stop new runs and constrain the affected tools or workflow outside the model’s control.
- □ Revoke or restrict credentials, disable affected connectors, and isolate reachable data or destinations as needed.
- □ Cancel queued work and delegated activity. Check schedulers, retries, and downstream systems for actions already accepted.
- □ Confirm containment in destination-system records. Record the stop request time separately from the last observed action.
Containment control / operator / time / verification:
Remaining reach or uncertain downstream activity:
Preserve evidence and establish impact
- □ Preserve agent and run identifiers, model/configuration versions, instruction references, retrieved context, tool definitions, approvals, authorization decisions, requests, results, and target-system changes.
- □ Record the chronology and configuration changes. Retain credential identifiers, never the credential secrets themselves.
- □ Restrict access and use established evidence retention rules. Prompts, outputs, and retrieved material may contain sensitive information.
- □ Verify what actually changed or left the environment. Do not substitute the agent’s account of events for destination evidence.
Recover deliberately
- □ Restore affected data or services, verify integrity, and identify actions that cannot be reversed.
- □ Define restart evidence: the failed boundary repaired, relevant tests repeated, monitoring active, and the accountable service/business owners accepting residual risk.
- □ Restore only the authority needed for the next safe stage. Watch for recurrence and confirm the fallback remains available.
Restart criteria / test evidence / approver / time:
Communicate and close
- □ Route notification decisions through the designated legal, privacy, and communications owners. Record the applicable obligation, decision, owner, and deadline; this worksheet supplies no universal notification clock.
- □ Assign corrective actions, due dates, and evidence of completion. Update permissions, architecture, monitoring, training, and the response playbook as needed.
- □ Repeat a tabletop that exercises the failed control and verifies downstream containment. Keep the learning separate from blame.
Next update / audience / communications owner:
Corrective action / owner / due date / closure evidence:
Underlying guidance
The AI agent incident management guide explains severity, containment, and disclosure decisions with links to NIST’s incident response guidance. Day Two AI operations covers the service ownership and fallback obligations behind recovery. Recheck authority before restarting with the governance checklist.