Claude
Skills
Sign in
Back

incident-response-lifecycle

Included with Lifetime
$97 forever

Incident response process management following the NIST 800-61 lifecycle. Covers severity classification, escalation matrices, role assignment, communication management, phased recovery coordination, blameless post-mortem facilitation, and 5-whys root cause analysis. Scoped to the process and coordination layer — for network-level evidence collection and forensic analysis, use incident-response-network instead.

General

What this skill does


# Incident Response Lifecycle

Structured process management for network incidents from detection through
post-incident review. This skill covers the organizational coordination
layer: severity classification, escalation, role assignment, stakeholder
communication, recovery coordination, and root cause analysis. It does not
cover technical evidence collection, device forensics, or containment
execution — use the incident-response-network skill for network-level
evidence gathering and forensic analysis.

The procedure follows the operational lifecycle shape: detect and classify
the incident, triage and escalate to the right people, coordinate the
investigation across teams, manage communications to all audiences, drive
resolution and recovery, then conduct a blameless post-incident review.

See `references/communication-templates.md` for notification templates by
audience and severity level. See `references/rca-framework.md` for
the 5-whys methodology, fishbone diagram guidance, and post-mortem
document structure.

## When to Use

- **Service-affecting incident declared** — a P1 or P2 event requires
  formal incident management with role assignment and communications
- **Escalation decision needed** — determining who to notify at what
  severity level and when to engage vendor support or management
- **Multi-team coordination required** — investigation spans network,
  security, application, and infrastructure teams needing a single
  command structure
- **Customer or regulatory notification required** — incident has
  external communication obligations (SLA breach, data exposure,
  regulatory reporting)
- **Post-incident review facilitation** — scheduling, structuring, and
  running blameless post-mortems with 5-whys root cause analysis
- **Incident metrics reporting** — collecting MTTD, MTTI, MTTR, and
  recurrence data for continuous improvement

## Prerequisites

- **Incident management authority** — the person initiating this process
  must have authorization to declare incidents and assign roles within
  the organization
- **Contact directory** — current on-call rosters, escalation contacts
  for management, vendor TAC numbers, and regulatory notification
  contacts must be accessible
- **Communication channels** — bridge call infrastructure (conference
  line or collaboration tool), status page access, and email
  distribution lists for each stakeholder group must be established
- **Incident tracking system** — a ticketing system to record the
  incident, track actions, and maintain the timeline of events
- **Defined severity criteria** — organizational agreement on what
  constitutes P1 through P4 severity (see Threshold Tables below for
  a reference framework)

## Procedure

Follow these six steps in sequence. Steps 3 and 4 run in parallel once
roles are assigned — investigation coordination and communication
management proceed simultaneously. Each step references templates from
`references/communication-templates.md` and methodology from
`references/rca-framework.md` where applicable.

### Step 1: Detection and Classification

Classify the incident by severity, type, and scope to determine the
appropriate response level.

**Severity assignment** — apply the P1–P4 taxonomy from the Threshold
Tables section. Base severity on the highest-impact criterion met.
When multiple criteria apply at different levels, the highest governs.

**Incident type classification** — categorize as outage (service
unavailable), degradation (reduced capacity), security (unauthorized
access or data exposure), or data loss (corruption or deletion).

**Scope determination** — assess whether the incident affects a single
device, a network segment, an entire site, or multiple sites. Scope
drives staffing, communication breadth, and recovery complexity.

**Initial impact assessment** — estimate affected user count, impacted
services and their business criticality, data at risk, and revenue
impact per hour. Record estimates in the incident ticket.

### Step 2: Triage and Escalation

Assign roles, notify stakeholders, and set response timeline
expectations based on the severity classification from Step 1.

**Role assignment** — every P1 or P2 incident needs four named roles:
Incident Commander (IC) — owns the incident end-to-end and makes
escalation decisions; Technical Lead — coordinates diagnostics and
synthesizes findings; Communications Lead — drafts stakeholder
notifications and manages the status page; Scribe — maintains the
real-time timeline and records bridge call decisions. For P3, IC and
Technical Lead may be combined. P4 uses normal operations workflows.

**Escalation matrix execution** — notify by severity:
P1 — all four roles plus engineering management, VP/director on-call,
vendor TAC if vendor equipment is involved, executive notification
within 30 minutes. P2 — all four roles plus engineering management
within 1 hour. P3 — Technical Lead plus team lead within 4 hours.
P4 — assigned engineer via normal ticket queue.

**Response timeline expectations:**
P1 — bridge in 15 minutes, first update in 30 minutes, then every
30 minutes. P2 — bridge in 30 minutes, first update in 1 hour, then
every 2 hours. P3 — initial assessment in 4 hours, daily updates.
P4 — acknowledgment within 1 business day.

**Vendor engagement criteria** — engage vendor TAC when the incident
involves hardware failure, software defects requiring patches, or when
internal triage has not identified root cause within the severity time
window.

### Step 3: Investigation Coordination

Coordinate the technical investigation across teams and evidence
sources. For network-level evidence collection (device state, routing
tables, interface data, log retrieval), reference the
incident-response-network skill — this step focuses on organizing the
investigation, not executing forensic commands.

**Evidence collection tasking** — assign team members to collect
evidence from relevant domains: network devices (via
incident-response-network procedures), application logs, infrastructure
metrics, and security tooling alerts. Each assignee reports findings
to the Technical Lead.

**Parallel investigation streams** — for complex incidents, run
multiple investigation threads simultaneously. Common parallel
tracks: (1) symptom analysis — what is failing and for whom,
(2) change correlation — what changed recently (deployments, config
modifications, maintenance), (3) external factors — upstream provider
issues, DDoS, DNS resolution failures.

**Hypothesis tracking** — maintain a running list of hypotheses with
current status (investigating, confirmed, ruled out). Each hypothesis
should have an owner and a validation method. Update the list on every
bridge call.

**Timeline of events (ToE)** — the Scribe maintains a running
chronological log of when events occurred, when they were detected,
what actions were taken, and what was discovered. The ToE becomes the
foundation for the post-incident review in Step 6.

**Subject matter expert engagement** — when investigation stalls or
enters an unfamiliar domain, escalate to specialists. Define clear
handoff: what has been tried, what data is available, and what
specific question needs answering.

### Step 4: Communication Management

Manage stakeholder communications throughout the incident. Use the
templates in `references/communication-templates.md` for consistent
messaging across audiences.

**Stakeholder notification by audience** — executive summary (business
impact, estimated resolution, customer exposure — no technical detail),
technical detail (root cause hypothesis, diagnostics, remediation plan
— delivered on bridge call), customer-facing (service impact, workaround
if available, estimated resolution — via status page), regulatory
(formal notification per compliance framework when required). Use
templates from `references/communication-templates.md`.

**Status update cadence** — follow severity-based cadence from Step 2.
Each update includes: current status, p

Related in General