incident-response-lifecycle
Incident response process management following the NIST 800-61 lifecycle. Covers severity classification, escalation matrices, role assignment, communication management, phased recovery coordination, blameless post-mortem facilitation, and 5-whys root cause analysis. Scoped to the process and coordination layer — for network-level evidence collection and forensic analysis, use incident-response-network instead.
What this skill does
# Incident Response Lifecycle Structured process management for network incidents from detection through post-incident review. This skill covers the organizational coordination layer: severity classification, escalation, role assignment, stakeholder communication, recovery coordination, and root cause analysis. It does not cover technical evidence collection, device forensics, or containment execution — use the incident-response-network skill for network-level evidence gathering and forensic analysis. The procedure follows the operational lifecycle shape: detect and classify the incident, triage and escalate to the right people, coordinate the investigation across teams, manage communications to all audiences, drive resolution and recovery, then conduct a blameless post-incident review. See `references/communication-templates.md` for notification templates by audience and severity level. See `references/rca-framework.md` for the 5-whys methodology, fishbone diagram guidance, and post-mortem document structure. ## When to Use - **Service-affecting incident declared** — a P1 or P2 event requires formal incident management with role assignment and communications - **Escalation decision needed** — determining who to notify at what severity level and when to engage vendor support or management - **Multi-team coordination required** — investigation spans network, security, application, and infrastructure teams needing a single command structure - **Customer or regulatory notification required** — incident has external communication obligations (SLA breach, data exposure, regulatory reporting) - **Post-incident review facilitation** — scheduling, structuring, and running blameless post-mortems with 5-whys root cause analysis - **Incident metrics reporting** — collecting MTTD, MTTI, MTTR, and recurrence data for continuous improvement ## Prerequisites - **Incident management authority** — the person initiating this process must have authorization to declare incidents and assign roles within the organization - **Contact directory** — current on-call rosters, escalation contacts for management, vendor TAC numbers, and regulatory notification contacts must be accessible - **Communication channels** — bridge call infrastructure (conference line or collaboration tool), status page access, and email distribution lists for each stakeholder group must be established - **Incident tracking system** — a ticketing system to record the incident, track actions, and maintain the timeline of events - **Defined severity criteria** — organizational agreement on what constitutes P1 through P4 severity (see Threshold Tables below for a reference framework) ## Procedure Follow these six steps in sequence. Steps 3 and 4 run in parallel once roles are assigned — investigation coordination and communication management proceed simultaneously. Each step references templates from `references/communication-templates.md` and methodology from `references/rca-framework.md` where applicable. ### Step 1: Detection and Classification Classify the incident by severity, type, and scope to determine the appropriate response level. **Severity assignment** — apply the P1–P4 taxonomy from the Threshold Tables section. Base severity on the highest-impact criterion met. When multiple criteria apply at different levels, the highest governs. **Incident type classification** — categorize as outage (service unavailable), degradation (reduced capacity), security (unauthorized access or data exposure), or data loss (corruption or deletion). **Scope determination** — assess whether the incident affects a single device, a network segment, an entire site, or multiple sites. Scope drives staffing, communication breadth, and recovery complexity. **Initial impact assessment** — estimate affected user count, impacted services and their business criticality, data at risk, and revenue impact per hour. Record estimates in the incident ticket. ### Step 2: Triage and Escalation Assign roles, notify stakeholders, and set response timeline expectations based on the severity classification from Step 1. **Role assignment** — every P1 or P2 incident needs four named roles: Incident Commander (IC) — owns the incident end-to-end and makes escalation decisions; Technical Lead — coordinates diagnostics and synthesizes findings; Communications Lead — drafts stakeholder notifications and manages the status page; Scribe — maintains the real-time timeline and records bridge call decisions. For P3, IC and Technical Lead may be combined. P4 uses normal operations workflows. **Escalation matrix execution** — notify by severity: P1 — all four roles plus engineering management, VP/director on-call, vendor TAC if vendor equipment is involved, executive notification within 30 minutes. P2 — all four roles plus engineering management within 1 hour. P3 — Technical Lead plus team lead within 4 hours. P4 — assigned engineer via normal ticket queue. **Response timeline expectations:** P1 — bridge in 15 minutes, first update in 30 minutes, then every 30 minutes. P2 — bridge in 30 minutes, first update in 1 hour, then every 2 hours. P3 — initial assessment in 4 hours, daily updates. P4 — acknowledgment within 1 business day. **Vendor engagement criteria** — engage vendor TAC when the incident involves hardware failure, software defects requiring patches, or when internal triage has not identified root cause within the severity time window. ### Step 3: Investigation Coordination Coordinate the technical investigation across teams and evidence sources. For network-level evidence collection (device state, routing tables, interface data, log retrieval), reference the incident-response-network skill — this step focuses on organizing the investigation, not executing forensic commands. **Evidence collection tasking** — assign team members to collect evidence from relevant domains: network devices (via incident-response-network procedures), application logs, infrastructure metrics, and security tooling alerts. Each assignee reports findings to the Technical Lead. **Parallel investigation streams** — for complex incidents, run multiple investigation threads simultaneously. Common parallel tracks: (1) symptom analysis — what is failing and for whom, (2) change correlation — what changed recently (deployments, config modifications, maintenance), (3) external factors — upstream provider issues, DDoS, DNS resolution failures. **Hypothesis tracking** — maintain a running list of hypotheses with current status (investigating, confirmed, ruled out). Each hypothesis should have an owner and a validation method. Update the list on every bridge call. **Timeline of events (ToE)** — the Scribe maintains a running chronological log of when events occurred, when they were detected, what actions were taken, and what was discovered. The ToE becomes the foundation for the post-incident review in Step 6. **Subject matter expert engagement** — when investigation stalls or enters an unfamiliar domain, escalate to specialists. Define clear handoff: what has been tried, what data is available, and what specific question needs answering. ### Step 4: Communication Management Manage stakeholder communications throughout the incident. Use the templates in `references/communication-templates.md` for consistent messaging across audiences. **Stakeholder notification by audience** — executive summary (business impact, estimated resolution, customer exposure — no technical detail), technical detail (root cause hypothesis, diagnostics, remediation plan — delivered on bridge call), customer-facing (service impact, workaround if available, estimated resolution — via status page), regulatory (formal notification per compliance framework when required). Use templates from `references/communication-templates.md`. **Status update cadence** — follow severity-based cadence from Step 2. Each update includes: current status, p
Related in General
modeling-omnistudio-epc-catalog
IncludedSalesforce Industries CME EPC product-modeling skill for Product2-based catalog creation. Use when creating EPC products, configuring product attributes, building offer bundles with Product Child Items, or reviewing EPC DataPack JSON metadata for product catalog changes. TRIGGER when: user creates or updates Product2 EPC records, AttributeAssignment payloads, AttributeMetadata/AttributeDefaultValues, Offer bundles, or ProductChildItem relationships. DO NOT TRIGGER when: designing OmniScripts/FlexCards/Integration Procedures (use building-omnistudio-omniscript, building-omnistudio-flexcard, or building-omnistudio-integration-procedure), implementing Apex business logic (use generating-apex), or troubleshooting deployment pipelines (use deploying-metadata).
relationship-science-coach
IncludedUse this skill for direct, practical adult relationship coaching: couples conflict, repair, trust, marriage, dating, flirting, attachment patterns, emotional connection, sex, desire differences, eroticism, kink negotiation, affection, love languages, breakups, and long-term passion. Draw on Gottman, EFT and Hold Me Tight, attachment science, modern sex research, Perel, Nagoski, Kerner, Schnarch, Love and Stosny, and flexible love-language tools. Be concrete and low-hedge. Redirect only for imminent danger, abuse, coercive control, minors, non-consent, self-harm, stalking, or medical/legal/psychiatric decisions.
building-sf-integrations
IncludedSalesforce integration architecture and runtime plumbing with 120-point scoring. Use this skill to set up Named Credentials, External Credentials, External Services, REST/SOAP callout patterns, Platform Events, and Change Data Capture. TRIGGER when: user sets up Named Credentials, External Services, REST/SOAP callouts, Platform Events, CDC, or touches .namedCredential-meta.xml files. DO NOT TRIGGER when: Connected App/OAuth config (use configuring-connected-apps), Apex-only logic (use generating-apex), or data import/export (use handling-sf-data).
venue-templates
IncludedAccess comprehensive LaTeX templates, formatting requirements, and submission guidelines for major scientific publication venues (Nature, Science, PLOS, IEEE, ACM), academic conferences (NeurIPS, ICML, CVPR, CHI), research posters, and grant proposals (NSF, NIH, DOE, DARPA). This skill should be used when preparing manuscripts for journal submission, conference papers, research posters, or grant proposals and need venue-specific formatting requirements and templates.
let-fate-decide
IncludedDraws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated. Interprets the spread to guide next steps. Use when the user says 'let fate decide', 'YOLO', 'whatever', 'idk', or other nonchalant phrases, makes Yu-Gi-Oh references, or when you are about to arbitrarily pick between multiple reasonable approaches. Prefer over ask-questions-if-underspecified when the user's tone is casual or playful rather than precision-seeking.
net-ops
IncludedCross-platform network troubleshooting (Windows, macOS, Linux) via local or remote shell. Use for: DNS broken, can't resolve hostnames, nslookup/dig works but apps fail, NRPT, WFP, scutil, /etc/resolver, systemd-resolved, /etc/resolv.conf, NetworkManager, VPN DNS leak residue (ProtonVPN/Mullvad/WireGuard/AnyConnect), AV/firewall blocking DNS or DoH, Tailscale DNS interaction, intermittent connectivity, remote diagnostics over SSH.