Modern QA2026Communicating During an Active Incident — tiles
Log inJoin
27 / 51 · 24 Technical Writing for QA · Incident Communication← prev⊞ allnext →☰ Read as one page

4.2Communicating During an Active Incident

Severity Levels

Before communicating about an incident, everyone needs to agree on its severity. This determines who is notified, how quickly, and through which channels.

Severity Definition Response Time Communication Cadence Notification Scope
Sev-1 (Critical) Complete service outage, data loss, or security breach Immediate (< 5 min) Every 15 minutes All stakeholders, customers, executives
Sev-2 (Major) Significant degradation, core feature broken, large user impact < 15 minutes Every 30 minutes Engineering, product, support, affected customers
Sev-3 (Minor) Partial degradation, non-core feature broken, limited user impact < 1 hour Every 1-2 hours Engineering, product
Sev-4 (Low) Cosmetic issue, minor inconvenience, workaround available Next business day As needed Engineering team

Status Update Template (During Active Incident)

## Incident Update -- [Timestamp UTC]

**Status:** Investigating / Identified / Monitoring / Resolved
**Severity:** Sev-[1/2/3]
**Incident Commander:** [Name]

### Current Situation
One to two sentences: what is happening right now.

### Impact
Who is affected and how. Be specific.

### Actions Being Taken
What the team is doing right now to resolve the issue.

### Next Update
When the next update will be provided.

### What Teams Can Do
Any actions other teams should take (e.g., "Do not deploy",
"Redirect customer inquiries to [link]").

Example:

## Incident Update -- 2026-02-09 14:30 UTC

**Status:** Identified
**Severity:** Sev-1
**Incident Commander:** Alice Chen

### Current Situation
Checkout is returning 500 errors for approximately 40% of users.
The root cause has been identified as a database connection pool
exhaustion caused by a slow query.

### Impact
Users cannot complete purchases. Approximately 3,200 users
affected. Revenue impact estimated at $500/minute.

### Actions Being Taken
- The slow query has been killed manually
- Connection pool is recovering (currently at 60% capacity)
- A database index fix is being deployed
- QA is verifying checkout functionality in staging

### Next Update
14:45 UTC or sooner if status changes.

### What Teams Can Do
- Support: Direct customers to retry in 15 minutes
- Marketing: Pause any active promotions
- Engineering: Do not deploy any changes until all-clear

Communication Channels During Incidents

Channel Purpose Who Uses It
War room (Slack channel or video call) Real-time coordination among responders Engineering, QA, DevOps
Status page Customer-facing updates Incident commander updates
Email to stakeholders Formal notification to executives and partners Incident commander or engineering manager
Support team channel Updates for customer support staff Incident commander or designated liaison
Social media Public acknowledgment of widespread issues Communications team

Communication Rules

  1. Update early and often. A status update that says "we are investigating" is better than silence.
  2. Underpromise and overdeliver. Say "we expect resolution within 2 hours" if you think it will take 1 hour.
  3. Be honest about what you do not know. "We have not yet identified the root cause" is better than speculation.
  4. Use consistent terminology. "Investigating, Identified, Monitoring, Resolved" -- everyone knows where you are.
  5. Separate facts from guesses. "We have confirmed that..." vs. "We believe that..."