76 / 95 · 05 Performance & Chaos Engineering · Game Days: Incident Response Testing← prev⊞ allnext →☰ Read as one page
12.3Game Day Checklist
Pre-Game (1-2 Weeks Before)
Scenario Design -- Define the failure scenario and success criteria
- What failure will you simulate?
- What is the expected system behavior?
- What is the expected human response?
- How will you measure success?
Participant Briefing -- Notify all participants
- On-call engineers for affected services
- Incident commander (rotating role)
- Communications lead
- Optional observers (management, new team members)
Safety Measures -- Define boundaries
- What blast radius is acceptable?
- What is the kill switch procedure?
- What hours will the exercise run?
- Is customer impact acceptable? If not, use staging.
Tool Readiness -- Verify all incident response tools work
- PagerDuty / Opsgenie configured and routing correctly
- Slack/Teams incident channels ready
- Monitoring dashboards accessible
- Runbooks up to date and accessible
During the Game
Inject the Failure -- Execute the chaos experiment
- Use Litmus, Gremlin, or manual intervention
- Record the exact time of injection
Observe the Response -- Track metrics without intervening
- Detection Time: How long until someone notices?
- Acknowledgment Time: How long until the on-call responds?
- Communication Time: How long until the team is assembled?
- Diagnosis Time: How long until root cause is identified?
- Recovery Time: How long until service is restored?
- Customer Impact: How many users were affected, for how long?
Document Everything -- A dedicated observer takes notes
- What happened and when (timeline)
- What went well
- What went poorly
- Gaps in tooling, monitoring, or runbooks
Post-Game (Within 48 Hours)
- Post-Game Review -- Blameless retrospective
- Review the timeline with all participants
- Identify action items (runbook updates, monitoring gaps, tool improvements)
- Assign owners and deadlines for each action item
- Share findings with the broader organization