Modern QA2026Game Day Checklist — tiles
Log inJoin
76 / 95 · 05 Performance & Chaos Engineering · Game Days: Incident Response Testing← prev⊞ allnext →☰ Read as one page

12.3Game Day Checklist

Pre-Game (1-2 Weeks Before)

  1. Scenario Design -- Define the failure scenario and success criteria

    • What failure will you simulate?
    • What is the expected system behavior?
    • What is the expected human response?
    • How will you measure success?
  2. Participant Briefing -- Notify all participants

    • On-call engineers for affected services
    • Incident commander (rotating role)
    • Communications lead
    • Optional observers (management, new team members)
  3. Safety Measures -- Define boundaries

    • What blast radius is acceptable?
    • What is the kill switch procedure?
    • What hours will the exercise run?
    • Is customer impact acceptable? If not, use staging.
  4. Tool Readiness -- Verify all incident response tools work

    • PagerDuty / Opsgenie configured and routing correctly
    • Slack/Teams incident channels ready
    • Monitoring dashboards accessible
    • Runbooks up to date and accessible

During the Game

  1. Inject the Failure -- Execute the chaos experiment

    • Use Litmus, Gremlin, or manual intervention
    • Record the exact time of injection
  2. Observe the Response -- Track metrics without intervening

    • Detection Time: How long until someone notices?
    • Acknowledgment Time: How long until the on-call responds?
    • Communication Time: How long until the team is assembled?
    • Diagnosis Time: How long until root cause is identified?
    • Recovery Time: How long until service is restored?
    • Customer Impact: How many users were affected, for how long?
  3. Document Everything -- A dedicated observer takes notes

    • What happened and when (timeline)
    • What went well
    • What went poorly
    • Gaps in tooling, monitoring, or runbooks

Post-Game (Within 48 Hours)

  1. Post-Game Review -- Blameless retrospective
    • Review the timeline with all participants
    • Identify action items (runbook updates, monitoring gaps, tool improvements)
    • Assign owners and deadlines for each action item
    • Share findings with the broader organization