Modern QA2026The Real World: A Day Without Observability
Log inJoin
1 / 2 · Book 6 · Why Testing Before Deployment Is Not Enough · drill: interview Q&A⊞ allnext →Get the book →

1.3The Real World: A Day Without Observability

Imagine this scenario. Your team deploys a new version of the checkout service at 2:00 PM on a Tuesday. The deployment passes all CI tests. The staging environment looks fine. The deployment goes out to all production instances simultaneously.

At 2:15 PM, a subtle bug begins manifesting: the new version has a race condition in the payment processing logic that occurs in approximately 1% of transactions. For 99% of users, everything works perfectly. For 1%, their payment is charged but the order is not created.

Without observability:

  • The first customer complaint arrives at 2:45 PM via the support chat
  • Support escalates to engineering at 3:15 PM
  • Engineering begins investigating at 3:30 PM
  • The bug is identified at 4:30 PM
  • A rollback is initiated at 4:45 PM
  • The rollback completes at 5:00 PM

Total impact: 2 hours and 45 minutes of exposure. With 100 transactions per minute and a 1% failure rate, approximately 165 users were affected.

With observability-driven testing:

  • At 2:00 PM, the deployment goes to 5% of traffic (canary)
  • At 2:05 PM, the automated canary analysis detects that the error rate on the canary is 1% vs. 0.01% on the stable version
  • At 2:06 PM, the canary is automatically rolled back
  • At 2:10 PM, the team receives an alert with full context: error rate spike, affected trace IDs, and the specific error message
  • By 2:30 PM, the bug is identified and a fix is in progress

Total impact: 6 minutes of exposure to 5% of traffic. Approximately 0.3 users were affected.

That is the difference observability-driven testing makes. Not "we test better before production" but "we detect and contain problems in production before they affect most users."