Skip to content
SupportLayer

Incidents

Escalation design is what an incident tests

During an incident nobody has time to invent process. What the organization has already designed is what it gets to use.

By the team · 2 min read

Incidents do not reveal how hard people work. They reveal whether decision rights, communication, and handoffs were designed before the pressure started.

Severity has to mean something operationally

Severity definitions are only useful when each level carries a specific response: who is paged, how fast, who leads, what cadence communication follows, and what authority the lead holds. Levels that describe impact without specifying response create debate at the worst possible moment.

Write severity so the frontline can assign it confidently in under a minute, with examples rather than adjectives.

Decision rights are the scarce resource

Most incident delay is not technical. It is waiting for permission. Decide in advance who can publish a customer statement, waive a policy, issue goodwill, pause a campaign, or extend a deadline, and set the limits within which they need no further approval.

  • One named incident lead per severity level
  • Pre-approved customer messaging for common scenarios
  • Spending and goodwill limits granted in advance
  • A defined path when the named owner is unavailable

Support and engineering need a shared surface

The most common structural failure is a support organization reporting customer impact into a channel engineering is not reading, while engineering resolves an issue nobody tells customers about. A single incident channel, with a required customer impact summary and a required status back to support, removes most of that friction.

Communication cadence is a promise

Customers tolerate problems far better than silence. A committed update interval, held even when the update is that there is no news, reduces contact volume and preserves trust. Unpredictable communication generates its own second wave of demand.

Recovery is part of the design

The incident ends when the customer relationship is restored, not when the system is stable. Decide beforehand what service recovery looks like, who is entitled to it, and how affected customers are identified and contacted, or that work will be improvised under fatigue.

Request a conversation

If this describes your situation, start there.

Tell us what is happening on your side of it and we will tell you whether an assessment is the right first step.