The incident severity classification decision record: why the severity tier definitions you chose determine your customer trust erosion rate and your on-call escalation calibration failure mode

The incident severity classification model — a P0 definition that uses "any customer cannot complete any workflow" as the criterion, conflating single-user edge cases with full-service outages and training the escalation chain to respond to all top-tier notifications with calibrated skepticism; a highest tier so narrowly scoped to data loss and security events that a complete service outage affecting 100% of users is classified P1 and treated with a 4-hour notification SLA while enterprise customers with 30-minute downstream obligations watch their own contracts expire; and no tier for degraded performance, so that a product running at 15x its normal response latency is either a P0 or a P2 bug depending on the incident commander's interpretation, with no notification threshold, no status page requirement, and no post-incident review — are severity classification decisions that are almost never made explicitly. They emerge from an incident response posture built during early development when every problem was small and the classification mattered less than the response; from a legal-compliance lens that equates "critical incident" with "data at risk" and treats availability events as operationally bad but not categorically different from feature bugs; and from a product launch environment where the team had never experienced degraded performance as a category distinct from outage, because at low traffic volume the performance floor and the functional floor are essentially the same point. Three failure patterns: the 38-person B2B project management SaaS that averaged 3.2 P0 incidents per week over its first 8 months of operation, produced 78 P0 notifications, conditioned its CTO and engineering manager to treat P0 notifications as noise, and then experienced a genuine full-service database outage where the incident commander's first escalation call was not returned for 31 minutes and the CTO did not engage until 47 minutes after the outage began; the 45-person data infrastructure SaaS that defined P0 as data loss or security breach only, classified a complete query-tier unavailability event affecting its largest enterprise customer as P1 because all data was intact and correctly buffered, and received a contract compliance notice from the enterprise customer 3 weeks later because the 4-hour P1 notification SLA had put the enterprise customer in breach of its own 30-minute downstream notification obligation; and the 52-person developer tools SaaS that had 12 degraded performance events over 18 months, classified every one as P2 because all features were technically functional, issued no status page updates for any of them, and discovered 16 months into operation that 8 of those 12 events had measurable NPS impact — an average of 8 points below baseline on post-event surveys — affecting cohorts who had experienced their first degraded performance event within their first 60 days of using the product.

A 38-person B2B SaaS company built a project management platform for construction contractors — tracking subcontractor schedules, document approvals, compliance certifications, and site inspection workflows. Their product served 140 paying accounts, ranging from one-person contractors to regional firms managing 80 simultaneous projects. Their on-call rotation covered 6 engineers. Their severity classification policy had been written in the first month of the product's life and specified a single top-tier criterion: "P0: any condition where any paying customer cannot complete any workflow in the product." The intent was to be customer-centric — if a customer was stuck, it was a P0. The policy had not been revised since the product launched.

By month 3, the P0 rate was 1.8 per week. By month 8, it had grown to 3.2 per week. Over those 8 months, the team had accumulated 78 P0 incidents. Analyzing the distribution retrospectively: 51 of the 78 P0 incidents were single-customer edge cases where one user at one account could not complete a specific workflow — a PDF larger than 50MB failing to upload for a customer using a non-standard PDF encoder; a date-format parsing failure for a customer whose regional settings used a dot-separated date format the validation code had not handled; a browser extension conflict causing the drag-and-drop interface to stop responding for one user at one account. Each of these satisfied the P0 criterion — a customer could not complete a workflow. Each triggered the P0 response: a Slack notification to the engineering-incidents channel, a customer success escalation to the affected account, and an obligation for an engineering post-incident review within 48 hours. Nineteen of the 78 P0 incidents were subset-of-customers issues: a feature unavailable for a specific browser version, a performance regression affecting customers using the API rather than the web interface, a payment processing failure during a 40-minute maintenance window for the company's payment processor. Eight of the 78 P0 incidents were genuine widespread outages: service unavailability or data pipeline failures affecting the majority of active users.

The behavioral conditioning began after approximately the 20th P0. The engineering manager started opening P0 Slack notifications on a delayed basis — checking the notification within 5 minutes if she happened to be at her desk, but not interrupting meetings or personal time for the notification itself. By the 50th P0, both she and the CTO had learned, through 4 months of direct experience, that a P0 notification had a 65% chance of being a single-customer edge case requiring a focused engineering fix rather than a crisis response. The CTO began treating P0 notifications as a signal to check context before escalating: reading the incident description before calling the on-call engineer, waiting for the incident commander's first status update rather than immediately joining the incident channel. This behavior was rational given the prior distribution. It was also exactly wrong when, in month 9, the product experienced a genuine full-service database connection pool exhaustion. The connection pool had been sized for 200 concurrent connections; a large customer's batch import of 5,200 historical project documents had opened and held 340 connections for 18 minutes, exhausting the pool. Zero percent of active users could load their dashboards. The incident generated a P0 notification identical in channel and format to the previous 78. The on-call engineer created the incident ticket and called the engineering manager. The engineering manager was in a sprint planning meeting. She saw the call, noted it was about a P0, and sent a text asking the on-call engineer to leave a voicemail with the impact summary so she could respond after the meeting in 15 minutes. The on-call engineer left the voicemail. The engineering manager finished the meeting and called back 31 minutes after the initial outage began. The CTO joined the incident channel 47 minutes after the outage began. The actual resolution — scaling the connection pool and terminating the blocked batch job — took 4 minutes once the engineering manager was engaged. The total outage duration was 51 minutes: 47 minutes of escalation delay, 4 minutes of resolution. The cause of the 47-minute delay was not negligence — it was a behavioral response correctly calibrated to 78 previous P0 notifications, 91% of which had not required immediate escalation. Connect this failure to the incident response playbook decision record: the response playbook specifies the actions taken after severity classification, but it cannot compensate for a classification that has been calibrated into noise by inflation; the response procedure assumes that a P0 notification carries a consistent urgency signal, and the severity classification decision must maintain the signal-to-urgency fidelity that makes the response procedure reliable; a classification definition that produces a 10:1 ratio of minor events to genuine crises under the same tier label is not a classification problem the response playbook can solve — it is a definition problem that must be solved at the classification tier level.

A 45-person data infrastructure SaaS built a real-time analytics pipeline platform — ingesting event streams from customer applications, running continuous aggregations, and serving query results to customer dashboards at sub-second latency. Their product served 62 paying accounts, including 11 enterprise accounts generating more than $20k/year each. Their severity classification had been written by the founding CTO with an explicit data-protection philosophy: "P0 is reserved for incidents where customer data is at risk. Complete availability events are operationally bad but fundamentally recoverable; data loss is not." The P0 tier required data loss, confirmed data corruption, or a confirmed security breach. The P1 tier covered "complete service unavailability — the majority of customers cannot access core functionality." The P2 tier covered everything else: partial feature unavailability, performance degradation, isolated customer issues.

Fourteen months after launch, the query tier — the component responsible for executing customer dashboard queries against the processed event data — experienced a complete failure. A memory allocation regression in a newly deployed aggregation engine caused the query coordinator process to OOM-kill under sustained load and fail to restart cleanly, entering a crash-loop with a 90-second period. During the crash-loop, 100% of real-time queries returned connection timeouts. The ingestion tier was unaffected: customer event data was being received and buffered correctly. No data was lost. No data was corrupted. No security event had occurred. The query tier was unavailable; the data was intact. The incident commander classified the event as P1 based on the severity criteria: complete service unavailability to the majority of users, but no data at risk. The P1 response was initiated: the on-call engineer began the diagnosis and escalated to the engineering manager. The P1 SLA specified a 4-hour maximum resolution target and a 4-hour maximum time to notify affected customers via status page update.

The incident commander updated the status page at the 2-hour mark — 2 hours into the incident, within the P1 SLA. Resolution came at 2 hours and 23 minutes. From an internal metrics perspective, the incident was managed correctly within the defined P1 parameters. The problem was the consequence of the 2-hour notification delay for one specific customer: a financial services firm paying $47k/year for the platform, which used the real-time analytics pipeline to power a compliance monitoring dashboard for their institutional clients. Their contract with their institutional clients specified a status notification obligation: in the event of a service disruption affecting the compliance monitoring dashboard, the financial services firm was required to notify its institutional clients within 30 minutes of the disruption onset. At minute 31 of the data infrastructure outage, the financial services firm still had received no notification from the data infrastructure platform. Their own engineers had detected the unavailability through their own monitoring; they knew their platform was down. But they could not notify their institutional clients because they did not know whether the outage was a transient condition that would resolve in seconds or a sustained unavailability event that would last hours — without the upstream provider's status assessment, they had no basis for the notification their contract required. At minute 37, a member of their operations team posted a public message on the data infrastructure platform's status page asking for an assessment of duration and impact scope. The incident commander saw the public post at minute 49 and updated the status page with an estimated resolution time. The financial services firm notified their institutional clients at minute 52 — 22 minutes past their contractual obligation. Their legal team received a contract compliance notice from one institutional client 3 weeks later citing the delayed notification. Connect this failure to the SLA design decision record: the SLA commitments made to enterprise customers depend on the severity classification decision to determine which events trigger which notification timelines; an SLA that commits to sub-1-hour notification for top-tier incidents without a corresponding severity classification that captures complete-service-unavailability in the top tier creates an internal notification SLA that is structurally incompatible with enterprise customer downstream obligations; the SLA design decision and the severity classification decision must be made together, with the enterprise customer's downstream notification timelines as an explicit input to both.

A 52-person developer tools SaaS built a code search and navigation platform — indexing repositories, serving real-time symbol search results, and providing cross-reference navigation across large monorepos. Their product served 230 paying accounts, primarily engineering teams at mid-size technology companies. Their severity classification had three tiers: P0 for complete service unavailability (no user can search or navigate), P1 for critical feature unavailability (search works but cross-reference navigation is broken, or indexing is completely stalled), and P2 for bugs, performance issues, or minor feature degradations that did not meet the P0 or P1 criteria. Performance issues were explicitly defined as P2: "degraded performance means the product is working but slower than expected — this is a P2 and should be addressed within the current sprint or the next sprint, depending on severity."

Over 18 months, the platform experienced 12 degraded performance events. Each event shared a common pattern: P99 response times for search queries rose from the baseline of 180–220 milliseconds to 8,000–15,000 milliseconds, remaining there for 40 minutes to 3 hours before resolving, usually as the underlying cause — a search index rebuild congesting I/O, a query cache eviction storm following a deployment, or a background indexing job consuming the CPU headroom allocated to the query tier — resolved itself or was addressed by the on-call engineer. During each degraded performance event, all features were technically accessible. Search returned results; they just took 8–15 seconds per query. Cross-reference navigation worked; it just took 10–20 seconds per click. Every incident commander who classified these events applied the tier definition literally and correctly: the product was working, therefore it was P2. No status page updates were issued for any of the 12 events. No customer notifications were sent. No post-incident reviews were required by the P2 policy. Each event was triaged as a performance bug and added to the engineering backlog.

At month 18, a recently joined engineering manager ran an analysis correlating incident timing with NPS survey responses. The platform sent automated NPS surveys to active users 7 days after their first use of a specific feature. She matched survey timestamps against the incident log and identified users who had completed an NPS survey within 24 hours of a degraded performance event. The results: users who submitted NPS responses within 24 hours of a degraded performance event rated the product an average of 8 points lower than the baseline NPS for the same cohorts measured outside degraded performance windows (baseline composite NPS: 52; post-degraded-performance composite NPS: 44, a −8 point change). Eight of the 12 events showed statistically significant NPS impact (p < 0.05) when tested with sufficient sample sizes. Three of the 12 events were too small or occurred at low-traffic periods to achieve significance. One event was ambiguous. The distribution of NPS impact was not uniform across user tenure: users in their first 60 days of using the product showed 14-point degradation (baseline 48, post-event 34); users with more than 6 months of tenure showed only 4-point degradation (baseline 57, post-event 53). New users who experienced degraded performance early in their adoption curve were forming their mental model of product reliability during those experiences. The 12 degraded performance events had never triggered a status page update. They had never generated a customer notification. They had never been reviewed in a post-incident analysis. They had been treated as P2 bugs and had accumulated in the engineering backlog at varying priority. From the external customer's perspective, the product had been effectively unusable 12 times over 18 months for periods of 40 minutes to 3 hours each, and the company had communicated nothing about any of them. Connect this failure to the alerting threshold decision record: the alert rule that detected each degraded performance event — a P99 query latency threshold alert — correctly identified the degradation and created an internal incident; the classification gap was not in the detection layer but in the tier assignment layer; the alert fired with the right signal at the right time, and then the tier assignment discarded the customer-impact significance of that signal by placing it in the lowest tier; the alerting threshold decision and the severity classification decision must be aligned: the threshold at which an alert fires must correspond to the tier at which the incident is classified, and the tier must correspond to the response and notification actions the customer-impact severity requires.

Structural properties set by the incident severity classification decision

Three structural properties are determined when an engineering team decides — or fails to explicitly decide — how to classify incident severity: what the impact scope definition determines about the escalation desensitization gap as top-tier inflation trains the escalation chain to respond with calibrated skepticism, what the tier completeness model determines about the performance degradation classification blind spot as slowdowns that feel like outages disappear into the bug backlog, and what the customer notification threshold determines about the enterprise SLA cascade gap as upstream notification timelines mismatch enterprise customers' downstream contractual obligations. None of these properties are typically labeled as decisions at the time the incident severity policy is written. The P0 definition is written to be customer-centric without explicitly specifying the scope boundaries that prevent single-user exceptions from satisfying it. The tier structure is derived from operational severity categories — down, degraded, bug — without explicit consideration of the customer experience impact that maps degraded performance to the same urgency as an outage from the customer's perspective. The notification threshold is set to a high tier to avoid over-communicating, without modeling the enterprise customer's downstream obligations that require notification at a lower tier.

Property 1: The impact scope definition and the escalation desensitization gap. The escalation desensitization gap is the behavioral distance between the urgency a severity tier is designed to signal and the urgency with which the escalation chain responds to that tier's notifications — a gap that opens when the tier is over-triggered by events that do not carry the urgency the tier was designed to represent. The mechanism is statistical: escalation recipients form a Bayesian prior about the urgency distribution of top-tier notifications from their experience of past notifications; if the distribution is skewed toward minor events, the prior reflects that skew and the behavioral response is calibrated accordingly; the calibration is accurate for the prior distribution but produces under-response when a genuine major event occurs, because the event is drawn from a different distribution than the one the prior represents. The scope definition is the structural input that determines whether the top tier is over-triggered: a definition using "any customer affected" as the scope criterion can be satisfied by a single-user exception as readily as by a service-wide outage, collapsing two orders of magnitude of actual urgency into one tier label; a definition using "greater than 25% of active users affected" or "complete core workflow unavailability for a named enterprise account" as the scope criterion requires a materially significant impact before the top tier fires, keeping the prior distribution aligned with the tier's intended urgency signal. The desensitization gap cannot be closed through response procedure changes alone: instructing the escalation chain to respond with maximum urgency to every top-tier notification is not effective if the notification rate creates a frequency that makes maximum-urgency response unsustainable; the gap must be closed at the definition level by ensuring the top tier fires infrequently enough and consistently enough for genuine major incidents that the escalation chain maintains the required urgency calibration. Connect this property to the on-call load management decision record: the on-call alert volume budget constrains the total number of pages the on-call engineer can receive before the rotation becomes unsustainable; the severity classification decision shapes the distribution of those pages across urgency levels; a too-broad P0 definition produces a high-urgency page distribution that exhausts the behavioral response budget and forces the escalation chain to triage internally — treating nominally identical pages as different urgency based on context — which is the same calibration degradation problem that the alert volume budget and actionability classification were designed to prevent in the monitoring context.

Property 2: The tier completeness model and the performance degradation classification blind spot. The performance degradation classification blind spot is a structural gap in the tier model where a materially significant customer impact category — the product working at 10–100x its normal response latency — has no tier that accurately represents its urgency and therefore defaults to the lowest available tier regardless of its actual customer impact. The gap exists because most tier models are derived from the availability binary: either the service is available or it is not. Within the availability binary, tiers distinguish between "partially available" and "completely unavailable," but they do not distinguish between "available at baseline performance" and "available at severely degraded performance." The customer's experience of the product does not follow the availability binary: a product that responds in 15 seconds to every action is not "available" in any meaningful sense from the user's perspective, even though it satisfies the technical definition of availability. The structural consequence of the blind spot is that performance degradation events are permanently misclassified: they cannot satisfy the functional impairment criteria for higher tiers, so they default to the lowest tier; the lowest tier carries response obligations — post-incident review, notification threshold, escalation chain — calibrated for minor bugs, not for events that materially damage user trust and NPS; the mismatch is invisible in the incident metrics (the incidents are classified and counted correctly for the tier they are assigned to) but visible in customer metrics (NPS impact, support ticket rate, retention curves for cohorts who experienced the events). The fix is to specify a "degraded performance" tier with quantitative criteria: P99 latency exceeding a defined multiple of the baseline, error rate exceeding a defined threshold, or throughput below a defined fraction of the rolling average, sustained for more than a defined minimum duration. The criteria must be quantitative and pre-specified so that the incident commander can apply them under incident pressure without ambiguity. The tier must carry its own response obligations: a status page update within 30 minutes, a post-incident review requirement, and a notification threshold set lower than the P1 threshold but higher than the P2 bug threshold. Connect this property to the observability strategy decision record: the observability signals that detect degraded performance — P99 latency, error rate, throughput — are the same signals that must flow into the severity classification tier assignment; a degraded performance tier without the corresponding alert rules that fire when the tier criteria are met will never be consistently applied in practice, because incident commanders under incident pressure classify from the alerts available to them; the observability strategy must include the alert configuration for each tier's quantitative criteria, so that the alert signal and the tier assignment are structurally aligned.

Property 3: The customer notification threshold and the enterprise SLA cascade gap. The enterprise SLA cascade gap is the time between the onset of an incident that exceeds an enterprise customer's downstream notification threshold and the moment the enterprise customer receives the information needed to fulfill their own downstream notification obligation — a gap that exists when the upstream provider's notification SLA is longer than the enterprise customer's downstream contractual notification deadline. The gap is a structural mismatch between two independently specified timelines: the upstream provider sets their notification SLA based on internal response capacity and incident resolution targets; the enterprise customer's downstream notification obligation is set by their contracts with their own customers, independently of the upstream provider's SLA. The gap is invisible at the time both SLAs are established because the enterprise customer's downstream obligations are not a standard input to the upstream provider's tier specification process. In enterprise B2B software, the cascade gap is a persistent contractual risk: enterprise customers typically have shorter downstream notification obligations than the upstream provider's response SLA for equivalent incident tiers, because the enterprise customer's customers are often paying for contractual uptime guarantees that require rapid communication when those guarantees are at risk. The structural fix requires making the cascade gap visible: during enterprise customer onboarding, collect the customer's downstream notification obligation and compare it against the provider's notification SLA for the relevant tiers; where the obligation is shorter than the SLA, the options are to change the notification SLA for that tier (improving response for all customers), to create a per-customer notification tier that triggers earlier for flagged enterprise accounts, or to include explicit contract language specifying that the enterprise customer is responsible for detecting availability events through their own monitoring and the upstream provider's SLA is the commitment for status updates, not for first notification. Connect this property to the SLA design decision record: the SLA design decision specifies the commitments made to customers at each tier of the product offering; the enterprise SLA tier carries the highest commitments and the most scrutiny; the severity classification decision must be an explicit input to the SLA design process — specifically, which tier of the severity classification triggers which SLA commitment — so that the commitments are achievable given the classification's notification thresholds; an enterprise SLA that promises sub-1-hour notification for "any significant service disruption" without a corresponding severity tier definition that captures "significant service disruption" above the notification threshold creates a commitment the classification system cannot fulfill.

The incident severity classification ADR: five sections

Section 1: Impact scope taxonomy and severity tier boundary specification. Begin the incident severity classification decision record by specifying the complete taxonomy of customer impact dimensions that will be used to define tier boundaries: the scope dimension (what fraction of the customer base is affected — one customer, a defined customer segment, the majority of customers, all customers) and the impairment dimension (what the affected customers cannot do — a secondary feature is unavailable, a primary workflow is impaired, the entire service is unreachable, data integrity is violated). Specify the severity tiers using both dimensions simultaneously: P0 requires a broad scope threshold (at least 25% of active users or at least one named enterprise account with tier-one coverage) AND a complete core impairment (no meaningful use of the service is possible, or data integrity is at risk); P1 requires either the scope threshold with partial impairment (the service is significantly degraded or a primary workflow is unavailable for a broad user population) or narrow scope with complete impairment (a single enterprise account cannot use the service); Sev-3 / Degraded Performance requires the scope threshold with a performance impairment that crosses the quantitative degraded-performance criteria (P99 latency exceeding 3x baseline, error rate above 1%, or throughput below 50% of the rolling 7-day average for more than 10 consecutive minutes); P2 covers everything that falls outside the scope threshold and impairment criteria for the above tiers. Specify the scope threshold numerically rather than qualitatively — "majority," "significant portion," and "many customers" are unenforceable under incident pressure; "greater than 25% of users who have been active in the last 7 days" is enforceable. Specify that the tier criteria are evaluated against the impact at the time of classification, not the potential impact at the time the causal condition was first detectable — the tier should reflect the current observed impact, not a precautionary escalation of a condition that might develop into a broader impact. Connect this section to the incident response playbook decision record: the response playbook's procedures are indexed by severity tier; each tier maps to a distinct procedure that specifies the escalation chain, the response time commitments, the communication actions, and the post-incident review requirements; the tier boundary specification in the severity classification decision is the index that the response playbook depends on; a boundary specification that produces inconsistent tier assignments across incident commanders will produce inconsistent response procedures, because the playbook cannot compensate for a classification input it receives incorrectly.

Section 2: Degraded performance tier specification and quantitative threshold definition. Specify the degraded performance tier as a distinct tier between the "partial functional impairment" tier and the "bug/minor issue" tier. The degraded performance tier is triggered when the service is functionally accessible but performance metrics cross quantitative thresholds that indicate material customer experience impairment. Define the thresholds explicitly: P99 response latency for core user-facing actions exceeding 3x the rolling 14-day P99 baseline, or sustained for more than 10 minutes; user-facing error rate (HTTP 4xx/5xx on product endpoints, excluding intentional rate limiting responses) exceeding 1% over a 5-minute rolling window; ingestion or processing throughput below 50% of the rolling 7-day average for more than 10 consecutive minutes. Specify that the degraded performance tier carries all three of the response obligations that distinguish it from a P2 bug: a status page update within 30 minutes of classification (the update must state that degraded performance is occurring, not that the team is investigating an unspecified issue); a notification to all enterprise accounts for whom the impaired surface is a contracted core feature; and a post-incident review within 5 business days with the review required to include a root-cause analysis and a prevention plan even if the event resolved without intervention. The 30-minute status page update obligation is the most operationally important distinction: status page updates during degraded performance events prevent the customer trust erosion that accumulates when customers experience performance impairment without receiving any communication — the experience of "the product was unusably slow and no one acknowledged it" is significantly more damaging to customer trust than "the product was unusably slow and I immediately received a status update explaining the situation." Specify that the degraded performance tier's post-incident review is required to evaluate whether the triggering condition was a previously undetected failure mode — and if so, add an alert for the underlying condition to the monitoring coverage at a threshold that would fire before the performance impact crosses the degraded performance criteria. Connect this section to the alerting threshold decision record: the degraded performance tier criteria must have direct alert coverage — a specific alert rule for each quantitative threshold — so that the tier is triggered by an automatic alert rather than by a customer complaint or a manual metric check; the alerting threshold decision specifies the sensitivity and routing of these alert rules; the severity classification decision specifies the tier the alert maps to; both must be written together, because a tier without alert coverage will not be applied consistently in practice, and an alert without a defined tier assignment will produce inconsistent classification under incident pressure.

Section 3: Customer notification model and enterprise SLA mapping. Specify the customer notification requirements for each severity tier: the notification channel, the maximum time from incident onset to first notification, the minimum information that must be included in the first notification, and the notification cadence during an ongoing incident. For the top tier: status page update within 15 minutes of classification, direct email notification to all customers (or enterprise customers for a product with tiered notification) within 30 minutes, minimum information = "We are experiencing [impact description] affecting [scope]. We are actively investigating and will provide an update within [next update interval]." For the P1 tier: status page update within 30 minutes, email to enterprise accounts within 60 minutes. For the degraded performance tier: status page update within 30 minutes, no required direct email to non-enterprise accounts unless the event lasts more than 60 minutes. Separately, specify the enterprise SLA mapping process: during enterprise customer onboarding and at each contract renewal, collect the enterprise customer's downstream notification obligation (the maximum time from their own service disruption to notification of their downstream customers or internal stakeholders). For each enterprise customer, calculate the cascade gap for each severity tier (the difference between the enterprise customer's downstream obligation and the tier's notification SLA). For customers where the cascade gap is negative (the downstream obligation is shorter than the notification SLA), implement per-customer accelerated notification: the incident management tooling generates an automatic notification to flagged customers when any incident at P1 or above is classified that affects their service area, regardless of whether the global notification threshold has been met. Document the per-customer notification logic in the severity classification decision record and in the enterprise customer's contract as an explicit SLA commitment, so that the obligation is visible to both the customer success team and the incident response team. Connect this section to the SLA design decision record: the customer notification model is a component of the SLA commitment, and the enterprise SLA mapping process closes the cascade gap that exists when the SLA is designed without enterprise customer downstream obligations as an explicit input; the SLA design decision record must cross-reference the severity classification decision record's notification tier specification, so that a change to either document triggers a review of the other.

Section 4: Severity inflation detection model and tier drift governance. Specify the severity inflation detection process: the leading indicators that a tier definition has drifted from its intended urgency calibration, the review trigger thresholds for each indicator, and the revision process for updating the tier specification when drift is detected. Define three leading indicators: (1) Top-tier incident rate — the number of P0 (or equivalent highest tier) incidents per month; threshold: if the rate exceeds 2 per month for a 30-day rolling window during a period when product usage has not increased proportionally, schedule a tier definition review within 15 business days; a well-calibrated top tier should produce 6–15 incidents per year for a typical SaaS product in stable operation, not 2 per week; (2) Escalation acknowledgment time — the time from P0 incident creation to first response action by the incident commander; measure P99 acknowledgment time on a rolling 90-day basis; if P99 exceeds 3x the tier's defined response time commitment, the tier is exhibiting desensitization symptoms and requires immediate review of the definition and the recent classification distribution; (3) Downgrade rate — the fraction of incidents that are reclassified to a lower tier during or after resolution; if more than 25% of top-tier incidents are downgraded at post-incident review, the classification criteria are being applied inconsistently, typically because the criteria are ambiguous or lack worked examples; threshold: any quarter with a downgrade rate above 25% triggers a definition review with the explicit goal of adding decision trees or worked examples that reduce ambiguity. Specify the tier revision process: tier definitions are not changed unilaterally by the on-call team or engineering management during an active incident; revisions happen through a documented review process outside of incident time, with changes recorded in the severity classification decision record as a superseding revision; the revision must include a rationale section explaining what signal prompted the review, what change was made, and what impact is expected on the leading indicator that triggered the review. Connect this section to the on-call load management decision record: severity inflation detection and alert noise management are structurally parallel problems — both involve a signal that should carry consistent urgency being over-triggered, degrading the behavioral response calibration of the recipients; the on-call load management decision governs the aggregate volume and actionability of monitoring alerts; the severity classification decision governs the urgency calibration of incident tiers; both must be reviewed together because a tier definition change that reduces P0 volume will reduce on-call escalation burden, and an alert volume change that reduces non-actionable pages will reduce the noise floor within which a P0 notification must be distinguished.

Section 5: Severity classification review cadence and decision record maintenance. Specify the review cadence for the severity classification decision record: the criteria that trigger an ad hoc review outside the scheduled cadence, the scheduled review interval, the ownership of the review, and the documentation requirements for each revision. Ad hoc review triggers: any incident where the post-incident review identifies a classification ambiguity or inconsistency; any incident where the actual response time exceeded the tier's SLA by more than 50%; any quarter where a leading indicator (top-tier rate, acknowledgment time, downgrade rate) crosses the review threshold defined in Section 4; any enterprise customer complaint citing an SLA breach that was caused by a notification timing gap rather than a response failure. Scheduled review interval: semi-annual review of the tier definitions against the prior 6 months of incident classification data; the review should include a distribution analysis (what fraction of incidents fell into each tier), an impact assessment (what customer-visible consequences did each tier's incidents produce — support tickets, NPS survey impact, customer escalations), and a calibration check (were there incidents where the post-incident consensus was that the initial classification was significantly wrong in either direction). Ownership: the incident severity classification decision record is owned by the engineering manager responsible for on-call operations, with input from customer success (who see the customer-impact distribution of incidents) and legal/contracts (who see the enterprise customer SLA compliance record). Documentation requirements for revisions: each revision must include the version number, the revision date, the trigger (ad hoc or scheduled), the specific change made, the rationale, and the expected impact on the leading indicators; revisions must be communicated to the on-call rotation and the customer success team within 5 business days of approval. Connect this section to the observability strategy decision record: the observability strategy specifies which signals are collected and at what fidelity; the severity classification review depends on those signals for the distribution analysis and the calibration check — specifically, the per-incident customer impact data (support ticket correlation, NPS survey timing correlation) that validates whether the tier assignment matched the actual customer-impact severity; an observability strategy that collects P99 latency and error rates but does not correlate incident timing with customer feedback signals will produce an incident distribution that looks correct internally but misses the customer-impact evidence that reveals classification blind spots; the observability strategy and the severity classification decision must be designed to produce the data the review process needs.

FAQ

How do you define P0 vs P1 in a way that doesn't produce severity inflation?

Define severity tiers using two independent dimensions: customer impact scope (what fraction of customers are affected) and functionality impairment type (what customers cannot do). P0 should require both a broad scope threshold and a complete impairment: a condition affecting more than 25% of active users with complete inability to perform a core workflow, or any condition involving data loss or security breach regardless of scope. P1 should cover either the broad scope threshold with partial impairment, or narrow scope with complete impairment for a named enterprise account. This two-dimensional definition prevents single-customer edge cases from satisfying the P0 criterion unless the edge case involves data loss, security, or complete inaccessibility for a named enterprise account. Specify the scope threshold numerically — 'majority of users' without a percentage allows incident commanders to apply different interpretations under pressure. Review the classification distribution quarterly: more than 20% of incidents at the top tier indicates the definition is too broad. A well-calibrated P0 definition should produce 6–15 incidents per year for a typical SaaS product, not 2–3 per week.

Should degraded performance be its own severity tier or folded into P1?

Degraded performance should be its own tier, distinct from P1. Degraded performance and partial functional unavailability have different response profiles — a feature that is unavailable requires a bug fix or infrastructure recovery; degraded performance may require scaling, query optimization, or cache eviction — different skills and tooling. If folded into P1, the response procedure must cover both cases with a generic process that is either too broad to be useful or requires mid-incident branching. A dedicated degraded performance tier (Sev-3, P1.5, or 'Degraded') allows a distinct response procedure, distinct notification threshold (status page within 30 minutes, lower urgency than P1 but higher than P2), and distinct post-incident review requirements. Define the tier quantitatively: P99 latency exceeding 3x the 14-day rolling baseline for more than 10 minutes, or error rate above 1% on user-facing endpoints, or throughput below 50% of the 7-day rolling average. Without quantitative thresholds, the 'degraded performance' classification is subjective and will not be applied consistently by different incident commanders under incident pressure. The key test: if customers regularly contact support reporting that the product felt broken or unusable, but the incident log shows no incidents above P2 for the periods they describe, you have a missing degraded performance tier.

How do you handle enterprise customers with tighter notification SLAs than your internal response SLA?

Map enterprise customer notification obligations explicitly during onboarding. Ask each enterprise customer: what is your contractual obligation to notify your downstream customers or internal stakeholders when your service is disrupted by an upstream provider? Record the obligation and calculate the cascade gap for each severity tier (the difference between the customer's downstream deadline and your notification SLA for that tier). For customers where the cascade gap is negative — the downstream obligation is shorter than your SLA — implement per-customer accelerated notification: the incident management tooling generates an automatic notification to flagged accounts when any incident at P1 or above is classified that affects their service area, independent of whether the global notification threshold is met. This requires tooling investment but preserves classification consistency (you do not change the tier definition for one customer) while fulfilling the contractual obligation. The alternative — changing the global notification SLA to match the tightest enterprise obligation — commits the entire escalation chain to a higher standard and may not be operationally achievable for all incident types. Document the per-customer notification logic in the severity classification decision record and in the customer contract, so that both the incident response team and the customer success team have a shared reference.

How do you detect severity inflation before it desensitizes the escalation chain?

Track three leading indicators: (1) Top-tier incident rate — the number of P0s per month; if the rate exceeds 2 per month in a rolling 30-day window during a period of stable usage, schedule a definition review; a well-calibrated top tier produces 6–15 incidents per year, not 24–36. (2) Escalation acknowledgment P99 — the time from incident creation to the incident commander's first documented action; a rising P99 acknowledgment time for the top tier is the behavioral signal of desensitization before it manifests in a missed major incident response; measure on a rolling 90-day basis with a threshold at 3x the tier's defined response commitment. (3) Downgrade rate — the fraction of top-tier incidents reclassified to a lower tier at post-incident review; above 25% indicates the criteria are ambiguous and require worked examples or decision trees; high downgrade rates mean incident commanders are over-classifying under pressure, which is the mechanism that produces inflation. Review all three monthly. When any indicator crosses its threshold, schedule a definition review with a specific mandate: not to raise the bar arbitrarily, but to identify the ambiguity or boundary case that is causing the over-trigger, and to add a worked example or numerical criterion that makes the boundary unambiguous. The goal is a P0 definition that fires reliably for genuine major incidents and fires rarely for anything else — not a definition that is easier to apply by firing less often across all cases.

Further reading

  • Incident response playbook decision record — the response procedure for each severity tier; the playbook is indexed by tier and specifies the escalation chain, communication actions, and post-incident review requirements for each classification; the severity classification decision is the input the playbook depends on, and a classification definition that produces inconsistent tier assignments will produce inconsistent response procedures regardless of the playbook's quality — both decisions must be made together with explicit cross-references.
  • SLA design decision record — the customer-facing commitments that the severity classification must support; the notification SLA for each tier is a function of both the internal response capacity (specified by the SLA design decision) and the tier's notification threshold (specified by the severity classification decision); enterprise SLA cascade gaps appear when these two decisions are made independently without mapping enterprise customer downstream obligations as a joint input to both.
  • Alerting threshold decision record — the sensitivity and routing of the alert rules that fire for each severity tier; the degraded performance tier criteria require specific alert rules configured at the same quantitative thresholds; a tier without alert coverage will not be applied consistently in practice because incident commanders classify from the alerts available to them; the alerting threshold decision and the severity classification decision must be aligned so that every tier has the alert coverage needed to trigger it reliably without manual metric inspection.
  • On-call load management decision record — the alert volume budget and escalation load model that a well-calibrated severity classification directly affects; severity inflation increases the escalation frequency at the top tier, consuming the on-call escalation budget and producing the same desensitization effect at the management escalation layer that alert noise produces at the on-call engineer layer; both decisions must be reviewed together when either the top-tier incident rate or the escalation acknowledgment P99 crosses a threshold.
  • Observability strategy decision record — the signal inventory that produces the data the severity classification review process depends on; correlating incident timing with customer feedback signals (NPS surveys, support ticket rate) requires that both signals exist in queryable form and are tracked with sufficient timestamp precision to match incident windows; an observability strategy that covers P99 latency and error rates but does not include customer feedback signal correlation will miss the NPS and retention evidence that reveals performance degradation classification gaps.
  • Open-source extractor — find the severity classification decisions buried in your AI chat history: the planning session where the P0 definition was drafted quickly to get the on-call policy finished before launch without modeling the scope dimension, the post-incident review where the engineering manager mentioned that the P0 rate felt high but no formal review was scheduled, and the enterprise customer onboarding conversation where downstream notification obligations were discussed but not recorded in the contract — each is a recoverable decision record that explains the structural gap that the next missed P0 response or enterprise SLA breach will exploit.