The on-call rotation compensation decision record: why the compensation model you configured determines your on-call burden accumulation surface and your senior engineer retention failure mode

On-call compensation is decided once, early — during a Series A headcount session, a first postmortem review, or an engineering all-hands — against an incident rate, team size, and customer base small enough that the burden the policy produces is invisible. The dollar amounts are right for the product that exists when the policy is written. What the founding on-call decision does not specify is what incident rate the model was calibrated against, what happens to the rotation structure when the team loses members from a high-incident sub-rotation, or how a secondary rotation designed for rare mentor-style escalation is supposed to work when commercial growth makes secondary pages a routine event. Three failure patterns develop from those omissions: the team that discovered its founding engineers were leaving in a cluster over on-call burden while the compensation policy that had been correct at 7 engineers and 3 pages a week had never been updated for 11 pages a week across an effective rotation of 4; the team that discovered two engineers who were carrying a high-incident sub-rotation alone for four months both resigned within a 3-month window; and the secondary rotation engineers who had been earning $0 additional compensation for 3am pages for eighteen months because the "rare" assumption in the founding decision had not been revisited since the commercial launch.

A 27-person SaaS company that built workflow automation for human resources teams had established its on-call compensation policy in its eighteenth month, after its first significant production incident. Before that point, on-call coverage had been informal — the founding engineering team of five had shared a single group Slack notification and handled incidents as they arose, without any compensation adjustment. As the team grew and the product expanded to its first twelve enterprise customers, the informal arrangement stopped working: engineers were unclear about who was responsible for overnight incidents, response times were inconsistent, and two incidents within a week had gone unacknowledged for over an hour because each engineer assumed someone else was watching the alerts.

The engineering manager at the time proposed a structured on-call rotation with explicit compensation. After a two-hour discussion at an engineering all-hands, the team settled on a weekly rotation across all 7 backend and infrastructure engineers, with a stipend of $200 per on-call week and $50 per page that resulted in a wakeup call between midnight and 6am. The policy was recorded in a Notion document titled "On-call compensation policy v1" and approved by the CFO the following week. The document specified the dollar amounts, the rotation schedule tooling (PagerDuty), the escalation path (on-call engineer → engineering manager → CTO), and the expectation that on-call engineers would acknowledge alerts within 15 minutes. What it did not specify was the incident rate the model was calibrated against, the maximum total weekly compensation at which the model remained fair without a policy review, or the trigger for revisiting the policy when the incident rate changed materially.

At the time the policy was written, the system generated 2–3 pages per week in total, of which 0–1 were wakeup-hour pages. The expected on-call week cost the company $200–$250 per engineer and the expected on-call experience for an engineer was: acknowledge 2–3 alerts during business hours, investigate 0–1 alerts at night, and remain available via phone for the remainder. The engineers who attended the all-hands discussion understood this framing. The policy document did not record it.

Over the following 22 months, the product grew from 12 enterprise customers to 47. The customer base included two financial services firms whose regulatory requirements drove a set of integration points with third-party payroll systems, two healthcare organizations whose HIPAA compliance requirements added a layer of audit-log and data-isolation complexity, and a cluster of mid-market customers whose usage patterns produced a predictable Monday-morning traffic spike. The system's incident rate grew with this complexity: at 22 months after the policy was established, the service was generating 8–11 pages per week, of which 3–5 were wakeup-hour pages. The $200 weekly stipend and $50 wakeup-page increment were unchanged. A bad on-call week — 11 pages, 5 wakeup calls — now cost the company $450 in compensation while consuming 12–15 hours of an engineer's personal and overnight time. A typical on-call week — 9 pages, 4 wakeup calls — cost $400.

Three of the five founding engineers who had built the core systems left within a 6-month window. Their exit interview responses cited the on-call burden consistently: not the raw dollar amounts, but the ratio of compensation to burden relative to what they could earn at companies with lower incident rates or larger rotations. The engineering manager had not flagged the compensation model as a retention risk because aggregate payroll analysis did not surface the on-call stipend as a material compensation gap and the incident rate growth had been gradual enough that no single month looked dramatically different from the prior month. The pattern was visible only in retrospect, when the three departure dates were plotted against the incident rate trend.

The company hired two replacement engineers and promoted an internal engineer from a junior role to fill the third position. The two new hires required a 12-week ramp before they could be effective on-call participants — they needed time to understand the system architecture, build familiarity with the most common failure modes, and develop the diagnostic intuition that reduced incident resolution time from 45 minutes to 12 minutes. During the 12-week ramp, the effective rotation shrank from 7 to 4 engineers. Each remaining engineer's on-call frequency increased from one week in seven to one week in four — a 75% increase in on-call frequency on top of the already-elevated incident rate. One of the four remaining engineers who carried the rotation during the ramp period left 3 months after the first cohort's departures.

A 31-person B2B SaaS platform that built data pipeline tooling for analytics teams had structured its on-call rotation by service ownership. The engineering team was organized into three service areas: the data ingestion pipeline (3 engineers), the processing and transformation cluster (3 engineers), and the customer-facing API and dashboard (3 engineers). Each service area had its own on-call sub-rotation: one engineer per service area was on-call each week, responsible for incidents in their area. Each engineer was on-call one week in three. The compensation model was flat — $300 per on-call week for all three sub-rotations — with no per-page increment.

The founding decision that established this structure was recorded in a configuration document in the team wiki. It specified the three service areas, the engineer assignments, the compensation amount, and the rotation schedule. It did not specify the minimum sub-rotation size below which the structure required a formal review. It did not record the incident volume distribution across service areas. It did not specify what would happen to the compensation model or the rotation structure if a sub-rotation member left and was not immediately replaced. These omissions were benign at founding, when the three-by-three structure was symmetric and the team understood informally that the ingestion pipeline was the most complex service area and that its incidents required the most specialized knowledge to diagnose quickly.

The informal understanding was accurate. The ingestion pipeline generated 70% of the total incident volume — its failure modes involved timing sensitivity in the handoff between source system webhooks and the internal message queue, edge cases in schema normalization for non-standard source formats, and a set of third-party source connector behaviors that were documented only in the institutional memory of the engineers who had debugged them. Incidents in the ingestion sub-rotation took an average of 34 minutes to resolve; incidents in the other two sub-rotations took an average of 11 minutes. The $300 flat compensation covered both without distinction, which was acceptable at founding because the system's total incident volume was low enough that the absolute burden in each sub-rotation was manageable.

Fourteen months after the rotation structure was established, two engineers left the ingestion pipeline sub-rotation within a 6-week period. One left for a company offering a larger equity package; the other left for a position that did not include on-call responsibilities. Their positions were posted and both were eventually backfilled, but the hiring process took four months. During those four months, the ingestion sub-rotation had 2 engineers instead of 3. The rotation schedule adjusted informally: the two remaining engineers alternated weekly, each on-call one week out of two rather than one week out of three.

The two remaining ingestion engineers carried 70% of the total incident volume at twice the frequency of the adjacent sub-rotations. Each ingestion engineer's on-call weeks included an average of 6.3 pages, 3.1 of which were wakeup-hour pages, at $300 flat compensation. Each API engineer's on-call weeks included an average of 1.8 pages, 0.4 of which were wakeup-hour pages, at $300 flat compensation. The sub-rotation equity gap was not reflected in any report, dashboard, or review because the company tracked on-call compensation as a total payroll line item and tracked incident volume as a system-wide aggregate. No per-engineer, per-sub-rotation burden metric existed. The two ingestion engineers raised the equity issue twice in engineering retrospectives. The discussion was acknowledged but the formal resolution was deferred pending the completion of the backfill hiring process.

One of the two remaining ingestion engineers left three months into the coverage period. The other requested a formal leave of absence six weeks after the first departure. The ingestion sub-rotation, which had started the year with 3 engineers, now had 0 active engineers. The team ran a flattened emergency rotation across all available engineers — including engineers from the processing and API sub-rotations who had limited familiarity with the ingestion system's failure modes — for eight weeks while the backfill process completed and the new engineers were ramped. Incident resolution time in the ingestion area averaged 67 minutes during the emergency period, up from 34 minutes.

A 39-person infrastructure platform company had built a shadow on-call rotation alongside the primary one: every primary on-call engineer had a paired secondary who would be escalated to automatically if the primary did not acknowledge a page within 15 minutes. The secondary rotation was staffed by the same engineering pool on an offset schedule — if engineer A was primary on-call in week 1, engineer B was secondary, and in week 2 engineer B was primary while engineer C was secondary. Each engineer was secondary one week in nine across the nine-person infrastructure team.

The secondary rotation was uncompensated. The founding decision — recorded in the same PagerDuty configuration document that specified the escalation ladder — noted that secondary on-call was a "mentor-style support role: secondary engineers are rarely paged and primarily learn from seeing how incidents are handled." The document specified the 15-minute primary acknowledgment window but did not specify the expected frequency of secondary pages, the expected escalation rate from primary to secondary, or the compensation model for secondary pages if they exceeded the mentor-style frequency assumption. The CFO had approved the primary rotation compensation ($250/week) and had been told that the secondary rotation required no additional budget because secondary engineers would rarely be paged.

For the first year after the rotation structure was established, the secondary frequency assumption held. The infrastructure team served an internal product — a Kubernetes orchestration layer used by the company's own engineering teams — and the primary on-call engineers had deep familiarity with the system's failure modes. Primary acknowledgment times averaged 4 minutes. Secondary escalations occurred 0–2 times per month across all primaries, and in most cases the secondary engineer's role was to join a call already in progress rather than to lead a cold start investigation. The mentor-style characterization was accurate.

In the platform company's third year, the internal orchestration layer was commercialized: the company began offering it as a managed service to external engineering teams at 18 customer organizations. The commercial product required contractual SLAs — 99.9% uptime for enterprise customers, 99.5% for standard — with defined incident response times. The alert configuration was updated to reflect the SLA thresholds. The customer onboarding introduced new failure modes: configuration drift from customer-specific customizations, edge cases in the multi-tenant namespace isolation that had not appeared in internal use, and traffic patterns from customer batch jobs that ran on schedules the internal team had not anticipated. The incident rate grew from 4 pages per week to 14 pages per week over 18 months.

With 14 pages per week and a 15-minute acknowledgment window, the primary engineers began managing their response timing differently. For alerts that fired at 11pm or 3am and were ambiguous — the alert condition was firing but the immediate user impact was unclear — the primary engineers adopted a pattern of acknowledging the alert at 14 minutes, reviewing the available metrics, and then deciding whether to escalate or handle the incident themselves. From the primary's perspective, the 14-minute review window was useful: it provided enough time to assess the scope of the incident before waking up the secondary. From the secondary's perspective, being paged at 3:14am for an incident the primary had been watching since 3:00am meant receiving the page when the immediate diagnosis pressure was highest and with no warm-up context about what had already been investigated.

Secondary engineers were paged 3–4 times per month in the second year of the commercial product. The mentor-style framing had inverted: secondary engineers were no longer observing resolved incidents but were often being handed unsolved incidents at the most acute point of the response. The $0 additional compensation for secondary pages had not been revisited since the commercial launch because the secondary rotation was not tracked as a separate compensation category and the individual secondary page events did not appear in any compensation analysis. Two secondary rotation engineers left within a 9-month period. Exit interviews cited the compensation fairness issue directly — one engineer described it as "being on-call for 3am infrastructure fires once a month for free."

The company redesigned the secondary rotation compensation during the month after the second departure. The redesign set secondary compensation at $125/week (50% of the primary stipend) and added a $30 secondary page increment for any page that occurred between midnight and 6am. The redesign was fair and was recognized as fair by the engineering team. It was implemented nine months after the commercial launch had changed the secondary rotation's effective burden from the mentor-style assumption to the managed service reality. During those nine months, the compensation model had accumulated uncompensated burden silently, producing two retention failures that were visible only in retrospect against the departure interview data.

Structural properties set by the on-call rotation compensation decision

Three structural properties are determined when a founding team configures its on-call rotation and compensation model: how fair the compensation will remain as the incident rate grows away from the calibration baseline, how equitable the rotation will be across sub-rotations with different incident volume distributions, and how much burden a secondary rotation accumulates before it appears in any compensation review. None are labeled explicitly in the founding on-call decision — they are properties that emerge from the assumptions the founding configuration embeds about what incident rate the model was designed for, which sub-rotations carry the most burden, and how often a secondary engineer is expected to be woken up at 3am.

Property 1: The compensation model and the effective hourly rate surface. On-call compensation is typically expressed as a fixed weekly stipend plus an optional per-page or per-wakeup increment, calibrated against a specific incident rate that reflects the product at the time the policy is written. As the product grows in complexity, customer base, and system surface area, the incident rate grows with it — and the compensation amounts stay constant while the effective hourly burden on each on-call engineer increases. The effective hourly rate of on-call time — the total weekly compensation divided by the hours of interrupted or constrained time the on-call week produces — declines without triggering any review mechanism because the compensation policy does not know what the current incident rate is; it only knows the original calibration. The founding engineers who participated in the compensation conversation understand the original calibration; the engineers who join later cannot assess whether the model still reflects the burden it was designed to compensate, because the policy document records the dollars but not the incident rate those dollars were calibrated against. The structural fix is to specify the incident rate baseline in the compensation decision record — the observed pages-per-week and wakeup-pages-per-week at the time the policy was established — and to specify the incident rate threshold at which a scheduled review is triggered automatically: not "if the rate increases significantly" but "when the trailing 30-day average pages-per-week exceeds [N], the compensation review fires within 30 days." The alerting threshold decision record is the upstream lever: the alerting configuration determines the incident volume that reaches the on-call engineer, and alert sensitivity choices that were correct at the founding incident rate may produce a flood of low-quality alerts as the system scales — the compensation model's incident rate baseline and the alerting threshold's sensitivity calibration should be reviewed on the same cadence, because the alert volume growth that drives compensation unfairness often originates in alerting thresholds that were not recalibrated as the system's scale changed. The incident escalation policy decision record connects at the burden distribution layer: the escalation thresholds determine how many incidents the on-call engineer resolves independently versus escalates, and a well-calibrated escalation policy is one of the most effective levers for reducing on-call burden without changing the incident rate; when the escalation policy and the compensation model are reviewed separately, the compensation review may recommend increasing the stipend when the underlying problem is over-alerting or under-escalation that a policy adjustment could address at lower cost.

Property 2: The rotation equity model and the sub-rotation coverage failure mode. When service ownership is the basis for rotation structure, the rotation is equitable only if the incident volume is approximately equal across sub-rotations and the headcount of each sub-rotation is approximately equal. In practice, both conditions are rarely met: some service areas generate significantly more incident volume than others, and team attrition is not uniformly distributed across service areas. A founding decision that specifies the rotation structure and the flat compensation rate without specifying the minimum viable sub-rotation size or the incident volume distribution creates an invisible equity accumulation surface: as the team composition changes, the equity gap between sub-rotations grows without appearing in any metric the company tracks, because on-call compensation is tracked as a flat per-engineer-per-week amount and incident volume is typically tracked as a system-wide aggregate rather than a per-sub-rotation burden metric. The failure mode is not merely financial — it is social. Engineers in high-incident sub-rotations who are aware of the equity gap but lack the documented basis to request a review often surface the concern in retrospectives, where it is acknowledged but deferred. The gap compounds through the deferral period. The structural fix is to specify three things in the rotation structure decision: the minimum viable sub-rotation size (three engineers for a weekly rotation, as a hard constraint with an explicit response protocol for when the size falls below the minimum), the incident volume distribution across sub-rotations (recorded quarterly and used as the basis for per-sub-rotation compensation calibration), and the equity review mechanism (a quarterly review that compares the per-sub-rotation burden distribution against the flat compensation model and triggers a per-sub-rotation compensation adjustment if any sub-rotation's burden exceeds the model's calibration by more than a defined factor). The service ownership model decision record is the structural parent: the on-call rotation structure is a direct consequence of the service ownership model, and the two decisions should be reviewed together; a service ownership model that assigns a disproportionate share of the incident-generating surface area to a single team should be reflected in either a higher sub-rotation compensation for that team or a rotation structure modification that distributes the burden more broadly. The engineering team growth decision record connects at the headcount constraint layer: the minimum viable sub-rotation size is constrained by the total engineering team size and the hiring velocity, and the growth decision record's assumptions about hiring timelines directly determine how long a sub-rotation below minimum viable size will remain in that state before backfill restores it.

Property 3: The secondary rotation design and the uncompensated burden surface. Secondary on-call rotations are typically designed against an escalation rate assumption — the fraction of primary pages expected to result in a secondary escalation. When that assumption holds, the secondary rotation functions as designed: secondary engineers are paged rarely, the compensation is either a small stipend or none, and the framing of secondary on-call as a learning opportunity is accurate. When the assumption is violated by product growth — an increasing incident rate, a more complex system that produces more ambiguous alerts, or a primary response SLA that creates incentives to delay acknowledgment until the secondary window is nearly exhausted — the secondary rotation accumulates uncompensated burden invisibly. The founding decision document that established the secondary rotation as uncompensated records the mentor rationale and the frequency assumption, but the frequency assumption is not a measured commitment; it is a qualitative estimate that is never revisited unless someone specifically flags it. Secondary pages do not appear in the primary on-call compensation calculation. Secondary engineers do not appear in the primary on-call burden report. The gap between the frequency assumption and the observed secondary page rate can persist for months before it surfaces in a retention event or an explicit complaint. The structural fix is to compensate the secondary rotation from the start at a defined fraction of the primary stipend — removing the "rare pages, therefore uncompensated" framing that creates the accumulation surface — and to track secondary page volume separately in the on-call burden report, with the same incident rate threshold that triggers a primary compensation review also triggering a secondary rotation design review when the secondary page rate crosses the design assumption by more than a defined multiplier. The incident escalation policy decision record is again directly relevant: the secondary rotation's page volume is a function of the primary response SLA and the primary engineers' acknowledgment behavior; an escalation policy that sets a 15-minute window creates an incentive for primary engineers to use the window to defer the decision about escalation; a 5-minute window reduces deferral but increases the primary engineer's burden; the design of the acknowledgment window must be specified with awareness of how it shapes the secondary's effective burden. The technical roadmap process decision record connects at the budget layer: secondary rotation compensation that was designed as zero budget may require a budget line when the commercial product makes secondary pages routine; the compensation budget change is a predictable consequence of the commercial product's incident volume growth and should appear in the technical roadmap's headcount and operational cost planning, not surface as an unbudgeted line item after two secondary engineers leave.

What the founding session records and what it omits

The founding session that establishes on-call compensation — typically a short discussion during an engineering all-hands or a leadership sync rather than a formal architecture decision — records the dollar amounts, the rotation schedule tooling, and the escalation ladder. If the team has thought carefully about equity, it records the decision to use a flat stipend versus a per-page model, and the rationale. If the team has a CFO involved, it records the total weekly budget for the rotation. What it does not record is the incident rate the model was calibrated against, the minimum viable sub-rotation size for the rotation structure, the escalation rate assumption behind any secondary rotation, or the conditions under which any of these should be revisited.

These omissions are structurally similar across the decision record series: they are benign at founding, when the omissions are covered by the founding team's shared context. The incident rate calibration is not an omission risk at a 5-engineer team with 3 pages per week, when the founding engineers who set the policy are also the engineers on the rotation and they can assess informally whether the compensation is still fair. The minimum sub-rotation size is not an omission risk when the rotation has 7 engineers and no one is leaving. The secondary rotation frequency assumption is not an omission risk when the product is internal and the primary engineers know the system well enough to resolve 95% of alerts without escalation.

The failure modes develop at different rates. The effective compensation collapse develops from gradual incident rate growth — which is a continuous process that produces no single discontinuous event that would trigger a review, and which is therefore easy to miss until a cluster of departure interviews names it explicitly. The sub-rotation equity failure develops from attrition that is not uniformly distributed across service areas — which is a discrete event (an engineer leaving) followed by a continuous accumulation of inequity during the backfill period. The secondary rotation burden accumulation develops from product commercialization — which changes the product's failure mode distribution, customer SLA requirements, and alert configuration in ways that were not anticipated when the secondary rotation was designed for a lower-volume, internally-understood system.

The on-call rotation compensation ADR closes these gaps by documenting the incident rate baseline, the sub-rotation structure constraints, and the secondary rotation design assumptions at the time the rotation is established — and by specifying the review triggers that fire when those assumptions are violated. The decisions never written down in the on-call compensation domain are not the dollar amounts — those are always in the policy document. They are the incident rate those dollars were calibrated against, the minimum viable sub-rotation headcount, the escalation rate assumption behind the secondary rotation design, and the conditions under which any of these requires a review. The new CTO onboarding problem in the on-call compensation domain is specific: the incoming technical leader finds the weekly stipend amount in the on-call policy document and the rotation schedule in PagerDuty, but cannot determine whether the compensation model reflects the current incident rate, which sub-rotation carries the highest burden, whether the secondary rotation was designed for the current escalation frequency, or what review mechanism exists to catch the gap between any of these assumptions and the current reality. The on-call compensation ADR makes those decisions explicit, auditable, and verifiable against current system state — the incident rate in the monitoring dashboard, the sub-rotation burden in the on-call report, and the secondary escalation rate in the PagerDuty escalation log. The WhyChose extractor finds the on-call compensation discussions in your AI session history — the conversation where the engineering manager thought through the stipend versus per-page tradeoff, debated whether the secondary rotation should be compensated, decided what the acknowledgment window should be, or reasoned about what incident rate the team could sustain — and surfaces those parameters so you can assess which assumptions still hold against the current system's incident volume, team composition, and customer SLA structure.

The on-call rotation compensation ADR: five sections

Section 1: Compensation model and incident rate calibration baseline. Specify the compensation model structure: the weekly availability stipend and its rationale (compensating the constraint of being reachable and available, regardless of whether pages fire), any per-page or per-wakeup increment and its rationale (compensating the actual interruption cost of a page that requires active response), and the total expected weekly cost per engineer at the calibration incident rate. Record the calibration baseline explicitly: the observed average pages-per-week and wakeup-pages-per-week at the time the policy is established, derived from the prior 30-day or 90-day incident history rather than from a qualitative estimate. Specify the review trigger: the pages-per-week threshold at which a scheduled compensation review fires automatically — for example, "when the trailing 30-day average pages per week exceeds twice the calibration baseline, a compensation review is scheduled within 30 calendar days." Specify the review scope: the review must re-examine the calibration baseline against the current incident rate, assess the effective hourly rate the current compensation produces against the current burden, and consider whether the incident rate growth originates in product complexity growth (which implies a compensation adjustment) or in alerting configuration drift (which implies an alerting threshold review per the alerting threshold decision record rather than a compensation adjustment). The compensation model should be decoupled from the incident rate wherever possible: a fixed stipend that compensates availability and a variable per-page increment that compensates interruption provides a model that automatically scales the interruption compensation with the incident rate, reducing the frequency at which a policy update is required while still creating a review trigger when the total weekly cost pattern indicates structural rate growth rather than a transient spike.

Section 2: Rotation structure, equity model, and sub-rotation minimum viable size. Specify the rotation structure: the number of engineers in the rotation, the basis for rotation assignment (full-team rotation, service-area sub-rotations, seniority tiers, or domain specializations), and the compensation model for each tier if multiple compensation levels exist. If the rotation uses service-area sub-rotations, record the incident volume distribution across sub-rotations at the time the structure is established — the fraction of total weekly pages attributable to each sub-rotation's service area — and specify the flat versus differentiated compensation rationale: if the sub-rotations are compensated equally despite unequal burden, record the ceiling above which the equity gap requires a compensation adjustment or a rotation restructuring. Specify the minimum viable sub-rotation size: three engineers for a weekly rotation structure, with an explicit response protocol for when the size falls below three — the response must include a timeline (the structure changes within 30 days), a temporary mechanism (merge with an adjacent sub-rotation, bring in a cross-area engineer with a temporary compensation adjustment, or flatten to a full-team rotation), and the restoration condition (the sub-rotation returns to the specialized structure when the headcount is restored to minimum viable size and the new members have completed a defined ramp milestone). Specify the sub-rotation burden review cadence: a quarterly review that produces a per-sub-rotation burden metric — observed pages-per-week, wakeup-pages-per-week, and average resolution time per sub-rotation — compared against the flat compensation amount, with a defined threshold above which a per-sub-rotation compensation adjustment is required. Connect this section to the service ownership model decision record: the on-call sub-rotation structure is a direct consequence of the service ownership model's domain assignments, and a change to the service ownership model (a new domain added, an existing domain split or merged) must trigger a review of the sub-rotation structure and the incident volume distribution before the change takes effect, not after the first on-call cycle reveals that the new domain assignment has changed the burden distribution.

Section 3: Secondary rotation design, escalation behavior expectations, and compensation model. Specify the secondary rotation structure: the pairing logic between primary and secondary engineers, the secondary page trigger (automatic escalation after the primary acknowledgment window, manual escalation by the primary, or both), and the primary acknowledgment window duration with its rationale. Specify the escalation rate assumption: the expected fraction of primary pages that result in a secondary escalation, derived from the observed escalation rate in the prior 30 or 90 days, or from a reasoned estimate that accounts for the primary engineers' system familiarity and the alert classification quality. Record this assumption explicitly, because it is the key variable that determines whether the secondary rotation's burden matches the design. Specify the secondary rotation compensation model: the weekly secondary stipend (recommended as 40–60% of the primary stipend regardless of the expected escalation frequency, to remove the "rare pages, therefore uncompensated" design) and any per-secondary-page increment for pages during wakeup hours. Specify the secondary escalation rate review trigger: when the observed monthly secondary escalation rate exceeds the design assumption by more than 2x for two consecutive months, a review is triggered that examines the primary acknowledgment window, the primary response effectiveness, and the alerting threshold configuration, in addition to the secondary compensation model — because a 2x excess in secondary escalation rate is typically a signal of primary escalation timing behavior or alerting quality, not just a compensation gap. Connect this section to the incident escalation policy decision record: the escalation ladder, the acknowledgment window, and the escalation trigger conditions are specified in the escalation policy, and the secondary rotation design's escalation rate assumption must be consistent with the escalation policy's design; if the escalation policy is changed (a shorter acknowledgment window, a modified escalation trigger), the secondary rotation design review must be triggered simultaneously, because acknowledgment window changes have a direct and predictable effect on secondary page frequency.

Section 4: Compensation review cadence, trigger conditions, and review documentation. Specify the standing review cadence: at minimum, an annual review that examines the incident rate baseline against the current observed rate, the sub-rotation burden distribution against the flat compensation model, and the secondary escalation rate against the design assumption. Specify the triggered review conditions separately from the standing cadence: the triggered reviews fire when a defined threshold is crossed rather than on a calendar schedule, and should include the incident rate review trigger (pages-per-week exceeds twice the calibration baseline), the sub-rotation equity trigger (any sub-rotation's burden-per-engineer exceeds twice any other sub-rotation's burden-per-engineer for two consecutive months), the secondary escalation trigger (secondary escalation rate exceeds 2x the design assumption for two consecutive months), and the rotation headcount trigger (any sub-rotation falls below the minimum viable size). Specify the documentation requirement for each review: the review must produce a written record that includes the current observed metrics against the baseline values in the decision record, the compensation adjustment decision (adjusted, held, or deferred with documented rationale), and the updated baseline values if the calibration is reset. This documentation serves both the compensation equity function and the retention risk management function: a written record of compensation reviews demonstrates to engineers that the model is actively maintained against the current burden, which is a materially different signal than a static policy document that has not been updated since the product's first year. Connect this section to the technical roadmap process decision record: on-call burden growth is a predictable consequence of product surface area growth, and the compensation model's review triggers should be included in the technical roadmap's operational cost planning so that compensation adjustments are budgeted ahead of the review rather than requiring emergency budget approval after a retention event surfaces the gap.

Section 5: Burden measurement, retrospective integration, and retention signal tracking. Specify the on-call burden metrics that must be tracked and reported: per-engineer pages-per-month, per-engineer wakeup-pages-per-month, per-engineer mean time to resolve, per-sub-rotation burden relative to the flat compensation model, and secondary pages per secondary rotation engineer per month. These metrics must be separated by engineer and by sub-rotation, not only reported as system-wide aggregates — a system-wide average of 4 pages per week can conceal a sub-rotation with 12 pages per week and an adjacent sub-rotation with 1 page per week, which produces the equity failure invisibly. Specify the retrospective integration requirement: on-call burden metrics must be a standing agenda item in the quarterly engineering retrospective, with a defined format — current metrics versus decision record baseline, any triggered review conditions that have fired since the last retrospective, any compensation adjustments made in the prior quarter. Connect this section to the dependency update policy decision record: dependency incidents — failures caused by library upgrades, third-party integration behavior changes, or infrastructure dependency version incompatibilities — are a distinct category of on-call burden that tends to spike after dependency update sprints; the on-call burden report should flag dependency-origin incidents separately so that the compensation review can distinguish between sustainable incident rate growth from product surface area expansion and spiky incident rate growth from dependency management gaps that a policy adjustment can address. Specify the retention signal integration: exit interview themes related to on-call burden should be captured and matched against the compensation review history to identify whether the retention event occurred during a period when a triggered review was overdue; a retention event during an overdue triggered review is evidence that the trigger threshold is too high or the review response time is too slow, and the decision record should be updated with a tighter threshold or a shorter response window. The compensation decision record and the retention data together provide the signal that neither provides alone: the compensation record shows what the model assumed; the retention data shows when the model's gap became a behavioral outcome.

FAQ

What should an on-call rotation compensation decision record specify beyond the dollar amount per week?

Four things. First, the incident rate baseline: the average pages-per-week and wakeup-pages-per-week the compensation model was calibrated against, and the maximum pages-per-week at which the model remains fair without a review — stated as a concrete threshold, not a qualitative judgment. Second, the rotation structure and minimum viable sub-rotation size: which engineers are in each rotation tier, which service areas they cover, and the minimum headcount below which the sub-rotation structure requires a formal review rather than an ad-hoc coverage arrangement. Third, the secondary rotation design assumptions: the expected frequency of secondary pages, the escalation behavior assumption, and the compensation model for secondary pages when they exceed the design frequency. Fourth, the compensation review trigger: the incident rate threshold, the rotation headcount change, or the calendar interval that requires a scheduled review — not a qualitative "if we feel it is necessary" clause but a concrete, measurable condition that fires the review automatically.

How do you design an on-call compensation model that remains fair as the incident rate grows?

Three structural choices. First, decouple the baseline from the variable: separate the weekly availability stipend from the per-page increment; the stipend compensates the constraint of being on-call regardless of incident volume; the per-page increment compensates actual interruptions; when the incident rate grows, the per-page component scales automatically. Second, specify the total weekly cap and the cap-triggered review: when the weekly total regularly exceeds a defined ceiling — say, 150% of the stipend alone — the review fires automatically, because consistent ceiling-hitting indicates the incident rate has grown past the model's calibration and the rotation, alerting thresholds, or service reliability requires attention. Third, include the incident rate in the compensation document and update it quarterly: record the observed average and p95 pages-per-week alongside the compensation amounts; an engineer reviewing the policy can then assess whether the current rate matches the calibration; a documented rate-versus-amount mismatch provides the basis for requesting a review before the gap produces a retention event.

How do you structure secondary on-call rotations to avoid uncompensated burden accumulation?

Three mechanisms. First, compensate the secondary rotation from the start at a defined fraction of the primary stipend — typically 40–60% — regardless of the expected frequency of secondary pages; this removes the "rare pages, therefore uncompensated" framing. Second, track secondary page volume separately and include it in the on-call burden report: secondary pages must appear in the same burden metrics as primary pages, separated by a flag rather than excluded; if the secondary page rate exceeds the design assumption by more than 2x in any quarter, the escalation policy, the primary response SLA, or the secondary rotation design requires review. Third, specify the escalation behavior assumption explicitly in the decision record: the secondary rotation design depends on an expected escalation rate; record this assumption and the observable metric that confirms or refutes it each quarter, so that a violation of the assumption triggers a review rather than accumulating silently until a retention event surfaces it.

What is the minimum viable sub-rotation size for a service-ownership-based on-call structure?

Three engineers is the practical lower bound for a weekly rotation cycle, and it should be specified as a hard constraint in the rotation structure decision record. At three engineers, each engineer is on-call one week in three. At two engineers, each engineer is on-call every other week with no buffer for leave or attrition. At one engineer, the sub-rotation has failed and the engineer is on-call continuously. The minimum-viable-size constraint should specify the trigger response: when a sub-rotation falls below three members, the structure changes within 30 days — either by merging with an adjacent sub-rotation, bringing in a cross-area engineer with a temporary compensation adjustment, or flattening to a full-team rotation. The decision record should also specify which signals constitute a sub-rotation headcount risk before the sub-rotation drops below minimum viable size: a planned departure, a leave request longer than two weeks, or any event that reduces the sub-rotation to two effective members within the quarter should trigger the response proactively.