The on-call rotation design decision record: why the rotation scope model you chose determines your informal escalation demand surface and your institutional knowledge transfer failure mode
The on-call rotation design is not a scheduling decision — deciding who is on call for which services and on what cadence. It is a set of founding decisions about what operational knowledge is required to resolve the incidents that occur in production, which engineers hold that knowledge, and whether the rotation scope matches the knowledge distribution in the engineering organization. When the rotation scope is narrower than the knowledge required — when engineers are on call for services they cannot diagnose independently — the formal rotation produces an informal escalation demand on the subject matter experts who are not scheduled. When the rotation scope is wider than a single engineer's knowledge — when everyone rotates across all services — the mismatch between who is paged and who has context produces an orientation tax in every on-call response that falls on an unfamiliar service. And when the rotation design relies on a primary-secondary escalation structure to transfer operational knowledge between engineers, it embeds an assumption about escalation rates that is almost never specified and almost never met. Three failure patterns: the 44-person infrastructure tooling company where a full-platform rotation assigns all engineers to all services and 14 months of incident data show a 37-minute MTTR gap between incidents where the paged engineer owns the service and incidents where they don't; the 37-person B2B SaaS where service ownership rotation produces single-engineer knowledge concentration that is invisible until the service owner leaves and the new owner's first critical incident takes 4.5 hours to resolve; and the 45-person fintech SaaS that builds a primary-secondary model designed to transfer knowledge through shared on-call experiences — and discovers 18 months later that the primary-to-secondary escalation rate is 3%, the secondary engineers have never independently managed a critical incident, and the knowledge transfer the rotation was designed to produce never occurred.
A 44-person B2B infrastructure tooling company built an internal developer platform — service deployment pipelines, secrets management, infrastructure provisioning, and observability tooling — for 340 customer engineering teams across financial services, healthcare, and manufacturing verticals. The engineering organization of 22 engineers was structured into four platform teams of four to six engineers each: the Deployment team, the Observability team, the Secrets team, and the Provisioning team. Each team owned three or four production services. In total, the platform ran 12 production services.
The on-call rotation was designed at the engineering organization's formation as a full-platform rotation: all 20 engineers above the junior level were added to a single rotation covering all 12 production services. The rotation was designed with two principles in mind. First, no engineer should be exempt from operational work — on-call coverage is a shared engineering responsibility, not a specialty. Second, broad operational exposure would make every engineer a generalist responder who could handle any production issue, reducing single-service dependencies. The rotation had 20 engineers, producing an average on-call frequency of approximately one week in twenty per engineer.
The platform ran incident tracking across all 12 services using PagerDuty. After 14 months of operation, a senior site reliability engineer ran an analysis of MTTR data across 347 pages. The analysis segmented each page by two variables: the service that generated the alert, and whether the on-call engineer at the time was a primary owner of that service — meaning a member of the team that built and maintained it. The results were stark. For incidents where the on-call engineer was a primary service owner, the median MTTR was 14 minutes. For incidents where the on-call engineer was outside the service's owning team, the median MTTR was 51 minutes. The 37-minute gap was consistent across all four service clusters and all four years of seniority levels represented in the rotation. It did not narrow as engineers accumulated on-call weeks — engineers who had been on call 8 or 9 times showed the same non-owner MTTR as engineers who had been on call twice, suggesting that the operational knowledge required to close the gap was not being accumulated through rotation exposure alone.
Investigation of the 51-minute non-owner incidents revealed a consistent pattern the SRE analyst called the orientation tax: the first 20 to 40 minutes of a non-owner incident response were spent on activities that required no diagnostic skill but consumed time — reading the service overview documentation to understand what the service did and what its expected operational state looked like; locating the service's primary monitoring dashboards (not linked from the runbook, which assumed the engineer already knew where they were); identifying which other services the failing service depended on and which services depended on it; and finding the name of the primary service owner to contact for help. This orientation activity was completed before any diagnostic work on the actual failure occurred. The 14-minute owner MTTR reflected incidents where none of these orientation steps were needed — the owner knew the service, knew the dashboards, knew the dependency graph, and began diagnosis immediately. Connect this pattern to the runbook quality decision record: the runbooks for each service had been written by the service owners, for engineers with the service owner's contextual knowledge; they specified diagnostic steps that assumed familiarity with the service's internal architecture, referenced dashboards by abbreviated names that were legible only to engineers who already knew which dashboards existed, and described failure modes in terms of internal service concepts that engineers without service context could not apply without additional background reading; the runbooks were accurate representations of the knowledge needed to resolve the incidents they described, but they were written for engineers who already had the knowledge, making them useless for the orientation task they were most needed for.
The informal escalation pattern was the second finding from the analysis. Across the 347 incidents, 112 included a direct message to a service owner who was not the on-call engineer. These contacts were not formal secondary escalations — the rotation had no secondary structure — they were informal Slack messages asking for help diagnosing an unfamiliar service under incident conditions. Seventy-one of the 112 informal contacts were to four specific engineers who were primary owners of the two highest-incident services. These four engineers received an average of 2.8 informal on-call contact messages per month during their off-call weeks — not pager alerts, but Slack messages requiring their attention during business hours, evenings, and weekends because a non-owner engineer was on call and needed their context. The four engineers had individually noted this burden in engineering retrospectives and performance reviews; none of the notes had been connected to the rotation design, and no rotation change had been made in response. The informal escalation demand was real, sustained, and concentrated on exactly the engineers who carried the most operational knowledge — but because it was not tracked in the incident management system and did not appear in the on-call metrics, it was invisible to the engineering leadership making rotation design decisions.
A 37-person B2B SaaS company built a contract lifecycle management platform for mid-market professional services firms — contract creation, redline tracking, approval workflows, and signature collection for law firms, accounting firms, and management consultancies with between 50 and 500 users each. The engineering team of 15 engineers was organized into three product teams of four to five engineers each: the Documents team, the Workflow team, and the Integrations team. Each team owned three to five production services; the platform ran 14 services in total.
The on-call rotation was designed as a service ownership rotation: each team had its own sub-rotation, and engineers were only on call for services owned by their team. The design logic was straightforward — a paged engineer who knows the service they are paged for will resolve the incident faster than a paged engineer who doesn't. The Documents team ran a four-engineer sub-rotation covering six services; the Workflow team ran a four-engineer sub-rotation covering five services; the Integrations team ran a three-engineer sub-rotation covering three services, including the payments service — the integration with Stripe that handled subscription billing and usage metering for the platform's 214 active customer accounts.
The payments service had been built primarily by a senior engineer, Marcus, over the course of 18 months. Marcus had been at the company since the founding team and was the technical author of the service's core billing calculation logic, the subscription state machine, and the Stripe webhook handling pipeline. Two other engineers on the Integrations team were also on the payments service sub-rotation, but in practice, both engineers worked primarily on the OAuth integrations and the CRM sync connectors; Marcus handled payments-related code reviews, architecture decisions, and the majority of on-call responses for payments incidents. The payments service generated approximately 40% of the Integrations team's total incident volume, and Marcus resolved 70% of payments incidents during his on-call weeks and an additional 20% informally during other engineers' weeks via Slack.
Marcus announced his resignation in the fourth year of the company's operation. The engineering team recognized the risk: the payments service was complex, had incomplete documentation, and was the service with the highest single-engineer knowledge concentration in the organization. The company ran a two-week handoff process — Marcus walked his replacement, Priya, through the service architecture in three documentation sessions, answered questions, and co-authored an updated runbook. The handoff felt thorough. Priya was a strong engineer with three years of payments integration experience at a prior company and deep familiarity with the Stripe API. The knowledge transfer sessions produced 22 pages of service documentation and an updated runbook covering the five most common incident classes Marcus had encountered.
Three weeks after Marcus's last day, Priya was on her first independent on-call week for the payments service. A P1 incident was triggered at 2:14am on a Wednesday: the subscription renewal batch job was failing with a foreign key constraint violation on the subscription_line_items table, affecting 47 customer accounts whose monthly renewal should have processed overnight. The total customer impact was approximately $34,000 in billing that had not been processed and 47 customer dashboard states showing "payment failed" for renewals that should have succeeded. Priya had not seen this specific failure mode in her two-week handoff. The updated runbook covered five incident classes — all five of which Marcus had specifically chosen as the most common; this class had appeared once in the prior three years and Marcus had resolved it in 19 minutes using knowledge of the service's internal state management that he had not thought to document because it was background knowledge rather than a discrete procedure. Connect this pattern to the on-call handoff decision record: the knowledge transfer that Marcus and Priya ran was thorough for the knowledge that Marcus was aware of — the explicitly known failure modes, the architectural decisions, the procedural runbooks; it could not transfer the implicit operational knowledge that Marcus held without knowing he held it — the failure modes he had resolved without opening a runbook because they were familiar, the service states that looked alarming but were normal under specific conditions, the Stripe webhook behavior that varied in ways that were not documented anywhere because Marcus had learned it empirically over 18 months and it had never surfaced as a question because no one else had needed to know it; knowledge transfer processes that are structured around explicit knowledge cannot capture implicit knowledge, because the possessor of implicit knowledge does not know what to include in the transfer.
Priya spent the first hour working through the runbook, identifying that the failure was not covered there, and escalating to the Workflow team's on-call engineer who had no payments service context. At 3:40am she reached a former Integrations engineer — now at a different company — via Slack. He recognized the failure mode immediately from a description of the constraint violation and directed her to a specific database migration that had changed the foreign key relationship in a way that left a narrow window for this class of failure on batch job retries. Priya reached a correct diagnosis at 3:58am and resolved the incident at 4:41am: 2 hours and 27 minutes of resolution time, plus 2 hours before the resolution path was identified, for a total of 4 hours and 27 minutes of customer-impact duration on a payment failure that affected 47 accounts. Marcus's historical MTTR for payments incidents in the same severity class was 22 minutes. The 4.5-hour vs. 22-minute gap was not explained by Priya's skill level — she was a strong engineer who resolved the incident correctly once she had the missing context — but by the knowledge concentration the service ownership rotation had allowed to accumulate undetected for 18 months.
A 45-person fintech SaaS company built a cash flow forecasting and treasury management platform for mid-market businesses — cash position dashboards, forecast modeling, automated reconciliation, and bank feed integrations for 310 business customers with between $2M and $50M in annual revenue. The engineering team of 19 engineers was organized into two product teams and one platform team. The platform ran 9 production services; 4 of the 9 services were in the payment processing cluster, which handled bank transaction ingestion, reconciliation, and the core forecast calculation engine.
The on-call rotation used a primary-secondary model. Each rotation week had a designated primary on-call engineer and a designated secondary on-call engineer. The secondary was on standby during the primary's week and was expected to engage when the primary escalated. The rotation was designed with an explicit secondary knowledge transfer rationale: by including a secondary engineer in the on-call schedule for every rotation week, the secondary would build operational knowledge of the services in the rotation through shared incident experience — observing the primary's incident responses and building the context needed to serve as primary in future weeks. The rotation covered all 9 services, primary and secondary both drawn from a pool of 12 eligible engineers across all three teams.
The primary-secondary model was documented in the on-call runbook as follows: "If you have been triaging an incident for 15 minutes and have not confirmed a root cause, consider paging the secondary for support." The word "consider" was chosen deliberately — the on-call process owner had discussed whether to make secondary escalation mandatory and decided against it, reasoning that requiring escalation would undermine engineer autonomy and that engineers should use their judgment. The escalation criteria were described as guidelines, not requirements. Secondary on-call engineers were not notified of incidents unless paged — they were not added to incident channels, did not receive status updates, and had no visibility into what was happening during the primary's on-call week unless they were explicitly engaged.
Over 18 months of operation, the rotation processed 312 primary on-call weeks. The secondary was successfully paged in 9 of them — a primary-to-secondary escalation rate of 2.9%. Reviewing the 9 escalations: three were escalations to secondary for services outside the primary's team ownership, where the primary had explicitly no context for the failing service; three were escalations triggered by the primary being unavailable (traveling, at a family event); two were escalations because the incident required code changes that required a second engineer; and one was a genuine knowledge gap escalation where the primary could not diagnose a payment processing failure. In 303 of 312 primary on-call weeks, the secondary was never engaged. The pattern was consistent across all rotation participants: primary engineers resolved incidents independently or informally contacted subject matter experts through Slack, but almost never paged the secondary through the formal escalation path. Post-incident interviews with primary engineers revealed a consistent rationale: paging the secondary felt like announcing to the team that you couldn't handle the incident; informal Slack contacts felt collaborative rather than performative; and the secondary escalation path was understood culturally as a last resort for engineers who were truly stuck, not as a standard tool for time-bounded triage decisions.
At the 18-month mark, the head of engineering conducted a rotation resilience review. The review asked a specific question: if the three engineers who had been primary on-call for the payment processing cluster the most frequently were to leave simultaneously, what was the secondary engineer pool's capacity to cover that cluster? The answer was that four engineers were designated as potential secondaries for the payment processing cluster. Each had been secondary on-call for that cluster between 2 and 4 times over the 18 months. Each had been engaged during a payment processing incident exactly zero times — because the three experienced primaries had resolved all payment processing incidents independently. The four designated secondaries had never independently triaged a payment processing incident, had never used the payment processing dashboards under incident conditions, had never run the diagnostic queries in the payment processing runbook, and had never read the payment processing postmortem archive. They were in the rotation for a service cluster they had no operational experience with, under the assumption that 18 months of secondary on-call coverage had produced knowledge transfer that in fact had not occurred at all. Connect this pattern to the incident postmortem culture decision record: the rotation's knowledge transfer failure was directly related to the engineering culture's treatment of escalation — a culture that treats escalation as an admission of competence failure will suppress escalations across every mechanism that uses escalation as a knowledge transfer vehicle; the primary-secondary model designed to share operational knowledge through incident co-response is a knowledge transfer vehicle that depends on escalations occurring, and a culture that suppresses escalations also suppresses the knowledge transfer; the postmortem culture and the rotation design are coupled decisions — an organization that has not explicitly established that escalation is expected and valued, and that has not specified the criteria under which escalation is required rather than optional, cannot rely on escalation-dependent knowledge transfer mechanisms in its rotation design.
Structural properties set by the on-call rotation design decision
Three structural properties are determined when an engineering organization decides — or fails to explicitly decide — the rotation scope, the escalation model, and the knowledge transfer mechanism for its on-call rotation. The properties are not visible in the rotation schedule. They are established by whether the knowledge required to resolve the incidents that occur in production is held by the engineers scheduled to respond to those incidents, and they determine the organization's capacity to maintain incident response quality through growth, attrition, and service evolution.
Property 1: The informal escalation demand surface. When the on-call engineer is paged for a service they lack the context to diagnose, they choose between two paths: formal escalation through the defined escalation structure, and informal contact with the subject matter expert most likely to know the answer quickly. The informal path is consistently preferred when the formal escalation path carries a social cost — the perception that escalating reveals an inability to handle the incident. The result is a dual on-call burden: the scheduled engineer carries the pager and the formal responsibility, while the subject matter expert carries an untracked informal co-coverage burden during their off-call weeks. The informal escalation demand is highest when the rotation scope is wide relative to the knowledge distribution: full-platform rotations that include engineers with uneven familiarity across services produce consistent informal escalation demand on the engineers with the deepest service knowledge. The informal demand is invisible in incident tracking systems that record only pager alerts and formal escalation events, which means the total on-call burden in a knowledge-mismatched rotation is systematically underestimated. The practical measure of informal escalation demand is not derivable from PagerDuty data alone: it requires counting direct messages to engineers who are not the scheduled on-call responder for incidents involving their services. When this count reveals that specific engineers are regularly being informally contacted during their off-call weeks, the rotation design is producing a second, untracked on-call rotation for those engineers — one that they cannot decline because the contacts arrive as informal requests for help rather than as pager alerts they can escalate or hand off. Connect this property to the on-call handoff decision record: the informal escalation demand surface is highest for services where the operational knowledge is concentrated in one or two engineers, and the on-call handoff is the moment when that knowledge concentration must be addressed — the outgoing primary on-call engineer's handoff note should include not just what happened during the shift but which informal escalation contacts occurred, because those contacts are the signal that the runbook does not cover the knowledge gap the incoming engineer will encounter; an on-call handoff process that does not capture informal escalation patterns is missing the most actionable signal about where the rotation design's knowledge coverage is insufficient.
Property 2: The institutional knowledge transfer failure mode. Every rotation design embeds an implicit assumption about how operational knowledge is transferred between engineers as the rotation evolves — as engineers join the team, change roles, and leave. Service ownership rotations assume that operational knowledge is captured in runbooks and documentation sufficient for any new owner to operate the service independently after a handoff. Primary-secondary rotations assume that knowledge is transferred through shared on-call experiences, with secondary engineers building operational familiarity by observing and participating in incidents during their secondary weeks. Both assumptions fail consistently, and for predictable reasons. The service ownership assumption fails because the knowledge most important for incident response — the failure-mode patterns the service owner knows empirically, the service states that are alarming but normal under specific conditions, the historical context that explains why the service behaves the way it does — is implicit in the owner's operational experience and is not transferred by documentation sessions that focus on explicit knowledge. The documentation captures what the owner knows they know; it cannot capture what they know without knowing they know it. The primary-secondary assumption fails because escalation rates are controlled by culture, not by rotation design, and a culture that treats escalation as a competence signal will suppress escalations to near zero, eliminating the shared incident experience that the knowledge transfer mechanism depends on. The institutional knowledge transfer failure mode is the gap between the operational knowledge the rotation design assumes is held by the engineers in the rotation and the operational knowledge those engineers actually hold. This gap is invisible until it materializes — in the form of a 4.5-hour MTTR on a service that should take 22 minutes to diagnose, or in the form of a secondary engineer who has been in the rotation for 18 months and has never operated the services they are designated to back up. Connect this property to the incident response playbook decision record: the incident response playbook specifies escalation criteria, escalation paths, and expected escalation time windows; a rotation design that relies on escalation for knowledge transfer requires that the playbook specify escalation as expected and non-optional for defined diagnostic windows — not as an option the engineer considers based on self-assessment; a playbook that says "consider escalating if you haven't resolved the incident" produces a culture of non-escalation; a playbook that says "page the secondary if you have not confirmed root cause within 15 minutes of initial triage" produces a measurable escalation rate and an audit trail for whether the escalation criterion is being met; the rotation design and the incident response playbook must specify the escalation trigger jointly, and neither document is complete without the other.
Property 3: The rotation scope-runbook quality coupling. The rotation scope determines the knowledge gap between the on-call engineer and the service they are covering, and the runbook quality requirement is a direct function of that gap. A service ownership rotation, where only owners are on call for their services, requires runbooks that capture the implicit knowledge the owner carries — because ownership transfers are the moments when the knowledge gap appears and the runbook is the only bridge. The critical failure of service ownership rotation runbooks is that they are written by owners for owners: the owner knows the service and writes the runbook as a reference for things they might forget under incident pressure, not as an orientation document for an engineer who has never seen the service before. The resulting runbooks are accurate but inaccessible to anyone without the prerequisite context the owner assumed. A full-platform rotation requires a different class of runbook: one written for the engineer with the least familiarity with the service who is nonetheless eligible to be paged for it. This class of runbook includes service overview context (what the service does, what its expected operational state looks like, and what customers experience when it fails), dashboard locations and interpretation guidance, dependency maps, and failure-mode descriptions that require no assumed knowledge of service internals. The gap between these two runbook classes is large, and most organizations' runbooks fall into the service ownership class because the engineers writing the runbooks are the service owners, writing for their own reference. A rotation design change from service ownership to full-platform scope, or from full-platform to cluster scope, is not complete without a corresponding runbook quality audit: every runbook in the new rotation scope must be reviewed against the knowledge level of the least-familiar eligible engineer, and the delta between the current runbook's knowledge prerequisites and that engineer's knowledge must be closed before the rotation scope change takes effect. Connect this property to the runbook quality decision record: the runbook quality decision record specifies the authorship model, the review cadence, and the quality standard for operational runbooks; the rotation design decision record must specify the rotation scope, and the rotation scope determines the runbook quality standard the runbook quality decision record must achieve; these two decision records are upstream and downstream of each other — the rotation scope is the upstream specification that sets the runbook quality requirement, and the runbook quality decision record specifies how that requirement is met; an organization that has a thoughtful runbook quality decision record but has not specified the rotation scope is measuring runbook quality against an unanchored standard; an organization that has specified the rotation scope but has not connected it to the runbook quality standard is designing for knowledge gaps rather than against them.
The on-call rotation design ADR: five sections
Section 1: The rotation eligibility and scope specification. Begin the on-call rotation design decision record by specifying who is eligible for on-call rotation and what service scope each eligible engineer covers. The eligibility specification has two components: the organizational scope (which engineers participate in the rotation: all engineers above a seniority threshold, engineers who have completed a specified operational readiness checklist for the services in their coverage scope, or engineers nominated by their team for on-call readiness) and the knowledge prerequisite per coverage scope (the specific competencies an engineer must demonstrate before being added to a given service cluster's rotation). The knowledge prerequisite is the more important specification and the one most consistently omitted: most rotation eligibility policies specify seniority or tenure ("engineers above L3 are on call") without specifying what an eligible engineer must know about the services they will be paged for. A seniority requirement does not address the knowledge gap between a senior engineer and a service they have never operated; an operational readiness checklist that requires the engineer to review the runbook and confirm they can complete the three most common diagnostic procedures does. The coverage scope specification must also be explicit about the boundary: which specific services are in the on-call engineer's scope for this rotation week, and what the engineer should do if they are paged for a service outside their specified scope (page the service owner directly, escalate to the secondary, or decline the alert and re-route). Connect this section to the runbook quality decision record: the operational readiness checklist for each coverage scope must be derived from the runbooks for the services in that scope; if the runbook for a service cannot be used to construct an operational readiness checklist — because the runbook assumes knowledge prerequisites that the eligibility check cannot verify — the runbook is not meeting the quality requirement the rotation scope demands; the eligibility check is the test of whether the runbooks are written for the rotation scope they are covering.
Section 2: The escalation structure and invocation criteria. Specify the escalation structure — primary, secondary, tertiary, service owner contact — and the explicit criteria under which escalation is required rather than optional. The most important word in the escalation model specification is "required" — not "consider" or "may escalate" but "page the secondary if the following criterion is met." The escalation criterion must be measurable: "no confirmed root cause hypothesis within 15 minutes of initial alert acknowledgement" is measurable; "if you can't handle it" is not. The specification must also address the social framing of escalation, because the escalation rate is controlled by culture and not by the escalation policy alone. An explicit statement in the rotation documentation that "escalating when the criteria are met is the correct action and is not a signal about the primary engineer's competence or knowledge level" does not eliminate the social cost of escalation, but it provides the primary engineer with organizational support for escalating and establishes a documented expectation that the engineering leadership has publicly committed to. The escalation criteria should also specify the reverse: the conditions under which the primary is expected to resolve the incident independently and should not page the secondary — because an escalation policy without a "do not escalate for" clause will produce over-escalation in some environments and the secondary will stop taking escalations seriously. Connect this section to the incident response playbook decision record: the escalation structure and invocation criteria in the rotation design decision record are one of the inputs to the incident response playbook; the playbook specifies what the on-call engineer does when a P1 alert fires, and that procedure includes when to escalate, to whom, and how; the rotation design decision record specifies the escalation model; the incident response playbook specifies the escalation procedure; both are required for the escalation model to be consistently applied under incident conditions, because the rotation design document is not typically consulted during an incident — the playbook is.
Section 3: The knowledge transfer model specification. Specify how engineers who are not primary service owners build operational knowledge for the services in their rotation scope, using a mechanism that does not depend on escalation invocations occurring. Three mechanisms that do not require escalations. First, mandatory incident shadowing: the secondary on-call engineer is automatically added to the incident channel for all P1 and P2 incidents during the primary's rotation week, receives the same alerts in a read-only mode, and is responsible for writing the first draft of the postmortem if the primary requests it. The secondary does not respond to the incident unless paged, but they observe the full incident response in real time and can identify knowledge gaps to address in the post-rotation review. Second, required dry-run sessions: before each new rotation cycle, the incoming primary and secondary engineer review the top three runbooks for the highest-incident services in their coverage scope together — walking through the diagnostic steps, identifying which steps require context not captured in the runbook, and updating the runbook based on the review. The dry-run session produces one knowledge transfer event per rotation cycle that is not dependent on a production incident. Third, post-rotation knowledge transfer notes: the outgoing primary is responsible for producing a structured knowledge transfer note within 24 hours of their rotation ending, covering three topics — incidents that occurred (summary of what happened, how it was diagnosed, and any context required to repeat the diagnosis that is not in the runbook), service state observations (anything about the service state at rotation end that the incoming primary should be aware of, including borderline conditions or recent changes with uncertain implications), and informal contacts received (any Slack messages from engineers outside the rotation asking for help with services in the rotation scope, which signal knowledge gaps in the runbook the recipient identified during the rotation). Connect this section to the on-call handoff decision record: the on-call handoff decision record specifies the format and required content of the shift transition between the outgoing and incoming on-call engineers; the knowledge transfer model in the rotation design decision record specifies how knowledge is built up over multiple rotation cycles, not just transferred at shift boundaries; the two decisions are complementary — the handoff decision record ensures knowledge is passed correctly at each transition, while the rotation design's knowledge transfer model ensures that the knowledge available to pass at each transition is growing over time; an organization that has a thorough handoff process but no knowledge transfer model will have well-executed handoffs of a static knowledge base that does not expand; an organization that has a knowledge transfer model but a poor handoff process will have engineers who have accumulated knowledge through the rotation cycle but cannot reliably transmit it to the incoming engineer at transition time.
Section 4: The rotation scope-runbook quality coupling specification. Specify the runbook quality standard for each coverage scope in the rotation, defined as the knowledge level of the least-knowledgeable engineer eligible for that scope's on-call rotation. For each service in the rotation, the runbook quality specification should include three elements. The knowledge prerequisite baseline: the service description written for an engineer with no prior exposure to the service — what it does, what its expected operational state looks like, what customers experience when it fails. A full-platform or cluster-scope rotation that includes engineers from outside the service's owning team requires this baseline; a service ownership rotation does not, but should include it anyway to support ownership transitions. The dashboard and tooling map: the locations of the primary monitoring dashboards, the interpretation guidance for the most important metrics on those dashboards, and the diagnostic queries or commands that are used in the majority of incident responses. These should be links, not names, because the engineer who is orientating to an unfamiliar service under incident pressure cannot reliably find dashboards from abbreviated names. The failure-mode index: a catalogue of the five to ten most common failure modes for the service, written as pattern-response pairs ("if you see X in the logs alongside Y in the dashboard, the cause is usually Z and the resolution is W") rather than as general diagnostic guidance. The failure-mode index should be built from postmortem histories and updated after each incident that reveals a new pattern. The rotation design decision record should specify the trigger for a runbook quality review: any change in rotation scope that adds a service to the coverage of engineers who are not primary owners of that service requires a runbook quality review for that service before the scope change takes effect. An ownership transfer — where a service's primary owner leaves or changes roles — also triggers a runbook quality review, because the new owner will be operating the service without the implicit knowledge the prior owner held, and the runbook must close the gap that the prior owner bridged implicitly.
Section 5: The concentration risk review cadence. Specify the cadence and methodology for identifying single-engineer knowledge concentration risks before they materialize as ownership transfer incidents. The concentration risk review should be conducted quarterly and should examine three signals. MTTR variance by engineer: calculate the median MTTR for each service cluster, segmented by whether the on-call engineer is a primary owner of the services in that cluster. A gap of more than 2x between owner and non-owner MTTR is a concentration risk signal — it indicates that the rotation scope includes a knowledge gap that is not being bridged by runbook quality or accumulated rotation experience. Informal escalation demand: review the Slack message history or any available informal contact records for the prior quarter, specifically counting direct messages to engineers who were not the scheduled on-call responder for incidents involving their services. The concentration of informal contacts identifies which engineers are carrying operational knowledge that the runbooks and formal escalation model are not capturing. Secondary escalation rate: for rotations with a primary-secondary structure, track the escalation rate over the review period. An escalation rate below 10% over any 3-month window is a signal that either the formal escalation criteria are too permissive ("consider escalating") or the culture is suppressing escalations — both of which mean the primary-secondary knowledge transfer mechanism is not functioning. The concentration risk review should produce a prioritized list of services and engineers where knowledge concentration risk is highest, and a set of specific actions to address each risk before the next review cycle. These actions may include runbook quality improvements, dry-run session scheduling, engineering rotation adjustments, or documentation sessions with high-concentration engineers. Connect this section to the incident postmortem culture decision record: postmortem findings are a direct input to the concentration risk review — a postmortem that identifies a knowledge gap in the incident response (the on-call engineer spent 30 minutes orienting to the service before beginning diagnosis, or the runbook did not cover the failure mode encountered) is a signal that the rotation scope or runbook quality has a gap that the concentration risk review should address; an organization whose postmortem culture produces systems attribution findings will produce postmortem findings that identify rotation scope and runbook quality as contributing conditions rather than attributing the delayed diagnosis to the individual engineer's lack of knowledge; the postmortem archive, reviewed through a rotation design lens, is the highest-fidelity source of knowledge gap signals available to the rotation design review.
FAQ
What are the main on-call rotation scope models and their tradeoffs?
Three primary rotation scope models are used in practice. Full-platform rotation: all eligible engineers cover all production services. Benefit: broad operational familiarity; no single-engineer concentration risk. Cost: orientation tax when engineers are paged for unfamiliar services; MTTR variance by engineer-service familiarity; runbooks must be written for engineers with no service context. Service ownership rotation: engineers are only on call for services their team owns. Benefit: the paged engineer is the expert most likely to resolve the incident quickly. Cost: single-engineer knowledge concentration that becomes acute on ownership transfer; the rotation is only as resilient as the team is large. Service cluster rotation: engineers cover a defined cluster of related services. Benefit: engineers develop operational depth across a meaningful service surface without full-platform scope; cluster expertise is more transferable than individual service expertise. Cost: requires explicit cluster boundary maintenance as the service topology evolves. Most companies start with service ownership rotation and discover the concentration problem reactively after an ownership transfer incident rather than proactively.
How should a primary-secondary escalation model be designed to actually transfer operational knowledge?
The most common failure mode of primary-secondary models is designing knowledge transfer as a side effect of escalation invocations. Escalation rates are controlled by culture, not rotation design — if the culture treats escalation as an admission of competence failure, the escalation rate will be near zero. Four design elements that produce knowledge transfer regardless of escalation rate: (1) mandatory incident shadowing — the secondary is automatically added to incident channels for P1/P2 incidents during the primary's week, observing without responding unless paged; (2) explicit escalation criteria that specify the conditions under which escalation is required ("if you cannot confirm diagnosis within 15 minutes of initial triage") rather than optional; (3) required dry-run sessions before each rotation cycle where primary and secondary walk through the top three runbooks together; and (4) post-rotation knowledge transfer notes where the outgoing primary documents incidents, service state observations, and any informal escalation contacts they received during their week.
How do you detect and address single-engineer knowledge concentration in an on-call rotation?
Three detection signals. First, MTTR variance by engineer: if MTTR for a specific service cluster varies by more than 2x between engineers who are primary owners and engineers who are not, the variation is explained by knowledge concentration rather than incident complexity variation. Second, informal escalation tracking: count direct messages to engineers who are not the scheduled on-call responder for incidents involving their services — this reveals the hidden co-coverage burden that knowledge concentration places on specific engineers. Third, runbook knowledge prerequisite audit: review runbooks against the knowledge level of the least-experienced engineer eligible for that service's on-call coverage; if the runbook assumes knowledge of service internals or dashboard locations that only a primary owner would have, it is written for knowledge concentration. Once detected, address concentration through required secondary participation in resolved incidents, runbook rewriting for the least-familiar eligible engineer, and explicit knowledge transfer sessions scheduled before any planned departure or role change.
What should an on-call rotation design decision record specify?
Five specifications. First, the rotation eligibility and scope: who is eligible, what service scope they cover, and the operational readiness knowledge prerequisite for each coverage scope. Second, the escalation structure and invocation criteria: primary-secondary-tertiary structure, measurable escalation criteria that require escalation rather than suggest it, and an explicit statement that meeting the escalation criteria is expected and not a performance signal. Third, the knowledge transfer model: the mechanism by which engineers build operational knowledge for services in their rotation scope in a way that does not depend on escalation invocations occurring — including mandatory incident shadowing, dry-run sessions, and post-rotation knowledge transfer notes. Fourth, the rotation scope-runbook quality coupling: the minimum runbook quality standard for each coverage scope defined as the knowledge level of the least-knowledgeable eligible engineer, and the trigger for a runbook quality review on scope changes or ownership transfers. Fifth, the concentration risk review cadence: a quarterly review examining MTTR variance by engineer, informal escalation demand, and secondary escalation rate to identify knowledge concentration risks before they materialize as ownership transfer incidents.
Further reading
- On-call handoff decision record — the rotation design decision record specifies how operational knowledge is built across rotation cycles; the on-call handoff decision record specifies how that knowledge is transmitted at each cycle boundary; a rotation design that produces knowledge growth through mandatory shadowing, dry-run sessions, and post-rotation notes requires a handoff process that captures and transmits that knowledge accurately at each transition — the two decisions must be specified together as a complete knowledge management system for the on-call rotation; the handoff format should include the outgoing primary's informal escalation contact log, because those informal contacts are the clearest signal of where the rotation design's knowledge coverage has gaps that the incoming primary will encounter in their first days on call.
- Runbook quality decision record — the rotation scope is the upstream specification that determines the runbook quality requirement; a full-platform rotation that includes engineers with no service-specific context requires runbooks written for engineers who are orienting to an unfamiliar service under incident pressure, with no assumed knowledge of service internals, dashboard locations, or failure-mode history; a service ownership rotation requires runbooks that capture the implicit operational knowledge the owner holds, so that ownership transitions do not produce knowledge gaps; the runbook quality decision record specifies how the runbook quality requirement is met and maintained, but it cannot specify what that requirement is without knowing the rotation scope; specifying the rotation scope in the rotation design decision record and the runbook quality standard in the runbook quality decision record as separate, connected decisions ensures that the two documents can be reviewed independently for quality and consistency.
- Incident response playbook decision record — the escalation structure specified in the rotation design decision record is one of the critical inputs to the incident response playbook; the playbook specifies what the on-call engineer does step by step when a P1 alert fires, and step N+1 of that procedure is "escalate to secondary if you have not confirmed root cause within 15 minutes of initial triage" — language that comes from the rotation design's escalation criteria; the rotation design document and the playbook must be consistent, and the playbook is the document the engineer consults under incident pressure; if the playbook says "consider escalating" and the rotation design says "escalation is required after 15 minutes without diagnosis," the playbook's language will dominate the engineer's behavior, because the rotation design document is not typically read during an incident.
- Incident postmortem culture decision record — the on-call rotation design and the postmortem culture are coupled through two mechanisms: the escalation culture and the knowledge gap feedback loop; the postmortem culture's attribution frame determines whether escalation is treated as a systems tool or an individual competence signal, and that determination controls the escalation rate, which controls whether primary-secondary knowledge transfer occurs; the postmortem findings — if the postmortem culture produces systems attribution findings — identify rotation design gaps and runbook quality gaps as contributing conditions in incidents where delayed diagnosis was a factor; an organization with a high-quality postmortem culture and a poorly designed rotation will produce postmortem findings that reveal knowledge concentration risks; an organization with a well-designed rotation and a poor postmortem culture will lack the feedback loop that identifies where the rotation design needs to be updated as the service topology evolves.
- Open-source extractor — find the on-call rotation design decisions buried in your AI chat history: the 1-on-1 where a senior engineer mentioned they keep getting Slack messages during their off-call weeks from engineers who don't know the payments service; the all-hands where someone proposed expanding the rotation to all engineers and the engineering manager said "great idea, broad operational ownership" without specifying what operational readiness meant for services outside each engineer's team; the architecture review where the rotation scope was discussed and someone said "the secondary will build knowledge over time" without anyone asking how often the secondary is actually paged; and the postmortem review where the 4.5-hour MTTR on the payments service was attributed to "Priya needing more time to ramp up" rather than to the knowledge transfer gap the service ownership rotation had allowed to accumulate; recovering these decisions from your AI chat history makes your rotation design a set of five specified dimensions with documented rationale, rather than a schedule that was set once and has drifted into whatever knowledge distribution the current team happens to have.