The engineering planning horizon decision record: why the sprint cadence and planning model you chose determines your production responsiveness ceiling and your roadmap commitment surface

The engineering planning horizon decision is not a productivity decision. It is not about making engineers work faster or reducing planning overhead or following a methodology. It is a set of founding decisions about how the engineering organization allocates capacity across competing claims — new feature development, production issue remediation, technical debt reduction, and infrastructure maintenance — and the mechanisms by which those allocations are adjusted as production state changes. The sprint cadence determines how frequently the organization can reset its capacity allocation. The planning model — locked-scope sprints, rolling backlogs, OKR-anchored quarterly planning, kanban with explicit WIP limits — determines the mechanism by which unplanned work enters and displaces planned work. These two decisions together set two structural properties that are rarely specified at the time the planning model is chosen: the production responsiveness ceiling, which is the maximum speed at which production issues can be scheduled into engineering capacity for remediation, and the roadmap commitment surface, which is the area of risk that accumulates between planning cycles when the assumptions behind commitments change faster than the planning cadence can absorb. Three failure patterns: the 43-person B2B project management SaaS whose sprint-integrity policy accumulates 14 active P1s with a median remediation age of 34 days before anyone asks how many open P1s exist; the 38-person developer API company whose quarterly OKRs are committed against team size and scope assumptions that the production environment makes stale before the first monthly checkpoint; and the 47-person compliance SaaS that implements 1-week sprints to fix a production responsiveness problem and discovers two years later that the sprint length was not the variable controlling remediation time.

A 43-person B2B SaaS company built a project management platform for mid-market professional services firms — project timelines, resource allocation, budget tracking, and client reporting for architecture, engineering, and management consulting firms with between 30 and 300 employees. The engineering team of 18 engineers was organized into three product squads of five to six engineers each, covering the Projects domain, the Resource domain, and the Reporting domain. Each squad ran independent 2-week sprints with a shared planning ceremony at the start of each sprint and a shared review at the end.

The sprint model had been designed with an explicit sprint-integrity policy: once a sprint was committed at the planning ceremony, the sprint scope was locked. Changes to sprint scope required approval from the engineering manager, and the bar for approval was set deliberately high. The rationale was sound: the company had previously run a model with no sprint locking, where the two most senior engineers were constantly being pulled off sprint commitments to handle support escalations and ad hoc feature requests from sales; by the end of each sprint, 30 to 40 percent of sprint commitments were incomplete and rolled over to the next sprint, creating perpetual schedule slip and a team that had stopped trusting its own commitments. The sprint-integrity policy was adopted after a team retrospective identified "constant context switching and scope creep" as the primary driver of missed commitments. The new policy was clear: P0 incidents (complete service outages affecting all customers) could interrupt the sprint and pull engineers off sprint work; all other issues — production bugs, customer escalations, P1 incidents — were triaged by the on-call engineer, documented as GitHub issues, and added to the product backlog for the next sprint planning session.

The sprint-integrity policy succeeded at its stated objective: committed sprint goals began landing at a 78% completion rate within three sprints of implementation, up from 60% under the previous model. The engineering team reported higher satisfaction with sprint predictability and the ability to finish tasks without interruption. The retrospective finding had been correct, and the remedy had worked.

Eight months later, the head of product asked the engineering manager a question she had not previously thought to ask: how many open P1 production issues were in the backlog? The P1 classification covered production bugs that degraded customer experience but did not cause complete service outages — incorrect calculations in the budget tracking module, slow report generation that timed out for customers with large project portfolios, a date parsing edge case that caused the resource allocation view to show wrong assignments for projects crossing a fiscal year boundary. The engineering manager pulled the count from GitHub. The answer was 14 active P1 issues in the backlog, with ages ranging from 7 days to 94 days. The median age was 34 days. Nine of the 14 had been in the backlog through at least two sprint planning ceremonies and had been passed over for sprint inclusion each time because feature work for the active roadmap was given higher priority in planning. Three of the 14 had been reported by the same enterprise customer — the company's largest paying account — whose success manager had filed three escalation tickets that had not produced a fix in 6, 11, and 22 days respectively.

The root cause was not malicious prioritization or poor judgment in sprint planning. Each individual sprint planning session had correctly weighed the value of new features against the P1 backlog items and decided that the roadmap commitments were higher priority in that specific sprint. The problem was that the sprint-integrity policy's routing mechanism for P1 issues — directly to the backlog — had no escalation path that responded to accumulated age. A P1 that had been in the backlog for 7 days and a P1 that had been there for 47 days entered each sprint planning session with equal standing, and the planning session's optimization for roadmap commitments consistently beat the P1 list. The production responsiveness ceiling was not the sprint length: it was the sprint length plus the expected number of sprints a P1 would wait in the backlog before being prioritized — which, given 14 issues accumulated over 8 months, was averaging 2.4 sprints beyond the initial sprint it missed. Connect this pattern to the incident response playbook decision record: the incident response playbook specifies the response procedure for P1 incidents — the triage process, the escalation path, and the communication protocol with customers; the planning model specifies the mechanism by which P1 remediation is scheduled into engineering capacity; these are two separate documents that are rarely written together, and the gap between them is the production responsiveness ceiling; a P1 incident response playbook that promises customers a resolution within 5 business days while the sprint planning model provides no mechanism for pulling P1 fixes into the current sprint has made a commitment the planning model cannot fulfill without explicit authority given to the engineering manager or product manager to interrupt the sprint.

The engineering manager and head of product made two changes after the conversation. First, the sprint planning model was updated to include a P1 queue review at the start of each planning session, with a policy that any P1 older than 14 days was automatically pulled into the next sprint regardless of roadmap priority, and any P1 older than 7 days was discussed in planning with the presumption of inclusion unless a specific rationale was given for deferral. Second, the incident classification was reviewed against the customer communication model: customers reporting P1 issues were being given automated acknowledgement and an informal expectation of resolution "as soon as possible," which the sprint planning model could not meet consistently; the customer-facing SLA was rewritten to reflect the new 14-day ceiling honestly, and the success manager was briefed on the expected resolution timeline for the three enterprise escalations. The 14 open P1s were resolved over the next three sprints. The sprint-integrity policy remained in place; the routing mechanism for P1 issues was changed. The planning model decision had been made correctly given its stated objective; the production responsiveness implication of the routing mechanism had not been specified when the decision was made.

A 38-person B2B developer tooling company built a continuous integration and code quality API — automated test orchestration, code coverage enforcement, and PR analysis for development teams at companies with between 10 and 200 engineers. The engineering team of 14 engineers used a quarterly OKR model: at the start of each quarter, the VP of Engineering and the product manager ran a two-day planning session to produce 6 to 8 OKRs for the quarter, with each OKR's key results owned by one of three engineering sub-teams. The planning session used a tool the VP of Engineering frequently ran through ChatGPT: he pasted the prior quarter's performance data, current headcount, and proposed roadmap items, and used Claude to help prioritize the OKR list and estimate per-OKR engineering scope. The output was a quarterly plan the team reviewed and committed to, with monthly "checkpoint" sessions designed to assess progress and surface blockers.

The company's product was growing rapidly. API call volume was increasing at approximately 40% month-over-month through the middle of the fiscal year, driven by a partnership with a large code hosting platform that had integrated the API as a first-party extension. Q2 had been the best quarter in the company's history: 7 of 8 OKRs landed, three new enterprise contracts signed, and the team had shipped two major features ahead of schedule. The Q3 planning session was conducted in a positive environment, and the commitments made were ambitious but grounded in Q2's demonstrated capacity.

Two variables changed in Q3 that the quarterly planning model could not absorb. The first was team composition: in the first six weeks of Q3, one senior engineer resigned and accepted a competing offer, and a second engineer went on extended medical leave. The team's effective capacity dropped from 14 engineers to 12 within six weeks of committing to OKRs that had been planned against 14. The monthly checkpoint at the six-week mark identified the team size reduction but the checkpoint format was a status review — teams reported progress against key results — not a planning reset; the OKRs were not renegotiated to reflect the capacity reduction, because the VP of Engineering believed the team would close the gap through the quarter's second half and did not want to lower expectations after Q2's strong performance. This is a pattern that the postmortem action item ownership decision record maps onto: the ownership model for planning commitments determines whether the person who owns the OKR has both the authority and the obligation to renegotiate it when the assumptions behind the commitment become invalid; an OKR ownership model that treats the key results as fixed commitments rather than planned outcomes removes the mechanism by which the plan can adapt when the inputs change.

The second variable was production infrastructure load. The 40% monthly growth in API call volume reached a threshold in week 7 of Q3 where the existing infrastructure configuration began showing latency degradation under peak load. The on-call engineer identified the issue — the job orchestration service's queue depth was exceeding the worker pool's processing capacity during peak hours, causing p99 latency for PR analysis jobs to spike from 4.2 seconds to 18 to 24 seconds. Resolving the issue required a combination of horizontal worker scaling and a configuration change to the job prioritization algorithm — work that the VP of Engineering estimated at 2 to 3 engineer-weeks of effort. The work was not in the Q3 OKRs. The on-call engineer and the infrastructure sub-team lead escalated the issue to the VP of Engineering at the month-2 checkpoint. The VP of Engineering authorized the infrastructure work, pulling two engineers from a Q3 feature OKR to address the latency issue. The feature OKR was not renegotiated at the checkpoint; it was marked as "at risk" in the checkpoint notes. It was not completed at quarter end.

Q3 results: 2 of 8 OKRs landed fully. The engineering team's post-quarter retrospective identified the capacity reduction and the infrastructure work as the two causes. The VP of Engineering noted in the retrospective summary that both causes were visible at the month-2 checkpoint and neither was acted on aggressively enough. The more structural finding — which emerged from the retrospective discussion but was not written into the decision record — was that the quarterly planning horizon had been set against assumptions (14-engineer team, current production load) that were both changing on timescales shorter than one quarter, and the monthly checkpoints were status reviews that did not provide the authority or the mechanism to reset the commitments when the assumptions changed. The quarterly OKR model was well-suited to the environment the company was in at the end of Q2; it was poorly suited to the environment the company was in during Q3, and the planning horizon had not been designed to detect that the suitability condition had changed.

A 47-person SaaS company built a compliance workflow management platform — GDPR data subject request handling, SOC 2 evidence collection automation, and internal audit workflow orchestration — for mid-market companies navigating regulatory compliance without dedicated compliance staff. The engineering team of 21 engineers was organized into three product squads and one platform squad, all running 2-week sprints.

At the end of the company's third year of operation, a P1 incident occurred in the GDPR request handling module: the automated email notification system that was supposed to send customers confirmation of data subject requests stopped sending emails after a dependency update changed the email service client's authentication behavior. The issue was reported by three enterprise customers in the first day and by eleven more over the following week. The engineering team triaged the issue and confirmed it was a P1 — no complete service outage, but customer-facing email notifications had stopped entirely, which was a compliance risk for customers who relied on the platform to document their request acknowledgements. The fix was straightforward: a one-line configuration change in the email service initialization and a regression test addition. The fix took one engineer 3 hours to write, review, and test.

It took 18 days from initial report to production deployment.

The root cause of the 18-day delay was not the sprint model. The fix was written and merged into the main branch within 4 days of the initial report — the current sprint had room for a small interruption, and the engineering manager approved pulling the fix into sprint scope. The 14-day delay between merge and production deployment was caused by the deployment approval process, which required four sequential sign-offs: the engineering manager, the VP of Engineering, the security team lead, and the head of compliance. The compliance sign-off was required because any change touching the GDPR request handling module was classified as "compliance-adjacent" and required compliance review regardless of change size. The head of compliance was not available for 9 of the 14 days between merge and sign-off, and the sign-off was not escalated as urgent.

The retrospective identified the approval process as the bottleneck, with two specific findings: the compliance-adjacent classification was too broad (a one-line configuration change should not require the same sign-off as a schema change in the request data model), and the approval SLA was undefined (there was no maximum time-to-sign-off specified, so the head of compliance did not know the fix was time-sensitive and processed it in her normal review queue). The retrospective action items were to narrow the compliance-adjacent classification criteria and to add a time-to-sign-off SLA for production fixes classified as P1. The VP of Engineering agreed with both items and assigned them to the appropriate owners.

Neither action item was implemented. The VP of Engineering noted this himself eight months later when reviewing open action items from prior retrospectives. The approval process was unchanged. The compliance-adjacent classification was unchanged. Connect this pattern to the postmortem action item ownership decision record: retrospective action items that require sign-off or process changes outside the engineering team's authority — in this case, changes to the deployment approval process and the compliance classification system, both of which required buy-in from the head of compliance and the VP of Engineering — are the most likely to remain unimplemented; the ownership model for retrospective action items must specify who has the authority to implement each action and what the escalation path is if the action remains open past the owner's availability to implement it; an action item owned by a person whose involvement requires coordination with stakeholders outside the engineering team requires a different follow-through model than an action item owned entirely within the team.

In month 12, a new VP of Engineering joined the company. After reviewing the production incident history, she proposed a change to the sprint cadence: moving from 2-week to 1-week sprints, on the rationale that faster sprint cycles would enable faster production of fixes and reduce the time-to-deploy for P1 issues. The proposal was discussed at the engineering leadership team meeting, and the primary argument in favor was the 18-day P1 incident from month 8: "if we had 1-week sprints, we could have gotten the fix in a week instead of two." The engineering team voted to implement 1-week sprints, and the change went into effect at the start of the next sprint cycle.

Twenty-six months later, the engineering team ran an annual planning retrospective. The metrics compared the 2-week sprint period (company months 1 through 12) and the 1-week sprint period (months 13 through 38). On production responsiveness: the median time from P1 report to production deployment in the 2-week sprint period had been 11 days (the 18-day incident being the worst case; the majority of P1s were resolved in the 5 to 14 day range). In the 1-week sprint period, the median was 9 days. The improvement was present but small. Review of the specific cases showed that the 2-day improvement was explained by more frequent sprint boundaries giving the engineering manager more opportunities to formally pull P1 fixes into sprint scope, not by any change to the deployment approval process, which remained the dominant delay factor for all fixes that went through the compliance-adjacent path. On ceremony overhead: in the 1-week sprint period, the four squads held planning, review, and retrospective every week, for a total ceremony cost of approximately 7 to 8 hours per engineer per week across the three events. In the 2-week period, the same ceremonies were held every two weeks, for a proportional cost of 3.5 to 4 hours per week. The 4-hour-per-week difference compounded across 21 engineers was 84 engineer-hours per week of additional ceremony overhead — equivalent to two engineer-weeks per month — that was not reflected in capacity planning. On complex-task throughput: engineers with tasks requiring 4 or more consecutive days of focused work reported in the month-24 retrospective that task splitting across sprint boundaries had become a consistent source of integration complexity; three of the five engineers working on the platform team — which handled multi-week infrastructure migrations — reported that the 1-week cadence made it difficult to sustain deep work across sprint boundaries and produced check-in pressure that interrupted focus at 2 to 3 day intervals. The annual retrospective finding: the 1-week sprint cadence had improved production responsiveness by a small margin while increasing ceremony overhead and degrading throughput on complex work, because the variable that controlled production responsiveness — the deployment approval process — was not addressed by the change.

Structural properties set by the engineering planning horizon decision

Three structural properties are determined when an engineering organization decides — or fails to explicitly decide — the sprint cadence, the planning model, and the mechanism for routing unplanned work. The properties are not visible in the sprint velocity dashboard or the OKR completion rate. They are established by whether the planning model provides a functional path for production issues to reach engineering capacity for remediation, whether the planning horizon is shorter than the period over which key capacity assumptions change, and whether the sprint cadence is matched to the type of work the team actually does rather than the type of work the team aspires to do.

Property 1: The production responsiveness ceiling. Every planning model creates a production responsiveness ceiling — a maximum speed at which production issues can be scheduled into engineering capacity, below which the planning model's default routing mechanism cannot operate. For a locked-scope sprint model with a 2-week sprint length, the minimum production responsiveness ceiling is the time remaining in the current sprint when the issue is reported; the practical ceiling is higher because P1 issues enter the backlog and wait for sprint planning cycles where they compete with roadmap features for inclusion. The ceiling is not the sprint length. It is the sprint length multiplied by the expected number of planning cycles the issue waits before inclusion, plus the time from inclusion to deployment. This product is never calculated when the sprint model is adopted because the backlog prioritization behavior is not specified at the same time as the sprint model. The ceiling becomes visible only when the product of those factors produces an outcome that is inconsistent with customer expectations — when an open P1 is 34 days old and a customer asks for a status update. The ceiling can be lowered without changing the sprint cadence: adding a P1 age-based escalation rule (any P1 older than 14 days is included in the next sprint unconditionally), designating one engineer per sprint as an interrupt handler for production issues, or adding a mid-sprint scope modification authority to the engineering manager for issues above a defined severity threshold. The sprint cadence determines the maximum frequency of the planning reset; the planning model determines the actual production responsiveness ceiling, and the two must be specified together. Connect this property to the on-call rotation design decision record: the on-call rotation and the sprint model are the two primary capacity allocation mechanisms in an engineering organization; the on-call rotation allocates one engineer's capacity to production monitoring and first-response during their rotation week; the sprint model allocates all engineers' remaining capacity to planned work; the interaction between the two mechanisms is where production issues that require more than one engineer to fix — or that require expertise outside the on-call engineer's scope — are routed into sprint capacity; if this routing mechanism is not specified, production issues that the on-call engineer cannot resolve independently enter an allocation gap where the on-call rotation is no longer responsible for them and the sprint model's default routing is to the backlog.

Property 2: The roadmap commitment surface. The planning horizon sets the granularity at which capacity commitments are made and the period over which those commitments are assumed to be stable. A quarterly planning horizon makes commitments against assumptions about team size, scope clarity, production load, and organizational priorities that are expected to remain stable for the duration of the quarter. When those assumptions change faster than the planning horizon — when a team member leaves in week 6 of a 13-week quarter, or when production load grows 170% between planning and the first checkpoint — the commitments made at the start of the quarter are now made against stale inputs, and the accumulated error between committed capacity and available capacity compounds across the remaining duration. The roadmap commitment surface is the area of risk created by this accumulated error. It is larger when the planning horizon is longer, when the environment is more variable, and when the checkpoint model is a status review rather than a planning reset. The commitment surface is invisible during the planning cycle — it becomes visible when OKRs are reviewed at quarter end and the post-hoc analysis reveals that the plan was made against inputs that were already diverging from reality at planning time. Reducing the commitment surface requires either shortening the planning horizon (accepting higher ceremony overhead) or improving the checkpoint model's authority and scope (moving from status review to replanning permission, with the VP of Engineering authorized to renegotiate key results when key assumptions change by more than a specified threshold). The choice between these approaches depends on whether the environment's variability is driven by external factors (production load, market changes) or internal factors (team size, scope clarity); external variability is better addressed by checkpoint authority than by cadence reduction, because shorter sprints do not reduce external variability and introduce ceremony overhead that compounds the capacity problem.

Property 3: The process overhead-to-value ratio. The sprint cadence sets the proportional overhead cost of the planning process. A 1-week sprint with a total ceremony cost of 7 to 8 hours per engineer per week is 18 to 20 percent of available capacity allocated to coordination and planning. A 2-week sprint with the same total ceremony cost is 9 to 10 percent. The difference matters not because ceremony time is wasted — planning, review, and retrospective are genuinely valuable — but because the value of those ceremonies is roughly constant per occurrence regardless of sprint length, while the overhead cost doubles as a percentage of capacity when the cadence doubles. The process overhead-to-value ratio also interacts with the type of work the team does: tasks that require 4 or more consecutive days of focused work are interrupted more frequently by sprint boundaries in a 1-week cadence, producing either task splitting across sprint boundaries (which creates integration overhead when the task must be assembled) or ceremonial sprint rollovers (where tasks are committed into sprint N and marked as "in progress" through sprint N+1 and N+2 without the sprint model providing any actual planning value for those tasks). The degradation in throughput on complex work is predictable and should be measured when the cadence change is proposed by calculating the proportion of the current backlog that requires more than 3 days of focused work per task and estimating the ceremony overhead increase. In practice, this calculation is almost never made when a cadence change is proposed, because the proposal is motivated by a specific production responsiveness failure that feels like it could have been prevented by faster sprint cycles — and the calculation that would reveal whether the sprint length was actually the bottleneck requires tracing the root cause of the specific failure through the full deployment path, which typically identifies the deployment approval process or the issue routing mechanism as the controlling variable rather than the sprint cadence.

The engineering planning horizon ADR: five sections

Section 1: The sprint cadence specification with explicit overhead accounting. Begin the engineering planning horizon decision record by specifying the chosen sprint cadence — the length in calendar days, the ceremony structure and estimated time cost per ceremony, and the total proportional overhead cost as a percentage of available engineer-week capacity at the current headcount. The overhead accounting serves two purposes. First, it provides a benchmark against which future cadence change proposals can be evaluated: if the current model costs 9 percent of capacity in ceremony overhead and a proposed 1-week sprint model costs 19 percent, the net capacity impact of the change is visible at proposal time rather than six months after implementation. Second, it establishes the work type profile that the cadence is designed for: specify the expected distribution of work types in the sprint (what percentage of sprint tasks are expected to require more than 3 days of focused work, what percentage are expected to require 1 day or less) and confirm that the sprint length provides enough contiguous work time for the complex tasks to be completed within a single sprint without task splitting. A 1-week sprint in which 40 percent of committed tasks require 4 or more days of work will produce systematic task splitting that is invisible to the sprint velocity metric but degrades throughput and integration quality. The cadence specification should also document what the sprint cadence is not designed to solve: if the motivation for the chosen cadence was production responsiveness, and the actual production responsiveness ceiling is controlled by the deployment approval process rather than the sprint length, this should be stated explicitly so that a future production responsiveness failure does not prompt a cadence change that will not address the root cause. Connect this section to the deployment rollback decision record: the deployment cadence and the sprint cadence are related but not identical; a team can run 1-week sprints and deploy to production on a monthly cadence, or run 2-week sprints and deploy to production continuously; the deployment cadence is the variable that controls the time between a fix being merged and a fix being available to customers; the sprint cadence controls when fixes are pulled into scope; both must be specified to reason accurately about the production responsiveness ceiling, and the planning horizon decision record should reference the deployment cadence and its controlling variables explicitly.

Section 2: The unplanned work intake mechanism. Specify the complete intake model for unplanned work — work that arises after the sprint scope is committed and requires engineering capacity. The intake model has three components. The first is the classification system: define the work classes that determine how unplanned work enters the planning model — at minimum, a system that distinguishes complete production outages (which should interrupt the sprint in all models), significant customer-impacting production issues (which may interrupt the sprint depending on model), and all other unplanned work (which enters the backlog). The boundaries between classes should be objective and auditable: "customer-impacting" should be defined as "one or more customers unable to complete a core workflow," not as "important in the engineer's judgment." The second component is the entry path for each class: the specific action taken for each classification — immediate sprint interruption, pull into sprint with engineering manager approval, add to next sprint planning with presumption of inclusion after N days in backlog, or add to backlog at standard priority. The third component is the authority required for each path: who can authorize pulling an issue into a locked sprint, under what time constraint (is the authorization synchronous or can it wait for the daily standup?), and what happens if the authorizing person is unavailable. A locked-scope sprint model with no specified exception authority creates a production responsiveness ceiling that cannot be lowered without violating the sprint model's rules, because there is no legitimate path for non-P0 issues to enter the current sprint. Specifying the exception authority and the approval path closes this gap without abandoning the sprint-integrity policy. Connect this section to the incident response playbook decision record: the incident response playbook specifies the triage procedure and escalation path for production incidents; the planning horizon decision record specifies what happens to the incident after triage — which planning mechanism routes the fix into engineering capacity; these two documents must be written in reference to each other, because the incident response playbook cannot honestly specify remediation timelines without knowing what the planning model's intake mechanism will do with the fix request, and the planning model's intake mechanism should be written with explicit knowledge of what severity classes the incident response playbook will route to it and at what expected volume.

Section 3: The commitment surface management model. Specify how capacity commitments are adjusted when key assumptions change after the planning session commits sprint goals or OKRs. The commitment surface management model has three elements. First, the trigger specification: the conditions under which a replanning event is warranted — team size changes exceeding a specified threshold (e.g., loss of more than 15 percent of committed capacity), unplanned infrastructure or production work exceeding a specified percentage of sprint or quarterly capacity (e.g., more than 20 percent of capacity allocated to work not in the original plan), or external dependencies failing to resolve within a window that makes committed OKR key results achievable. The trigger must be measurable at the time it occurs, not retrospectively at quarter end. Second, the replanning authority: who has the authority to call a replanning event and renegotiate committed goals when a trigger condition is met — whether this is the engineering manager, the VP of Engineering, or requires product management participation — and the mechanism for doing so (a scheduled mid-sprint replanning session, an async document, or a modification to the sprint scope in the project tracking tool that is visible to all stakeholders). Third, the communication model: how replanned commitments are communicated to the stakeholders who were given the original commitment — product managers, executives, sales, customers — because a planning model that allows renegotiation without clear communication creates two parallel expectations about what engineering will deliver, which are more damaging to trust than the original miss would have been. Connect this section to the on-call rotation design decision record: unplanned production work is one of the primary triggers for a mid-cycle commitment surface management event; when a production issue requires more than one engineer-week of remediation effort, it consumes a meaningful portion of sprint or quarterly capacity that was committed to planned work; the on-call rotation design determines how much of the engineering team's capacity is allocated to production response each week, but it does not account for production issues that require additional capacity beyond the on-call engineer; the planning horizon decision record must specify the mechanism by which the commitment surface is adjusted when a production issue requires capacity above the on-call allocation, and this specification requires knowing the on-call model's allocation assumptions.

Section 4: The production responsiveness protocol. Specify the production responsiveness targets for each production issue class and the mechanism by which the sprint model meets those targets. The production responsiveness protocol has two components. First, the target specification: the maximum time-to-remediation for each issue class, stated as calendar time from initial report to production deployment of the fix. The target must be stated in calendar time, not sprint time, because customers and business stakeholders measure time in days, not sprint cycles. A target of "resolved in the next sprint" is not a calendar time target — it is a planning model artifact that translates to 0 to 28 days depending on when in the sprint cycle the issue is reported. The target should reflect what the engineering organization believes it owes customers and what the business requires to maintain customer trust. Second, the mechanism audit: verify that the existing sprint model can meet each target, given the sprint length, the sprint-integrity policy, and the deployment cadence. If a P1 target of 5 business days is not achievable under the locked-sprint model because the sprint length plus backlog queue depth routinely exceeds 5 days, the target is not achievable without changing either the sprint model's intake mechanism or the deployment approval process. The mechanism audit is the step that most consistently reveals whether the planning model is designed to meet the production responsiveness target or whether the target has been set against the aspiration rather than the reality. It is almost never done when the sprint model is adopted, because the production responsiveness target is typically stated by a different team (product management, customer success) in a different document (SLA, customer agreement) than the sprint model decision. Connect this section to the incident postmortem culture decision record: postmortem findings that identify planning model failures — the P1 was in the backlog for 34 days; the fix was merged but waited 14 days for deployment approval — are the primary signal that the production responsiveness protocol has a gap; postmortem culture determines whether these findings are written as systems attribution findings (the sprint model's backlog routing mechanism has no escalation path for P1 age) or individual attribution findings (the engineering manager did not prioritize the P1 fix); systems attribution findings produce planning model changes; individual attribution findings produce engineer behavior change requests that do not address the planning model gap.

Section 5: The planning horizon review cadence. Specify the cadence and methodology for reviewing the planning model itself — not the sprint outcomes but the model that produces those outcomes. The planning horizon review should be conducted at least annually and should examine four signals. First, production responsiveness ceiling: calculate the actual median and p90 time from P1 report to production deployment for the prior year; compare against the stated target; identify the largest contributors to delay in the cases where the target was missed — sprint model routing, deployment approval process, or engineer availability — and assess whether the sprint model change being considered (if any) would address the largest contributor or a smaller one. Second, OKR or sprint goal miss rate: calculate the percentage of sprint goals or quarterly OKRs missed over the prior year; for each miss, identify whether the cause was planning model failure (the commitment was made against assumptions that changed mid-cycle) or execution failure (the team had the capacity and the correct assumptions but did not complete the committed work); an OKR miss rate above 25 to 30 percent sustained over multiple cycles with planning model failure as the primary cause indicates that the planning horizon is longer than the period over which key assumptions are stable. Third, ceremony overhead as a percentage of capacity: calculate the actual time cost of all planning ceremonies (planning, review, retrospective, grooming) per engineer per week and compare it to the benchmark established in section 1; if ceremony overhead has grown as a percentage of capacity — because team growth has made the planning sessions longer, or because additional ceremonies have been added — assess whether the sprint cadence needs to be adjusted to restore the original overhead ratio. Fourth, complex-task throughput proxy: review the sprint history for the prevalence of task splitting across sprint boundaries — tasks committed in sprint N that are first completed in sprint N+1 or later; a high rate of cross-sprint task splitting in a 1-week sprint model indicates that the cadence is shorter than the natural work unit size for the team's work type, and the resulting overhead and integration cost should be quantified and compared against the responsiveness benefit of the shorter cadence.

FAQ

What is the right sprint cadence for a SaaS engineering team?

There is no universally correct sprint cadence, but the tradeoffs between cadence options can be specified. Two-week sprints are the most common starting point because they balance planning overhead against the frequency of plan refresh; ceremony cost is roughly 3 to 4 hours per engineer per 2-week sprint, approximately 6 to 8 percent of available capacity. One-week sprints double the proportional ceremony overhead and are most appropriate for teams whose work units are small, well-defined, and do not require more than 2 to 3 days of focused work; they consistently underperform for teams doing complex architecture, major feature development, or system migrations where individual tasks require 3 to 5 consecutive days of deep work. Four-week sprints reduce ceremony overhead but extend the production responsiveness ceiling and create commitment accuracy problems in environments where scope understanding changes faster than once per month. The planning model — how unplanned work enters, how sprint scope can be modified, and what authority is required for scope changes — is as important as the cadence, because the cadence sets the maximum refresh frequency while the model determines whether that frequency is actually used to improve production responsiveness.

How should a team handle production P1 incidents that occur mid-sprint?

The handling model for mid-sprint production incidents should be specified explicitly in the planning model decision record, not determined case-by-case. Four common models. First, P0-only interruption: only complete production outages interrupt the sprint; all P1 and below enter the backlog. This maximizes sprint integrity but creates a production responsiveness ceiling equal to the sprint length plus backlog queue depth. Second, time-box with severity threshold: define a time threshold (e.g., 4 engineer-hours) and a severity threshold above which the sprint is interrupted immediately; below both thresholds, the issue enters the backlog. Third, on-call engineer as dedicated interrupt handler: one engineer per sprint handles all unplanned production work; their capacity is not committed to sprint goals. Fourth, rolling backlog without sprint locking: treat the sprint goal as directional and tasks as a prioritized rolling queue that can be reordered at any time by a specified authority. Each model has a different production responsiveness ceiling, a different sprint integrity guarantee, and a different ceremony cost. The model must be specified before the first sprint under the new planning horizon.

What are the signs that the planning horizon is creating commitment accuracy problems?

Four leading signals. First, OKR or sprint goal miss rate above 30% consistently over 3 or more planning cycles — a sustained high miss rate indicates commitments are made against assumptions that don't match delivery conditions. Second, planning session attendance declining and retroactive scope renegotiation increasing — when the team stops investing in planning and starts renegotiating scope after commitments are made, the horizon is longer than team confidence in commitments is stable. Third, unplanned work consuming more than 20% of actual capacity over multiple cycles without the planning model absorbing it — the planned work is being delivered at the cost of untracked work the model isn't capturing. Fourth, open production issues accumulating in the backlog without being resolved — an increasing count of P1 or P2 issues consistently bumped to the next planning cycle is the most direct signal that the production responsiveness ceiling is too high and the planning model's prioritization mechanism is failing to surface them fast enough against roadmap commitments.

What should an engineering planning horizon decision record specify?

Five specifications. First, the sprint cadence and rationale: the chosen cadence, the work types it is designed for, and the ceremony overhead as a percentage of engineer-week capacity. Second, the unplanned work intake mechanism: the classification system for unplanned work, the entry path for each class, and the authority required to override the default entry path. Third, the commitment surface management model: the trigger conditions for a mid-cycle replanning event, the authority to call a replan, and the communication model for renegotiated commitments. Fourth, the production responsiveness protocol: the time-to-remediation target for each production issue class, verification that the sprint model can meet those targets, and the mechanism that bridges the sprint model's routing to the target requirement. Fifth, the planning horizon review cadence: the specific signals — OKR miss rate, unplanned work percentage, ceremony overhead, cross-sprint task splitting rate — that trigger a review of the planning model itself, and the authority required to change the sprint cadence or planning model when a review recommends it.

Further reading

  • Incident response playbook decision record — the incident response playbook and the planning horizon decision record are two documents that are almost never written in reference to each other, and the gap between them is the production responsiveness ceiling; the playbook specifies the response procedure and the customer communication timeline for production incidents; the planning model specifies the mechanism by which fixes are scheduled into engineering capacity; a P1 playbook that promises customers resolution within 5 business days while the sprint model routes P1 fixes to the backlog has made a commitment the planning model cannot meet; writing both documents with explicit cross-references to each other — specifically, including the sprint model's intake mechanism for P1 issues in the playbook's remediation section, and including the playbook's remediation timeline targets in the planning model's production responsiveness protocol — closes the gap that the three failure patterns above illustrate.
  • Postmortem action item ownership decision record — postmortem action items that require changes to the planning model are the most likely category of action items to remain unimplemented, because they require cross-functional authority (changing the deployment approval process, changing the sprint-integrity policy, changing the OKR renegotiation mechanism) that is not always held by the engineering team alone; the action item ownership model must specify who has the authority to implement each class of action item and the escalation path when the owner cannot implement unilaterally; the planning horizon review cadence in section 5 is the mechanism by which postmortem findings about planning model failures are aggregated and acted on at a level of authority that can actually change the model, rather than being recorded in individual postmortem documents and never surfaced to engineering leadership at a cadence that produces model change.
  • Deployment rollback decision record — the sprint cadence and the deployment cadence are related but independently specified; a team can run 1-week sprints and deploy to production every two weeks, or run 2-week sprints and deploy to production on every merge; the variable that controls the time from a fix being merged to a fix reaching production is the deployment cadence and the deployment approval process, not the sprint cadence; the 47-person compliance SaaS shortened its sprint cadence from 2 weeks to 1 week to address a P1 remediation time problem and saw a 2-day improvement over two years, because the deployment approval process was the bottleneck, not the sprint length; the planning horizon decision record should specify the deployment cadence and deployment approval process alongside the sprint cadence, explicitly connecting the two to the production responsiveness ceiling so that future proposals to change the sprint cadence as a production responsiveness remedy can be evaluated against the correct bottleneck.
  • On-call rotation design decision record — the on-call rotation and the sprint model are the two primary capacity allocation mechanisms in engineering, and they must be specified together to reason accurately about production responsiveness; the on-call rotation allocates one engineer's capacity to production monitoring and first-response; the sprint model allocates all remaining capacity to planned work; the interaction point — where production issues that require more than the on-call engineer's capacity enter the sprint model for remediation — is where the production responsiveness ceiling is either closed or left open; the on-call rotation design specifies the escalation model for issues the on-call engineer cannot resolve alone; the sprint model's intake mechanism specifies how those escalated issues enter engineering capacity; these two specifications must be written in reference to each other, specifically addressing the question: what happens to a P1 issue that the on-call engineer cannot resolve within their rotation shift and that requires 3 to 5 engineer-days of remediation work from a second engineer?
  • Open-source extractor — the engineering planning horizon decision is one of the most consistently buried decisions in AI chat history; planning sessions — especially the quarterly OKR sessions, the planning model retrospectives, and the sprint cadence change proposals — are almost always conducted partly through ChatGPT or Claude, where the VP of Engineering or CTO pastes current performance data and asks for help structuring the OKR list, estimating capacity, or evaluating the tradeoffs between sprint cadence options; the decision to implement 1-week sprints, the decision to add a sprint-integrity policy, the decision to adopt quarterly OKRs — these conversations happen in AI chat tools because they are complex, open-ended planning problems that benefit from a thinking partner; the result is that the rationale for the planning model, the tradeoffs considered, and the assumptions made at planning time are buried in AI chat history while the planning document records only the conclusion; when the VP of Engineering who made the original planning model decision leaves, or when the sprint cadence change proposal is made and the team needs to understand why the previous cadence was chosen, the reasoning is unretrievable from the planning documents alone; the extractor surfaces those buried decisions from the AI chat export, making the planning model's rationale findable by the engineering manager who inherits it and the team that must decide whether to change it.