The deployment rollback decision record: why the rollback strategy you chose determines your mean time to recovery ceiling and your partial-deployment blast radius
The deployment rollback strategy — whether rollback is achievable at all after a database migration runs, whether the trigger criteria for initiating rollback are pre-specified or decided under incident pressure, and whether the rollback boundary accounts for schema-application version pairing in blue-green and canary deployments — determines two structural properties of every deployment-induced incident: the ceiling on mean time to recovery, and the blast radius of the rollback operation itself. Most teams assume rollback is available. The reality is that backward-incompatible database migrations make rollback equivalent to forward-patching, that unspecified trigger criteria convert a 3-minute technical operation into a 143-minute decision process, and that schema-application version pairing in partial deployments leaves the system in a broken intermediate state when rollback is applied only to the application layer. Three failure patterns: the 38-person B2B SaaS where a NOT NULL migration renders the rolled-back application unable to insert new rows — the rollback expanded the blast radius rather than containing it; the 44-person developer tools SaaS with a tested 3-minute rollback procedure but no pre-specified trigger criteria — 143 minutes of incident bridge debate preceded 3 minutes of actual recovery; and the 52-person B2B SaaS that correctly initiated rollback during the canary phase but rolled back the application code without rolling back the migration, leaving old-version application code writing NULL values into a precision column the schema had already made required for downstream processing.
A 38-person SaaS company built a contract management and vendor compliance platform for procurement teams at mid-market retail and manufacturing companies — contract lifecycle management, vendor onboarding workflows, and financial reconciliation for transactions in 14 currencies. The engineering team had been running a 2-week sprint cadence for 11 months, with deployments every Tuesday and Thursday. The deployment pipeline ran automated tests, built a new application binary, and executed any pending Alembic database migrations before cutting over the application version. No documented rollback procedure existed; the team's implicit model was that if a deployment caused problems, they would roll back the application binary and the migration would not be an issue because migrations typically added columns or indices without modifying existing data.
In the 43rd deployment, a senior engineer on the financial reconciliation team shipped a feature that stored exchange rate precision metadata for each transaction: a new `exchange_rate_precision` column on the `transactions` table, typed as INTEGER NOT NULL DEFAULT 6 (6 decimal places as the default, matching the ISOstandard for most currency pairs). The migration added the column with a DEFAULT clause, which populated the column for all existing rows with the integer 6. The application code in the new deployment version wrote the precision value at INSERT time, explicitly setting the column for every new transaction. The deployment proceeded normally: migration ran, application binary was swapped, tests passed in production smoke check.
Fourteen minutes after deployment, a reconciliation engineer in a different team noticed that the hourly multi-currency reconciliation report was producing incorrect totals for transactions involving three currency pairs that used non-standard precision — the new code was applying 6-decimal precision to currency pairs that required 8-decimal precision, producing a rounding error visible only in high-value transactions. The error was not a logic bug in the exchange rate precision feature; it was a data error in the precision configuration table that the feature read from, which had been seeded with incorrect values for the three non-standard pairs during a prior configuration migration. The error affected 23 transactions in the 14-minute window since deployment.
The on-call engineer initiated rollback: the application binary was reverted to the previous version in 4 minutes. The 23 affected transactions were identified and queued for correction. Then the first new transaction attempt arrived. The previous application version's INSERT statement for the `transactions` table did not include the `exchange_rate_precision` column — the column had not existed when the previous version was written. PostgreSQL rejected the INSERT with a NOT NULL constraint violation: the column was NOT NULL with a DEFAULT clause, but the DEFAULT clause applies to rows inserted without specifying the column only if the INSERT does not specify a column list. The previous application version's ORM generated INSERT statements with an explicit column list that did not include `exchange_rate_precision`, which bypassed the DEFAULT and triggered the NOT NULL constraint. Every new transaction attempt failed. The contract management platform stopped recording new vendor transactions 4 minutes after rollback completed.
The team spent 110 minutes writing, testing, and deploying a forward patch: a new application version that read the precision configuration table correctly for all currency pairs and wrote the precision value for every INSERT. The 23 incorrect transactions from the 14-minute regression window were corrected manually. The exchange rate precision feature went live correctly in the forward patch. Total incident duration from first incorrect transaction to full recovery: 127 minutes. The duration of the original regression: 14 minutes affecting 23 transactions. The duration of the rollback-induced outage: 110 minutes during which no new transactions could be recorded. Connect this failure pattern to the production change freeze decision record: the production change freeze decision record addresses the policy for which deployments are allowed during high-risk periods and what the risk classification threshold is; the deployment rollback decision record is the recovery path when a deployment that was allowed through produces a regression; the two decision records are complementary controls — freeze prevents the highest-risk deployments from reaching production, and rollback bounds the impact of the deployments that get through; but the rollback decision record must address the migration reversibility question that the change freeze decision record does not: a deployment that passes the risk classification threshold and is allowed during a non-freeze period can still contain a forward-only migration that makes rollback unavailable if a regression is detected after the migration runs.
A 44-person SaaS company built a developer platform for continuous integration pipeline observability — build time analytics, test flakiness detection, artifact dependency tracking, and pipeline cost attribution for DevOps teams at 310 customers in the software development tools segment. The engineering organization had made a deliberate investment in deployment infrastructure after a difficult incident 8 months earlier: blue-green deployment with automated health checks during the traffic promotion ramp, a documented and tested rollback procedure (route traffic back to the blue deployment via a single load balancer rule change, tested monthly with a production chaos drill), and a 99.9% deployment success rate over the prior 8-month period. The rollback procedure was genuinely good: traffic routing reversion took an average of 3 minutes from decision to confirmation in the monthly drills, and the health checks during traffic promotion had caught two regressions automatically and prevented them from reaching full traffic.
What the deployment infrastructure did not specify was the rollback trigger criteria — the conditions under which the rollback procedure should be initiated, who had the authority to initiate it, and what the maximum time limit was for investigation before rollback was required. The deployment runbook contained the rollback steps in detail. It did not contain a decision framework for when to execute them. The implicit model was that "if the deployment causes a regression, roll back" — but what constituted a rollback-worthy regression, as opposed to a regression worth investigating and patching forward, was not specified. Connect this gap to the incident response playbook decision record: the rollback trigger criteria are an incident response decision that has the same structural importance as the escalation criteria and the communication templates — a decision that is made once at calm time and applied automatically at incident time, rather than a decision that is remade under incident pressure each time the situation arises; the incident response playbook that specifies the escalation triggers in detail but does not specify the rollback trigger criteria has specified how the team communicates about incidents but not how the team recovers from the most common class of deployment-induced incidents.
In month 9 of the improved deployment infrastructure, a release was deployed at 2:14 PM on a Tuesday: a refactoring of the artifact dependency tracking service that changed the internal data model for dependency graph traversal to improve query performance for customers with deep monorepo dependency trees. The deployment health checks passed during the 10% traffic ramp, and the release was promoted to 100% traffic. Fourteen minutes after full promotion, the error rate on the artifact dependency tracking API increased from 0.31% (baseline, measured over the prior 7 days) to 1.87% — a 6x baseline increase, representing approximately 340 additional errors per hour. The errors were concentrated on a specific query path: dependency graph traversal queries that crossed more than 3 depth levels, which the new data model handled with a different traversal algorithm than the old model. The new algorithm had a correctness error for graph structures with circular dependency references — a legal structure in some monorepo configurations.
The incident bridge convened at 2:29 PM with 7 engineers: the on-call engineer, the deployment team lead, two members of the artifact dependency tracking team, the engineering manager, and two engineers from adjacent teams who joined when the incident was declared. The deployment team believed the error rate increase was concentrated on a specific graph structure that affected a small percentage of customers and that a targeted patch to the circular dependency handling in the new traversal algorithm could be deployed in 30–45 minutes. The on-call engineer noted that the error rate was at 1.87% and trending toward 2% and asked whether the team should roll back while the patch was prepared. The deployment team lead said they were close to the root cause and preferred to patch forward. No one on the bridge had authority specified in any document to override that preference with a rollback decision. The maximum investigation time before rollback was required was not specified in any runbook or policy.
At 3:31 PM — 77 minutes after the error rate first exceeded baseline — the first patch attempt was deployed. It resolved the circular dependency error for the most common circular dependency structure but did not resolve it for a second, less common structure. Error rate dropped from 2.1% to 1.4% and then increased again to 1.9% as traffic patterns rotated toward customers with the second circular dependency structure. The on-call engineer again asked whether rollback should be initiated. The deployment team believed the second structure was now identifiable and that a second patch could be prepared in 15 minutes. The engineering manager agreed to continue the forward path. The second patch was prepared and deployed at 3:57 PM. Error rate returned to 0.31% within 6 minutes. Total incident duration: 103 minutes from error rate threshold breach to recovery, or 143 minutes from deployment completion to full recovery. Rollback execution time if the rollback had been initiated at the 15-minute threshold (error rate sustained above 3x baseline for 15 minutes): 3 minutes, confirmed by 8 monthly drills. The investigation and two patch attempts collectively added 140 minutes to the incident duration.
The post-incident review identified the trigger criteria gap as the structural cause of the extended incident duration. The rollback procedure was correct. The rollback infrastructure was reliable. The engineers on the bridge were competent. The extended duration was produced by the absence of a pre-specified criterion that removed the rollback-versus-forward-fix debate from the incident bridge entirely. With a trigger criterion of "error rate greater than 3x baseline sustained for 15 minutes triggers rollback unless the deployment team can identify a root cause confirmed as self-resolving within the 15-minute window" — a criterion that the team agreed in the post-incident review was appropriate for this signal pattern — the rollback would have been initiated at 2:44 PM and completed at 2:47 PM, and the circular dependency patch could have been prepared and deployed against the blue deployment during the 110-minute window that was instead spent in incident bridge debate and two patch attempts against the live production system.
A 52-person B2B SaaS company built a human resources and payroll management platform for mid-market professional services firms — payroll processing, benefits administration, and compliance reporting for 215 customers with employees in 8 countries. The engineering organization had adopted canary deployments 6 months earlier, after a full-traffic deployment of a payroll calculation update had produced incorrect paychecks for 12% of the customer base before the error was caught. The canary model: deploy to 10% of traffic for 30 minutes with automated accuracy checks on a representative set of payroll calculations, then promote to 100% traffic if no anomalies are detected, or roll back to the blue deployment if anomalies are detected. The canary model had successfully caught two regressions in the 6 months since adoption, preventing both from reaching the full customer base.
The deployment in question added a multi-currency precision improvement to the payroll line item calculation engine: a `calculation_precision` column on the `payroll_line_items` table (INTEGER, NOT NULL, no default — the application code set the value at INSERT time based on the currency pair configuration for the employee's pay structure), plus a refactoring of the line item calculation service that used the precision column to determine the rounding behavior for each calculation step. The migration ran before the canary traffic routing change. At canary time, the database had the `calculation_precision` column. The green (new version) application wrote the column for every payroll line item INSERT. The blue (old version) application, which was still serving 90% of traffic, did not write the column — the column had not existed when the blue version was written, and the ORM's explicit column list did not include it. But the blue application's INSERTs were not failing: the migration had been written without a NOT NULL constraint, as the engineer writing it had intended to add the constraint in a follow-up migration after verifying the data population was complete. The `calculation_precision` column was NOT NULL only in the engineer's intent, not in the schema — the schema had the column as nullable, and the blue application's INSERTs were succeeding with NULL values for the column.
Thirty-one minutes into the canary period, the automated accuracy checks detected a rounding discrepancy in payroll calculations for employees paid in currency pairs with non-standard precision — specifically, calculations where the green deployment's precision value was producing a result that differed from the historical calculation by more than the accepted tolerance. The discrepancy was a real regression: the precision configuration table had been seeded incorrectly for 3 of the 8 supported currency pairs, the same category of error that had affected the contract management SaaS in the first failure pattern. The canary rollback was initiated correctly: traffic routing was reverted from 10% green / 90% blue to 100% blue in 4 minutes. The green deployment was taken offline. The on-call engineer confirmed that the automated accuracy checks were no longer firing anomalies and marked the incident as resolved.
Two hours later, the hourly payroll batch processing job ran. The batch job read all `payroll_line_items` rows inserted in the previous hour and applied the `calculation_precision` value to finalize the rounding for each line item. Of the 4,200 rows inserted by the batch job's processing window, 3,640 had been inserted by the blue application during the canary period and the post-rollback period — rows where `calculation_precision` was NULL. The batch processing code that read the `calculation_precision` column had been written by the green deployment team to handle the column as NOT NULL (the column's intended constraint, not its actual constraint). The batch code treated a NULL `calculation_precision` value as a calculation error and flagged the affected line items for manual review. 3,640 payroll line items were flagged — 86% of the hour's batch — across 31 customer accounts whose employees had been active during the canary period and the 2-hour post-rollback window before the batch ran.
The remediation required three steps. First, a compensating migration to populate `calculation_precision` with the correct precision value for all rows where the column was NULL — 3,640 rows, requiring a join against the currency pair configuration table to determine the correct precision for each row's currency pair. Second, a backward-compatible code change to the green deployment that handled NULL `calculation_precision` values gracefully rather than flagging them as errors, so that the column could be populated incrementally without the batch processing code blocking on NULL values. Third, adding the NOT NULL constraint to the `calculation_precision` column after the NULL rows had been remediated. Total time from canary rollback decision to full batch processing recovery: 4.5 hours. The canary rollback itself: 4 minutes. Connect this failure pattern to the build artifact provenance decision record: the build provenance model that verifies each production artifact against the deployment package it was built from — including the migration files — provides the artifact integrity guarantee that rollback depends on when the rollback boundary includes the schema state; in the absence of an explicit record of which migrations ran before the canary rollback was initiated, the on-call engineer who initiated the rollback did not have a clear inventory of the schema state changes that were not reversed by the traffic routing rollback, and the batch processing failure 2 hours later was the first signal that the rollback boundary had been incomplete.
Structural properties set by the deployment rollback decision
Three structural properties are determined when a team decides — or fails to explicitly decide — the rollback strategy for their deployments: whether rollback is achievable at all after any given migration runs, what the decision latency cost of unspecified trigger criteria is, and what the blast radius of rollback in a partial-deployment architecture is. None of these properties are visible in the deployment pipeline configuration. A deployment pipeline that reliably runs migrations, builds application binaries, and routes traffic according to a documented procedure does not reveal whether the procedure is executable under rollback conditions until a rollback is needed. The result is that teams discover the limits of their rollback strategy during incidents — at exactly the moment when the limits have the highest cost.
Property 1: The forward-only migration surface and the rollback availability ceiling. The rollback availability ceiling for any deployment is determined by the reversibility of the migrations that ran before rollback was initiated, not by whether the application code can be reverted. A migration is forward-only when the schema change it makes is incompatible with the previous application version's behavior — the most common case is adding a NOT NULL column without a DEFAULT clause, or adding a NOT NULL column with a DEFAULT clause that the previous application's INSERT bypasses by specifying an explicit column list. When a forward-only migration runs and a regression is detected afterward, the available recovery paths are: (1) a forward-fix deployment that corrects the regression in the new application version, which may take significantly longer than rollback; (2) an application rollback with old-code-against-new-schema operation, which may produce a new failure mode depending on whether the old code can function correctly against the new schema; or (3) a compensating migration that reverses the schema change, which requires writing and testing a migration that may involve data loss or data transformation under incident time pressure. None of these paths has the recovery time of a traffic routing rollback to a clean previous state. The migration reversibility classification — forward-only, conditionally reversible, or fully reversible — converts the rollback availability question from a post-incident discovery into a pre-deployment decision that determines which recovery path is available before the deployment is executed. Connect this property to the production change freeze decision record: the risk classification that determines whether a deployment is allowed during a freeze period should include the migration reversibility classification as an explicit input; a forward-only migration in an otherwise low-risk deployment is a risk factor that the change freeze decision record should surface, because the absence of a clean rollback path means the incident recovery time for any regression is bounded by the forward-fix path rather than by the rollback execution time.
Property 2: The rollback trigger specification and the decision latency cost. The mean time to recovery for deployment-induced incidents has two independent components: the time to decide to roll back and the time to execute the rollback. The execution time is bounded by the technical rollback procedure — deterministic, measurable, and testable in non-production drills. The decision time is bounded by the rollback trigger criteria — and if the criteria are not pre-specified, the decision time is bounded only by when the incident bridge reaches consensus, which is a social process that is systematically slower under incident pressure than the threshold-detection that pre-specified criteria would provide. The structural mechanism that makes unspecified criteria produce long decision times is the sunk cost dynamic: once an investigation has been ongoing for 30 minutes, each additional 15 minutes of investigation appears cheaper than a rollback that would make the prior 30 minutes wasted; the investigation horizon keeps receding as the incident continues. Pre-specified trigger criteria eliminate the sunk cost dynamic by converting the rollback decision from a judgment call about whether the current situation warrants rollback into a measurement check: is the error rate above the threshold, has it been sustained for the trigger duration, and has the deployment team provided a confirmed root cause? If yes, rollback. If no, continue. The on-call engineer does not need to convince the deployment team that rollback is warranted; the deployment team's only lever is demonstrating a confirmed, self-resolving root cause within the specified trigger window. This converts a negotiation under pressure into a pre-specified protocol. Connect this property to the incident response playbook decision record: the rollback trigger criteria must be specified in the incident response playbook at the same level of detail as the escalation criteria — who initiates the rollback call, what the concurrence requirement is, what the maximum investigation time is, and what the override path is for a deployment team that believes the forward-fix is closer than the trigger window allows; a playbook that specifies escalation in detail but leaves rollback triggers unspecified has specified the communication side of incident response but not the recovery side.
Property 3: The schema-application version pairing and the partial-deployment rollback boundary. Blue-green and canary deployments introduce a deployment state that does not exist in full-traffic deployments: a period during which the database schema is at the post-migration state and the application traffic is split between old-version instances and new-version instances. This state is intentional and, during normal deployment, temporary — the traffic split exists only during the promotion ramp, and once promotion completes, all application instances are on the new version and the schema-application version pairing is clean. But rollback in this state — routing traffic back to the old version — leaves the database at the post-migration schema and the application at the pre-migration version. The behavioral consequence of this pairing depends on what the migration changed: additive migrations that added columns the old code ignores are safe; migrations that added NOT NULL columns the old code does not write produce constraint violations on every new insert; migrations that changed column semantics produce silent calculation errors that old code applies to new column definitions. The rollback boundary decision — whether rollback means reverting only the traffic routing or reverting both the traffic routing and the schema migration — must be specified by the engineer writing the migration, who knows what the migration changed, not determined by the on-call engineer initiating rollback, who typically does not have the information required to evaluate the schema-application compatibility at incident time. Connect this property to the on-call handoff decision record: the deployment state at the start of an on-call shift — specifically, whether any migrations ran during the current shift that affect rollback availability for deployments made before handoff — is a piece of context that the outgoing engineer has and the incoming engineer needs; a handoff protocol that includes current deployment state and migration reversibility for any deployments in the prior 24 hours gives the incoming on-call engineer the schema-application compatibility information they need to evaluate rollback boundaries correctly if a regression is detected during their shift.
The deployment rollback ADR: five sections
Section 1: The migration reversibility classification. Begin the deployment rollback decision record by specifying the migration reversibility classification system and the requirement that every database migration be classified before deployment. The classification system has three tiers. Fully reversible: the migration is additive only — it adds columns with DEFAULT values that the old application can handle correctly, adds indices, or adds tables that the old application ignores; the rollback migration is trivial (DROP COLUMN, DROP INDEX, DROP TABLE) and requires no data recovery. Conditionally reversible: the migration makes a change that the old application cannot handle correctly without a compensating migration — adding a NOT NULL column without a DEFAULT that the old application's INSERT must include, changing column semantics that the old application relies on, modifying enum values or foreign key constraints; the deployment package must include a rollback migration written and tested before deployment, and the rollback path includes executing the rollback migration before traffic routing reversion. Forward-only: the migration involves data transformation, data deletion, or schema changes that cannot be reversed without data loss or data reconstruction that exceeds the incident recovery time constraint; the rollback path is not available, and the deployment's incident recovery plan must specify the forward-fix path and its estimated time. The classification must be specified by the engineer who wrote the migration, reviewed by a second engineer as part of the code review, and recorded in the deployment package manifest that the on-call engineer can read at rollback time. The migration reversibility classification does not need to be elaborate — a three-field annotation in the migration file header (reversibility: fully-reversible | conditionally-reversible | forward-only, rollback migration: path/to/rollback.sql or N/A, forward-fix estimate: N/A or estimated-minutes) captures the information the on-call engineer needs to evaluate the rollback boundary. Connect this section to the production change freeze decision record: the change freeze risk classification process should include an explicit check for forward-only migrations in the deployment package; a deployment containing a forward-only migration during a freeze period is a higher-risk deployment than the same code change without the migration, because the forward-only migration removes the rollback path for the deployment window; the risk classification should require explicit sign-off from the engineering lead for forward-only migrations deployed during freeze periods, with a forward-fix estimate documented before deployment proceeds.
Section 2: The rollback trigger criteria and decision authority. Specify the rollback trigger criteria that govern when the rollback procedure is initiated for deployment-induced incidents. The criteria must address four elements. First, the automatic rollback thresholds: the metric values that trigger automatic rollback without a manual decision — error rate above a specified multiple of the pre-deployment baseline sustained for a specified window (the 3x-for-15-minutes threshold is a reasonable starting point; adjust based on the product's normal error rate variance), latency P99 above a specified multiple of baseline, any customer-visible data integrity error regardless of scope. Automatic rollback thresholds should be implemented in the deployment pipeline's health check system so that the rollback is initiated without requiring a human to be on a bridge when the threshold is crossed. Second, the manual rollback trigger criteria: the conditions under which the on-call engineer initiates rollback when the automatic thresholds have not been crossed but the signal is ambiguous — error rate below the automatic threshold but above a specified lower threshold with no confirmed root cause, latency degradation concentrated on a specific endpoint that is customer-critical even at low volume, or any anomaly that the on-call engineer cannot explain within a specified investigation time. Third, the maximum investigation time: the time limit after which rollback is required unless the deployment team can demonstrate a confirmed root cause that is either self-resolving or patchable within a specified additional window. The maximum investigation time should be specified in the rollback trigger criteria document, not left to the judgment of the incident bridge. A 30-minute default with a 15-minute extension contingent on a confirmed root cause is a reasonable starting structure; the specific values should be calibrated to the product's normal deployment risk profile and customer SLA requirements. Fourth, the decision authority: who can initiate rollback, who must concur, and whether the deployment team's preference to continue investigating can override the maximum investigation time trigger. Connect this section to the incident response playbook decision record: the rollback trigger criteria must be in the incident response playbook, not only in the deployment runbook; on-call engineers access the incident response playbook during incidents, not the deployment runbook, and the rollback trigger criteria that are in the deployment runbook but not the incident playbook will not be applied during the incidents where they matter most.
Section 3: The partial-deployment rollback boundary specification. Specify the rollback boundary for each deployment type used by the engineering team: full-traffic deployment (rollback reverts the application binary; schema rollback is executed if the migration reversibility classification is fully-reversible or conditionally-reversible with a rollback migration in the deployment package), blue-green deployment (rollback reverts the traffic routing to the blue deployment; schema rollback must be evaluated against the migration reversibility classification for any migrations that ran before the rollback is initiated), and canary deployment (same as blue-green, with the additional consideration that the blue application may have been writing data against the post-migration schema during the canary period). The specification must answer three questions for each deployment type: What is the schema state after the traffic routing rollback completes? What application version is serving traffic after rollback completes? Are these two states compatible — does the application version that will serve traffic after rollback produce correct behavior against the schema state that will exist after rollback? The compatibility evaluation is the migration engineer's responsibility, specified in the migration reversibility classification at migration write time. The on-call engineer's responsibility is reading the classification and executing the specified rollback sequence: for a conditionally-reversible migration, the sequence is (1) execute the rollback migration, (2) revert traffic routing to the previous application version, (3) verify that the application functions correctly against the rolled-back schema; for a forward-only migration, the sequence is (1) evaluate whether traffic routing rollback without schema rollback will leave the old application in a functional state, (2) if yes, revert traffic routing; if no, proceed to the forward-fix path. Connect this section to the on-call handoff decision record: the handoff context for any shift where a deployment occurred must include the migration reversibility classification for the deployed migration and the rollback boundary specification for the deployment type used; the incoming on-call engineer who receives this context can evaluate whether a regression detected during their shift is rollback-compatible without reading the migration code, which may not be accessible or interpretable during incident time.
Section 4: The rollback test protocol. Specify the frequency and scope of rollback tests in non-production environments that validate the rollback procedure before it is needed in production. The rollback test protocol has three tiers. Monthly traffic routing rollback drill: execute the full traffic routing rollback procedure in the staging environment — route traffic from the green deployment to the blue deployment using the production rollback runbook, confirm that the blue application is functioning correctly against the schema state, and measure the rollback execution time. The monthly drill validates the mechanical rollback procedure and produces a reliable execution time estimate for use in the maximum investigation time specification. Quarterly schema-plus-routing rollback drill: execute a full rollback including a conditionally-reversible migration rollback — deploy a migration classified as conditionally-reversible with a pre-written rollback migration, execute the rollback migration, revert traffic routing, and confirm that the blue application functions correctly against the rolled-back schema. The quarterly drill validates the schema rollback procedure for the conditionally-reversible classification and confirms that the rollback migration produces the expected schema state. Annual forward-only migration recovery drill: select a historical forward-only migration and execute the forward-fix path against a production-data snapshot in a staging environment — implement the compensating migration, verify data integrity, and measure the recovery time. The annual drill validates that forward-only migrations have documented forward-fix paths that are executable within the incident recovery time constraint, and surfaces forward-only migrations that have forward-fix paths that are undocumented or impractically slow. The drill results should be recorded in the rollback decision record as evidence that the rollback procedure is validated for each deployment type and migration reversibility class. Connect this section to the build artifact provenance decision record: the rollback test protocol must include verification that the artifact identity of the blue deployment — the specific binary or container image that will serve traffic after rollback — is confirmed before the rollback drill, not assumed to be the same artifact that was deployed before the green deployment; build provenance that records the artifact identity at each deployment stage provides the verification input for the rollback drill and ensures that rollback reverts to a known-good, verified artifact rather than to whatever binary happened to be in the blue slot at rollback time.
Section 5: The post-rollback investigation requirement. Specify the mandatory post-rollback investigation process that converts each rollback event into an improvement to the deployment risk model. The investigation has four required outputs. First, the root cause identification for the deployment regression: the specific code change, configuration change, or data error that produced the regression, and the deployment review process gap that allowed it to reach production — missing test coverage for the affected code path, insufficient canary period for the regression to manifest in the 10% traffic sample, or a data dependency that was not verified before deployment. Second, the deployment process change that prevents the same regression class: the specific test, review step, or data verification that would have caught the regression before deployment, and the deployment checklist item or pipeline gate that enforces the check going forward. Third, the rollback trigger criteria validation: whether the rollback trigger criteria specified in Section 2 caused the rollback to be initiated at the right time — too early (rollback triggered before the deployment team had a reasonable opportunity to identify a self-resolving root cause), appropriate (rollback triggered at approximately the time the maximum investigation time expired), or too late (rollback was appropriate earlier than it was initiated, indicating that the trigger threshold or maximum investigation time should be tightened). The validation output is a proposed update to the rollback trigger criteria in Section 2, reviewed in the post-incident retrospective. Fourth, the rollback boundary adequacy evaluation: whether the rollback boundary specification for the deployment type correctly predicted the schema-application compatibility after rollback, and whether any unexpected interactions between the rolled-back application and the post-migration schema were discovered during or after rollback. The adequacy evaluation output is a proposed update to the rollback boundary specification in Section 3 if any gap was discovered. The post-rollback investigation outputs are the mechanism through which the rollback decision record improves over time — each rollback event is evidence about where the trigger criteria are miscalibrated, where the reversibility classification is incomplete, or where the boundary specification does not account for a schema-application interaction that the migration engineer had not anticipated. Connect this section to the postmortem action item ownership decision record: the four investigation outputs — root cause, process change, trigger criteria update, boundary specification update — are postmortem action items with the same ownership and completion rate challenges as all postmortem action items; specifying that rollback investigation outputs are classified as High priority with a 14-day completion target and assigned to named individuals rather than teams at the close of the post-rollback retrospective ensures that the rollback decision record actually improves after each rollback event rather than accumulating findings that are documented once and never executed.
FAQ
Why is rollback sometimes not possible after a database migration?
A database migration is not reversible when the schema change it makes is incompatible with the previous application version — the application code that ran before the migration cannot function correctly against the schema that exists after the migration runs. The most common forward-only migration pattern is adding a NOT NULL column without a DEFAULT that the old application's INSERT must include: after rollback, every new INSERT from the old application is rejected by the NOT NULL constraint because the old code does not know to provide a value for the new column. The second pattern is removing a column the previous application version reads: the previous application throws errors on any code path that reads the dropped column. The third pattern is changing column semantics — renaming a column, changing its type, or modifying enum values — where the previous application reads or writes the column with the old interpretation, producing constraint violations or silent data corruption against the new schema. Migration reversibility is determined at migration write time, not at rollback time: the engineer writing the migration knows whether the schema change is forward-only and can write a rollback migration before deployment if the change is conditionally reversible. The on-call engineer initiating rollback at incident time cannot make a forward-only migration reversible — they can only choose between a forward-fix path and a traffic routing rollback that operates against the post-migration schema, which may produce a new failure mode if the old application code cannot handle the new schema correctly.
What should trigger an automatic rollback versus a manual rollback decision?
Automatic rollback should be triggered by metric thresholds that are unambiguously connected to the deployment and indicate customer-visible impact without requiring investigation: error rate on the new deployment above 3–5x pre-deployment baseline sustained for 10–15 minutes, latency P99 above 2x baseline for the same window, or any customer-visible data integrity error regardless of scope. These should be implemented as pipeline health checks that initiate rollback without a human decision. Manual rollback decisions apply when the signal is ambiguous — error rate elevated but within a range seen previously without a deployment, or the errors are isolated to a single client or endpoint — and the deployment team believes a root cause is identifiable within a specified window. The manual decision should still be governed by pre-specified criteria: a maximum investigation time (30 minutes is a reasonable default), a second engineer concurrence requirement to continue past the maximum, and an explicit rule that the deployment team's preference to investigate cannot override a sustained threshold breach. The maximum investigation time specification is the most consequential element: without it, the incident bridge will investigate beyond the point where rollback is correct, because each investigative step appears to make progress while the cumulative time cost of the investigation already exceeds what rollback would have cost.
How do you handle rollback during a blue-green or canary deployment?
Rollback during a blue-green or canary deployment has two components that must be evaluated together: the application traffic routing rollback (routing traffic back to the old version) and the database schema rollback (reverting the migration). Schema rollback is only possible if the migration was classified as conditionally reversible and a rollback migration was written before deployment. If the migration was forward-only, the rollback boundary is limited to traffic routing — returning old-version application code to a new-version schema — and the on-call engineer must evaluate whether the old application will produce correct, degraded, or broken behavior against the new schema before initiating traffic routing rollback. The evaluation: does the old code read any column the migration removed or renamed? Does the old code insert rows into any table where the migration added a NOT NULL column the old code does not populate? Does the old code rely on enum values or foreign key relationships the migration changed? If any answer is yes, traffic routing rollback without schema rollback will produce a new failure mode. In that case, the forward-fix path — patching the new version to resolve the regression while leaving the schema in the post-migration state — is the correct recovery path even if it takes longer than the rollback execution time.
What should a deployment rollback decision record specify?
Five specifications. First, the migration reversibility classification system and the requirement that every migration be classified before deployment: fully reversible (additive changes, trivial rollback migration), conditionally reversible (rollback migration written and included in the deployment package), or forward-only (no rollback path, forward-fix estimate documented). Second, the rollback trigger criteria: automatic rollback thresholds implemented as pipeline health checks, manual rollback trigger criteria for ambiguous signals, maximum investigation time before rollback is required, and decision authority (who initiates, who must concur, whether deployment team preference can override the trigger). Third, the rollback boundary specification for each deployment type: the sequence for executing schema rollback and traffic routing rollback for conditionally-reversible migrations, the compatibility evaluation for old-code-against-new-schema when schema rollback is not available, and the forward-fix trigger for forward-only migrations where traffic routing rollback is not viable. Fourth, the rollback test protocol: monthly traffic routing drill, quarterly schema-plus-routing drill, and annual forward-only recovery drill, with execution time and success criteria recorded in the decision record. Fifth, the post-rollback investigation requirement: root cause identification, deployment process change to prevent recurrence, trigger criteria validation, and rollback boundary adequacy evaluation — four action items with named owners and completion targets specified at the close of each rollback retrospective.
Further reading
- Production change freeze decision record — the change freeze decision record and the rollback decision record are complementary controls for the production change pipeline: the freeze decision record specifies which deployments are allowed during high-risk periods and what the risk classification threshold is; the rollback decision record specifies how to recover from deployments that are allowed through and produce regressions; the two decision records intersect in two places — the migration reversibility classification should be an explicit input to the change freeze risk classification (a forward-only migration during a freeze period is a higher-risk deployment than the same code change without the migration), and the forward-only migration forward-fix estimate should be evaluated against the freeze period's tolerance for extended incident recovery time before the deployment is approved; teams that have the change freeze decision record but not the rollback decision record have specified the prevention side of deployment risk but not the recovery side, which leaves the recovery time for deployments that pass the freeze risk threshold unbounded.
- Incident response playbook decision record — the rollback trigger criteria are an incident response decision that has the same structural importance as the escalation criteria and the communication templates — a decision that is made once at calm time and applied automatically at incident time, rather than remade under incident pressure each time a deployment-induced incident occurs; the incident response playbook that specifies escalation in detail but leaves rollback triggers unspecified has specified how the team communicates about incidents but not how the team recovers from the most common class of deployment-induced incidents; the JIT access procedure for production systems during an incident is a closely related specification — an engineer initiating rollback who does not have standing access to the load balancer or the deployment pipeline will need JIT approval at the moment when rollback execution time is the incident recovery metric, and the JIT approval latency must be shorter than the rollback execution time or the trigger criteria measurement is meaningless.
- Build artifact provenance decision record — rollback to a known-good previous version requires that the known-good version is identifiable and that the artifact in the blue slot is the same artifact that was deployed before the current green deployment; build provenance that records the artifact identity — the content hash of the binary or container image, the build inputs, and the pipeline run that produced it — provides the integrity guarantee that rollback depends on; without artifact provenance, the on-call engineer initiating rollback must trust that the blue slot contains the expected previous version, which is an assumption that has failed when deployment automation has rotated the blue slot contents, when a prior rollback was not properly recorded, or when infrastructure changes have altered the artifact in place; the rollback test protocol that validates the blue artifact identity before the monthly drill is an exercise in provenance verification that confirms the rollback assumption before it is needed under incident pressure.
- On-call handoff decision record — the deployment state at the start of an on-call shift — specifically, whether any forward-only or conditionally-reversible migrations ran during the prior shift, and what the rollback boundary specification is for any deployments made in the prior 24 hours — is a piece of context that the outgoing engineer has and the incoming engineer needs; a handoff protocol that includes current deployment state and migration reversibility classification for recent deployments gives the incoming on-call engineer the schema-application compatibility information they need to evaluate rollback boundaries correctly if a regression is detected during their shift; without this context, the incoming engineer must read the migration code to determine whether traffic routing rollback is safe against the current schema, which is an analysis that requires understanding the migration's intent and the old application's behavior — information that the outgoing engineer has from the deployment and that the incoming engineer must reconstruct under incident pressure.
- Open-source extractor — find the deployment rollback decisions buried in your AI chat history: the engineering all-hands where someone asked "what happens if we need to roll back a migration?" and the CTO said "we'd figure it out at the time" and moved on; the incident postmortem where the root cause was a forward-only migration that made rollback impossible and the action item was "write rollback migrations for all future schema changes" and the item sat open for four months; the pre-deployment planning session where a senior engineer said "this migration is NOT NULL without a default — if we need to roll back we'll have insert failures" and the deployment was approved with a note that "we'll monitor closely" instead of with a documented forward-fix path; and the quarterly deployment retrospective where someone observed that rollback took 3 minutes but the decision to roll back took 2 hours and no one could explain why the trigger criteria had not been specified — a meeting where the right question was asked and the answer was never written into a document; recovering these decisions from your AI chat history makes your rollback strategy a deliberate set of specifications with documented rationale rather than an accumulation of assumptions that only becomes visible when a deployment-induced incident requires rollback and the recovery path is different from what the team assumed it would be.