The SLA design decision record: why the availability commitment you made determines your contractual exposure surface and your measurement methodology gap

The customer-facing availability commitment tier, maintenance exclusion policy, and SLA measurement methodology are decisions that are almost never made explicitly — they emerge from a competitor's marketing page copied into a sales deck, a contract template borrowed from legal, and an internal SLO dashboard that measures something different from what the SLA contract says. Three failure patterns: the team whose "99.9% uptime" commitment was set without modeling the support cost structure that enterprise SLA enforcement requires, discovering the mismatch twelve months and $47,000 in credit obligations later; the team whose maintenance exclusion clause created 28 window-classification disputes in a single year, producing $180,000 in potential credit liability the finance team discovered during contract renewal; and the team whose internal SLO showed 99.92% while a major customer's annual audit showed 99.71%, because each was measuring something different.

A 33-person B2B SaaS company built a project planning and resource allocation platform for professional services firms — a tool that managed client engagement timelines, tracked consultant utilization rates, generated project budget forecasts, and produced the weekly status reports that account managers sent to clients. The platform had grown from a small-team product to one where enterprise customers depended on it for client-facing deliverables produced under contract with their own clients. When the sales team began closing deals with enterprise consulting firms and managed services providers, the prospects began asking about uptime guarantees. The sales team reviewed a dozen competitor websites, found that most listed "99.9% uptime SLA" in their enterprise tier feature tables, and included the same language in the enterprise pricing deck. Two enterprise deals were closed under contracts that referenced the deck's 99.9% commitment. Three more enterprise deals were closed over the next six months with similar SLA language embedded in procurement-drafted contracts.

The contracts specified different credit structures, negotiated individually by each customer's procurement team. The first two contracts offered a 10% monthly fee credit for each full calendar month where uptime dropped below 99.9%. The third contract offered a 25% monthly credit for any month below 99.9%, and a full month's credit for any month below 99%. The fourth contract offered a flat $5,000 credit per incident causing more than 4 hours of downtime. The fifth contract did not define an uptime measurement methodology, stating only that availability would be "measured in good faith by both parties." The sales team had signed all five contracts under a 99.9% uptime SLA without reviewing the credit clause differences, without modeling the maximum credit exposure under each structure, and without establishing an internal procedure for tracking uptime against the commitment.

Twelve months after the first enterprise contract was signed, the finance team conducted a contract obligation audit during annual renewal preparation. They found three things. First, the company had experienced 14 hours and 23 minutes of unplanned downtime across four incidents during the measurement period. The 99.9% SLA threshold allowed 8 hours and 45 minutes of downtime per year. They were in SLA breach on three of the five enterprise contracts — the two with monthly credit structures based on per-month availability had months where a single multi-hour incident placed that month below 99.9% availability even though the annual average was above the threshold. Second, none of the four incidents had been tracked against the SLA commitment in real time. The engineering team's post-incident reviews had calculated internal SLO impact, but no one had checked whether the incident duration placed any contract month below the contractual threshold. Third, the credit obligations required by the three contracts in breach totaled $47,000 — a number the finance team had not budgeted for and that the sales team had not disclosed to finance as a potential liability when the contracts were signed.

Beyond the credit obligations, the audit revealed a structural gap. The 99.9% commitment implied a support model that the company did not have: a customer-accessible status page with real-time incident updates, a process for notifying enterprise customers during incidents, a designated escalation contact for enterprise accounts during outages, and a procedure for calculating and issuing credits at period end. None of these existed. The engineering on-call rotation had no process for notifying enterprise accounts during incidents — the on-call engineer's job was to resolve the incident, not to communicate its status to customers. The status page had been set up informally and was updated inconsistently. The credit calculation process had never been defined, which meant that the first credit request from a customer would be processed ad hoc under the pressure of the customer relationship rather than through a documented procedure. The 99.9% uptime SLA had been committed in five enterprise contracts without any of the operational support structure that making that commitment mean something to an enterprise customer required. The decision had been made as a sales positioning decision — "match what competitors offer in their enterprise tier" — without being evaluated as an operational commitment that required corresponding operational capacity to fulfill.

A 28-person DevOps tooling platform built a continuous integration and deployment orchestration service for engineering teams at mid-market technology companies. The platform managed build pipelines, automated test orchestration, deployment promotion policies, and rollback procedures. Because the platform sat in the critical path between code commit and production deployment, enterprise customers placed substantial availability requirements on it — a CI/CD platform outage meant engineering teams could not deploy to production, which had direct consequences for customers' own release schedules. The sales team had committed to 99.95% monthly availability in enterprise contracts and had included a standard maintenance exclusion clause borrowed from a contract template that the company's outside legal counsel had used for a previous client.

The clause read: "Scheduled maintenance windows, announced at least 72 hours in advance via the service status page, are excluded from SLA availability calculations." The intent was to exclude the routine platform maintenance the infrastructure team conducted — database vacuuming, certificate renewals, Kubernetes version upgrades, infrastructure migrations. The engineering team conducted maintenance regularly: the infrastructure was complex, the Kubernetes version upgrade cadence was quarterly, and the team's practice was to perform maintenance during low-traffic windows at 2am on weekends. They conducted 47 maintenance windows during the first full year of enterprise customer operations. The maintenance windows averaged 38 minutes each and were announced via status page update, typically 72 to 96 hours in advance.

Eight months into the year, the infrastructure team performed an emergency upgrade to the PostgreSQL instance that backed the pipeline state database after discovering a performance regression that was causing build queues to back up during peak hours. The upgrade took 3 hours and 40 minutes, during which the platform was fully unavailable. The team posted a status page update 4 hours before the maintenance, below the 72-hour contractual threshold. The largest enterprise customer — a 900-person financial technology company that ran 200 to 300 pipeline runs per day and had a deployment freeze window that made the outage particularly disruptive — disputed the maintenance exclusion and requested SLA credit under the 99.95% commitment. Their contract was valued at $18,000 per month. A 4-hour outage in a 720-hour month represented 0.55% unavailability, well below the 0.05% threshold. The credit owed under the contract terms, treating the outage as non-excluded downtime, was $9,000 — 50% of monthly fee for a month below 99.9% availability, as specified in the tiered credit clause their procurement team had negotiated.

The dispute triggered a legal review of the maintenance exclusion clause. The review produced two findings. First, the clause as written did not define "scheduled maintenance" with enough specificity to distinguish it from "emergency maintenance" — the clause required 72 hours of advance notice but said nothing about what circumstances, if any, exempted maintenance from the notice requirement. Emergency maintenance performed with 4 hours of notice did not meet the contractual definition, and the customer's position — that the exclusion did not apply — was legally defensible. Second, and more concerning: the legal team reviewed all 47 maintenance windows against the 72-hour notice requirement and found that 28 of them had been announced between 48 and 72 hours in advance, not 72 or more hours in advance as the clause required. The infrastructure team's actual practice was to post status page announcements roughly when they thought about it, not against a formal deadline tied to the contractual requirement. Twenty-eight maintenance windows had been conducted with insufficient notice to qualify as excluded under the strict contract reading. The credit exposure attached to those windows, assuming customers could demonstrate downtime during them, was calculated at up to $180,000 across four enterprise contracts — a number that had not been contemplated when the maintenance exclusion clause was copied from a template into the sales contracts.

The maintenance exclusion clause had been included as a standard protective measure, and from the sales team's perspective it was protective — it was designed to ensure that necessary maintenance didn't consume SLA credits. But the clause created a classification requirement — every maintenance window was either excluded or SLA-counted — and the operational practice for ensuring windows met the classification requirement had never been established. The infrastructure team conducted maintenance according to their engineering judgment about timing, not against a checklist that verified 72-hour notice, documented the maintenance scope, and confirmed the window qualified as excluded. The gap between what the clause said and what the operational practice was had been invisible until a customer with strong commercial motivation to dispute a classification applied the clause strictly to a year's worth of maintenance history.

A 41-person B2B SaaS company built a financial reporting and compliance documentation platform for accounting firms and corporate finance teams. The platform generated regulatory filings, audit trail exports, and management reporting packages that were used directly in client deliverables. Because the platform processed time-sensitive financial workflows with firm submission deadlines, enterprise customers negotiated SLA terms with meaningful credit structures. The company had committed to 99.9% monthly availability in enterprise contracts, a commitment the engineering team considered conservative given their internal SLO tracking, which showed 99.92% availability over the trailing twelve months, well above the 99.9% commitment.

The internal SLO was measured using synthetic HTTP checks that ran from a single monitoring region — an AWS us-east-1 availability zone — at five-minute intervals. The checks called the platform's main API endpoint and the document generation service's health endpoint, verifying HTTP 200 responses within a 10-second timeout. The engineering team calculated monthly availability as the percentage of checks returning 200 within timeout, excluding periods where the status page showed "investigating" or "identified" status on the theory that these periods represented known outages that the team was actively working to resolve, not new availability failures to be counted separately.

The company's largest enterprise customer was a 240-person accounting firm with offices in New York, London, Singapore, and Frankfurt. The firm used the platform for all client reporting workflows across all four offices. Their contract included a 99.9% monthly availability commitment with a 50% monthly credit for any month below 99.9% and a full monthly fee credit for any month below 99.5%. The monthly contract value was $22,000. At the annual contract renewal, the accounting firm's IT director requested a formal SLA compliance report for the previous twelve months as a condition of renewal at the current pricing tier. The engineering team produced a report showing 99.92% average monthly availability across the twelve-month period, with two months dipping to 99.87% and 99.84% respectively — both above the 99.9% monthly threshold, no credits owed.

The accounting firm's IT director provided the company's own monitoring report: 99.71% average availability across the twelve-month period, with four months below 99.9% and one month at 98.94%. The gap between the two reports was 0.21 percentage points on an annual basis but six-tenths of a percentage point in the worst month and more than one full percentage point in a fifth month that the company's internal report showed above the commitment threshold. The reconciliation effort that followed identified three sources of divergence. The first was measurement geography: the firm's monitoring ran from their Singapore and Frankfurt offices in addition to New York, and the firm's European and Asian users regularly experienced elevated response times and connection timeouts during three incidents where the US-East-1 health checks showed no failure — a CDN edge node degradation in Frankfurt, a routing anomaly affecting Asia-Pacific traffic, and a brief misconfiguration in the platform's EU-region reverse proxy that didn't affect US traffic. The company's synthetic checks from US-East-1 returned 200 for all three incidents. The firm's users in Frankfurt and Singapore experienced 40 to 90 minutes of complete unavailability during each.

The second source of divergence was the "investigating" status exclusion. In three incidents, the company's on-call engineer had set the status page to "investigating" within two to eight minutes of the incident start, and the engineering team's SLO calculation excluded the "investigating" window from the downtime count on the theory that the team was aware of and actively resolving the issue. The excluded "investigating" windows totaled 73 minutes across the three incidents. The firm's SLA contract said "unavailable to users" — not "undetected by the provider's monitoring" or "before a status page update." The 73 minutes of customer-facing unavailability counted as downtime under the contract language regardless of when the status page was updated. The third source of divergence was check granularity: the company's five-minute synthetic check interval meant that a 3-minute outage could be invisible to the monitoring system if it fell between check windows, while the firm's continuous monitoring from actual user workstations captured sub-minute availability events. Two of the four months the firm's report showed below threshold had been elevated to near-threshold by several such sub-minute events that fell between the company's five-minute check windows.

The credit obligation calculated using the firm's monitoring data was $71,500 for four months below the 99.9% threshold. The credit obligation calculated using the company's internal SLO data was zero — the internal measurement showed no months below threshold. The SLA contract did not specify which measurement methodology governed in case of a dispute. The contract renewal negotiation became a measurement methodology dispute, with the company's outside counsel arguing for internal measurement and the firm's counsel arguing for customer-side measurement, neither position clearly supported by the contract language because the SLA commitment had been made without specifying how availability would be measured. The decision to commit to 99.9% monthly availability had been made against an internal SLO that was calculated in ways the SLA contract did not recognize, and the gap between the two measurement methodologies had been invisible for twelve months — it was only revealed when a customer with the commercial motivation and internal monitoring capacity to audit the commitment applied the contract language rather than accepting the provider's measurement.

Structural properties set by the SLA design decision

Three structural properties are determined when an engineering team establishes — or fails to establish — an SLA design decision record: what availability commitment tier the team can reliably sustain and support at enterprise scale, what operational behaviors the maintenance exclusion clause defines and enforces, and whether the internal SLO measurement methodology produces numbers that are comparable to a customer-side measurement. None of these are labeled as decisions when the first enterprise contract is signed — they emerge from the availability percentage written in the sales deck, the contract template clause borrowed from legal, and the monitoring dashboard configuration that the engineering team built to track internal reliability. The consequences are also invisible until a customer with commercial motivation and their own monitoring data audits the commitment.

Property 1: The customer-facing availability commitment tier and the support cost surface. An availability commitment tier defines two costs simultaneously: the reliability investment required to stay above the threshold (the engineering and infrastructure cost of the reliability level implied by the commitment), and the support cost structure required to fulfill the commitment operationally (status page maintenance, customer notification procedures during incidents, credit calculation procedures, and the on-call escalation path for enterprise customers). A team that sets an availability commitment tier by matching a competitor's marketing page has modeled neither cost. The reliability cost may be within reach — the team may actually achieve the committed availability level — but the support cost structure required to fulfill the commitment at enterprise scale is the operational investment that makes the commitment mean something to an enterprise customer. Enterprise customers sign SLA contracts not primarily because they expect to receive credits — most enterprise customers prefer to avoid the commercial disruption of a credit dispute — but because the SLA commitment signals that the provider has built the operational infrastructure to take enterprise availability seriously: real-time status communication during incidents, a designated escalation path for high-severity issues, and a credible mechanism for accountability when the commitment is missed. A team that has committed to 99.9% enterprise availability without building the supporting operational structure has made a commercial promise without the operational foundation that makes the promise credible. The decisions never written down in the SLA domain frequently include the support cost modeling: the sales team that included "99.9% uptime SLA" in the enterprise pricing deck had a reasonable motivation (match competitive positioning) but never wrote down the operational requirements that commitment created or evaluated whether the company was prepared to fulfill them. The SLO and error budget decision record connects at the reliability investment layer: the internal SLO target and error budget policy determine the reliability investment the engineering team makes to stay above the SLA threshold; the SLA commitment tier should be set at a level the engineering team can stay above with its current reliability investment, plus a buffer that accounts for the fact that SLA breaches have commercial consequences where SLO misses have only internal accountability consequences.

Property 2: The maintenance exclusion policy and the contractual exposure surface. A maintenance exclusion policy defines the operational boundary between downtime that the SLA commitment covers and downtime that the team has reserved the right to conduct for necessary infrastructure maintenance. The boundary has four dimensions: advance notice requirement (how much notice, delivered to whom, through what channel), window scope definition (which services are affected, what the maximum duration is, what happens if the window extends beyond the announced time), emergency maintenance definition (what conditions exempt the team from the advance notice requirement, what constitutes a qualifying emergency, what post-hoc notification is required), and dispute resolution procedure (how window-classification disputes are resolved, what evidence each party produces, and what the resolution timeline is). An exclusion clause that does not define all four dimensions leaves each undefined dimension as a classification surface — a space where reasonable people applying the contract language can reach different conclusions about whether a specific window qualifies as excluded maintenance or SLA-counted downtime. The classification surface becomes a dispute surface when a customer has commercial motivation to challenge a window classification. The commercial motivation scales with contract value and the credit obligation attached to the disputed window — a 50% monthly credit on a $22,000 contract produces a $11,000 credit, which creates strong motivation for a customer's procurement team to audit the prior year's maintenance windows for classification disputes. The incident response playbook decision record connects at the emergency maintenance definition layer: the conditions that qualify as an emergency for the purposes of the maintenance exclusion should be the same conditions that trigger the incident response playbook's escalation procedure — a security incident, imminent data loss, or infrastructure failure that makes delayed maintenance worse than immediate maintenance; aligning the emergency maintenance definition with the incident response playbook's escalation triggers produces a consistent operational definition that the team applies in real time and can document in the contract as the applicable standard. The observability strategy decision record connects at the notice verification layer: the advance notice requirement in the maintenance exclusion clause must be enforced operationally, which means the team needs a mechanism to verify that a planned maintenance window was announced with the required advance time before the maintenance is conducted — a status page management workflow that includes a confirmation step, not an informal convention where engineers post announcements when they think about it.

Property 3: The SLA measurement methodology and the internal-external monitoring gap. An SLA commitment is measured simultaneously by the provider's internal monitoring systems and by the customer's own monitoring and user experience — and when these measurements diverge, the contract determines which measurement governs the credit calculation. A team that sets its SLA commitment tier based on internal SLO measurement without auditing whether the internal measurement would produce the same result as a customer-side measurement is committing to a number derived from a methodology the SLA contract may not recognize. The three most common measurement divergences are location (provider monitors from one geography, customer traffic originates from several), method (provider uses synthetic checks shallower than the customer's actual workload, customer measures availability from actual user transactions), and status classification (provider excludes periods classified as "investigating" from SLO calculations, customer's contract language counts all customer-facing unavailability regardless of provider status). Each divergence produces a situation where the provider's internal measurement shows a different result than the customer's measurement — and the first time the discrepancy is discovered is typically when a customer with commercial motivation to audit the commitment applies the contract language to their own monitoring data. The alerting threshold decision record connects at the monitoring alignment layer: the alerting system that detects incidents should be calibrated to detect the same availability failures that customer-side monitoring would detect — if the alert fires after 5 minutes of synthetic check failures but customer users experience availability failures within 30 seconds, the alerting system is not detecting availability failures at the granularity the SLA contract covers; aligning alert thresholds with the detection granularity that matters for SLA credit calculations ensures that incidents are detected and the status page is updated before the customer's monitoring has accumulated significant downtime that the provider is unaware of. The WhyChose extractor finds the SLA design decisions buried in your AI session history — the initial enterprise sales preparation session where someone said "let's match what Competitor X offers in their enterprise tier" and the 99.9% commitment was written into the deck, the post-renewal legal review session where the maintenance exclusion language was scrutinized for the first time, and the customer success session where an account manager discussed the measurement methodology dispute with legal and the resolution was never captured as a decision that should change the SLA design for future contracts.

The SLA design ADR: five sections

Section 1: Customer-facing availability commitment tier and calibration methodology. Specify the availability commitment tier offered in customer contracts, the methodology used to select the tier, and the reliability investment and support cost structure that sustaining the tier requires. The tier selection methodology should include: (1) the trailing twelve-month internal SLO measurement at the chosen measurement granularity (monthly, quarterly, or annual), with the commitment tier set below the trailing SLO at a margin that provides a reliability buffer accounting for incident frequency variance; (2) the engineering cost of the reliability investment required to sustain the tier — the on-call coverage, the incident response capacity, the infrastructure redundancy and failover capability, and the testing and change management processes required to maintain the reliability level; and (3) the support cost structure required to fulfill the commitment at enterprise scale — the status page maintenance process, the customer notification procedure during incidents, the credit calculation procedure, and the escalation path for enterprise customers during high-severity incidents. Document the tier decision as a product and commercial decision (what tiers are offered at what pricing), not just as a technical reliability target, because the operational consequences of the commitment are primarily commercial rather than technical. Connect this section to the SLO and error budget decision record: the internal SLO target should be set above the customer-facing SLA commitment tier, with the difference treated as the engineering team's buffer — the SLO is the internal accountability target, the SLA is the commercial consequence threshold, and the buffer is the margin that absorbs normal incident frequency variance without generating SLA breach events.

Section 2: Maintenance exclusion policy and operational enforcement procedure. Define the maintenance exclusion policy with four required components: the advance notice requirement (specific hours, delivery channels, and recipient definition), the window scope definition (affected services, maximum duration, and extension handling), the emergency maintenance definition (qualifying conditions with explicit examples of what constitutes and does not constitute an emergency, post-hoc notification requirement and timeline, and per-contract-period cap on emergency windows), and the dispute resolution procedure (evidence requirements, resolution timeline, and escalating resolution path from account manager to legal if needed). Establish an operational enforcement procedure for the advance notice requirement: a planned maintenance window should not be conducted until a team member has verified that the status page announcement was posted at least the required number of hours before the window start. This verification should be a checkpoint in the maintenance window runbook, not an informal convention. Document the classification decision for each maintenance window in a maintenance log: date, duration, advance notice timestamp, customer notification channel, classification (scheduled or emergency), and if emergency, the qualifying condition. The maintenance log provides the evidence base for classification disputes and also surfaces patterns — if emergency maintenance windows are occurring at high frequency, the emergency maintenance definition may need tightening or the infrastructure investment in preventing emergency situations may need increasing. Connect this section to the incident response playbook decision record for the emergency maintenance trigger conditions and the post-hoc notification procedure that the playbook should include as a required step for any incident-triggered maintenance.

Section 3: SLA measurement methodology and customer-side alignment audit. Specify the measurement methodology that governs SLA credit calculations — the monitoring source (synthetic checks, real user monitoring, or customer-side measurement), the measurement geography (which regions, corresponding to what fraction of customer traffic), the check frequency and granularity, and the status classification policy (what periods are excluded from downtime calculations and on what basis). Conduct a measurement alignment audit at least annually and before each major contract renewal: run the internal measurement methodology and an approximation of a customer-side methodology against the same trailing twelve months and compare the results. If the methodologies produce different results, identify the source of divergence (location, method, or status classification) and either align the internal methodology to match a customer-side measurement or explicitly document the difference in the contract and agree with the customer on which methodology governs. For new enterprise contracts, specify the measurement methodology in the SLA clause itself — including the monitoring source, the check frequency, the geographic measurement points, and the status classification policy — so that a customer with their own monitoring cannot produce a different availability number using a different methodology and claim the contract language supports their measurement. Connect this section to the alerting threshold decision record for the alignment between the alerting system's detection sensitivity and the SLA measurement granularity: if the SLA is measured at 5-minute granularity but the alerting system fires alerts after 15 minutes of consecutive failure, there is a window where customer-facing downtime is accumulating against the SLA threshold before the team is aware of the incident.

Section 4: SLA credit structure and maximum liability modeling. Specify the credit structure offered in enterprise contracts — the credit percentage per tier of availability shortfall, the credit cap per billing period, the credit claim process (submission window, evidence requirements, and provider response timeline), and the maximum potential liability under the credit structure across the enterprise customer base. Model the maximum credit liability scenario: if every enterprise customer simultaneously had a month at the worst-case availability level that triggers the maximum credit, what is the total credit obligation? Compare that number to monthly recurring enterprise revenue — the maximum credit obligation should be a fraction of monthly revenue that the business can absorb without financial disruption, typically between 50% and 100% of the highest-value enterprise contract's monthly fee. Credit structures where the maximum liability scenario exceeds this threshold should be renegotiated or the credit cap should be explicitly set in the contract. Document the credit claim process in an internal procedure before signing the first enterprise contract that includes a credit clause: who receives a customer credit claim, who calculates the credit using the agreed measurement methodology, who approves the credit, and what timeline the customer can expect for response and credit issuance. Connect this section to the SLO and error budget decision record for the relationship between error budget burn rate and credit obligation accrual: if the team's error budget burn rate is high enough to produce months below the SLA threshold with some regularity, the credit structure should be stress-tested at the expected breach frequency to model the annual credit obligation rather than treating the credit clause as a worst-case edge case.

Section 5: SLA breach detection and enterprise customer communication procedure. Specify the procedure for detecting SLA threshold breaches in real time and communicating with enterprise customers during and after incidents that affect or may affect the SLA commitment. SLA threshold breach detection requires tracking cumulative downtime against the monthly threshold in real time, not only through post-incident retrospective review. A monitoring dashboard should display the current month's accumulated downtime for each applicable measurement tier and highlight when cumulative downtime approaches or crosses SLA thresholds — this gives the engineering team and customer success team visibility into developing SLA risk before the period ends. The enterprise customer communication procedure should define three communication events: (1) the incident notification — when the status page is updated to "identified," enterprise customer designated contacts are notified via email or a dedicated communication channel with the incident scope, affected services, and estimated resolution time; (2) the incident update — periodic updates (typically every 30 to 60 minutes during active incidents) to designated contacts with current status and revised estimated resolution time; and (3) the post-incident SLA report — within a defined period after the incident (typically five business days), a written report to designated contacts documenting incident timeline, root cause, downtime duration against the SLA measurement methodology, whether the incident affects the monthly threshold, and what credit, if any, will be applied. The post-incident SLA report closes the credit determination loop before the customer requests it, which is almost always preferable to waiting for the customer to audit the period and submit a credit claim. Connect this section to the observability strategy decision record for the monitoring infrastructure that makes real-time SLA threshold tracking possible, and to the incident response playbook decision record for the enterprise customer notification steps that should be embedded in the incident response runbook so that customer communication happens as a standard incident response action rather than as an afterthought once the technical issue is resolved.

FAQ

What is the difference between an SLO and an SLA, and why does it matter for the SLA commitment tier?

An SLO (Service Level Objective) is an internal target that the engineering team uses to measure and improve reliability — it is a commitment from the engineering team to itself about how reliably the system should operate. An SLA (Service Level Agreement) is a contractual commitment from the company to a customer that defines consequences — typically service credits, termination rights, or remediation obligations — when the commitment is not met. The distinction matters for the commitment tier because SLOs and SLAs measure different things and have different consequences when they are missed. An SLO that is missed produces internal accountability and a reliability improvement task for the engineering team. An SLA that is missed produces a credit obligation, a customer conversation, and potential contract renegotiation. Teams that set SLA commitment tiers by matching their internal SLO target — or by copying a competitor's SLA marketing page — often set commitments without modeling the consequence structure attached to those commitments. The SLA commitment tier decision requires modeling both the engineering cost of maintaining the commitment (the reliability investment required to stay above the threshold, including on-call coverage and incident response capacity) and the commercial consequence of missing the commitment (credit calculation at the contracted rates, potential customer churn, and the operational overhead of the credit calculation and dispute resolution process). A 99.9% SLA commitment on an enterprise contract with a 10% monthly credit for each 0.1% below the threshold looks different from a 99.5% SLA commitment on the same contract — both in terms of the reliability investment required and the credit exposure if reliability degrades.

What should a maintenance exclusion clause include to avoid classification disputes?

A maintenance exclusion clause should define four things with enough specificity that any downtime event can be unambiguously classified as excluded or not excluded without requiring negotiation. First, the advance notice requirement: the specific amount of advance notice required for a window to qualify as scheduled maintenance (typically 48 to 72 hours), the channel through which the notice must be delivered (email to designated contacts, status page post, or both), and whether notice to the customer's designated technical contact is sufficient or whether notice must reach a specific individual. Second, the window scope definition: what the maintenance window covers — the specific services affected, the maximum duration of the window, and what the team will do if the maintenance extends beyond the announced duration (whether the extension is covered by the same exclusion or converts to SLA-covered downtime). Third, the emergency maintenance definition: the explicit conditions under which maintenance can be performed without the advance notice requirement — typically, active security incidents (with a definition of what constitutes a security incident), imminent data loss risk, or infrastructure failures where delayed maintenance would cause greater downtime than performing maintenance immediately. Emergency maintenance should require post-hoc notification within a specified period (typically 24 hours) and should be capped at a maximum number of windows per contract period or a maximum duration per window. Fourth, the dispute resolution procedure: who can request classification review of a disputed window, what evidence each party is expected to produce, and what the resolution timeline is. Maintenance exclusion clauses that omit any of these four elements leave the corresponding classification question open, and open classification questions become disputes at renewal time, particularly when the commercial stakes are high.

How should a team align internal SLO measurement with customer-side SLA measurement?

Internal SLO measurement and customer-side SLA measurement diverge for three categories of reasons: measurement location, measurement method, and status classification. Measurement location divergence occurs when the provider's synthetic health checks run from a single cloud region or availability zone while customer traffic originates from multiple geographies — a regional network event or CDN edge failure visible to customers in one geography may be invisible to health checks running in a different region. The mitigation is running synthetic checks from at least the same geographies as significant customer traffic concentrations, or contracting on the basis of customer-side measurement rather than provider-side measurement. Measurement method divergence occurs when the provider measures availability as "synthetic check success rate" and the customer measures availability as "transaction success rate from the customer's application" — a service can pass synthetic health checks while returning errors on a subset of transaction types, particularly if the synthetic check is shallower than the customer's actual workload. The mitigation is aligning the synthetic check depth with the transaction types the SLA is meant to cover, or using real user monitoring data rather than synthetic checks as the SLA measurement basis. Status classification divergence occurs when the provider excludes periods classified as "investigating" on their status page from SLO calculations — on the assumption that "investigating" indicates the team is aware of the issue and working on it, not that the issue is confirmed — while the customer's SLA contract language counts any period where the service is unavailable to users as downtime regardless of the provider's status page classification. The mitigation is aligning the status page classification policy with the SLA contract language before signing contracts that use customer-facing availability as the measurement basis, and auditing both measurement paths on at least a quarterly basis to verify they would produce the same result for recent incidents.

How should SLA credit structures be designed to align incentives without creating unbounded liability?

An SLA credit structure should accomplish three things: create a credible incentive for the provider to maintain the committed availability level, create a proportionate remedy for the customer when the commitment is missed, and create a bounded worst-case liability that the provider can model and plan for. A credit structure that accomplishes all three is based on a percentage of the customer's monthly recurring fee — typically 5% to 30% of the monthly fee per SLA period missed, up to a maximum of one month's fee per contract period or per billing period. This structure ties the credit to the contract value (larger contracts receive larger credits in absolute terms, which creates a stronger incentive for high-value customers), caps the worst-case liability at a predictable multiple of monthly revenue (typically 100% of one month's fee), and avoids the unbounded liability structure of per-incident flat credits or liquidated damages clauses. The credit tier structure — what availability levels trigger what credit percentages — should be calibrated against the team's actual reliability history and reliability investment capacity. A common structure is: above commitment, no credit; 1 to 2 sigma below commitment, 10% monthly credit; 2 to 3 sigma below commitment, 25% monthly credit; more than 3 sigma below commitment (or below a hard floor like 99.5% for a 99.9% commitment), 50% monthly credit. The credit structure should also specify the claim process: the customer submits a credit request within a defined window (typically 30 days of the SLA period end), with what evidence (monitoring data, incident timestamps), and the provider responds within a defined window (typically 30 days). Credit structures that omit the claim process leave the credit calculation procedure undefined, which means the first claim reveals the gap in the process under the pressure of a customer dispute rather than under the lower-pressure conditions of contract design.

Further reading

  • SLO and error budget decision record — the internal reliability target and error budget policy that the SLA commitment tier should be set above, with the gap as the engineering team's buffer.
  • Incident response playbook decision record — the enterprise customer notification procedure that should be embedded in the incident response runbook so communication happens as a standard incident action.
  • Observability strategy decision record — the monitoring infrastructure that makes real-time SLA threshold tracking and multi-geography synthetic check coverage possible.
  • Alerting threshold decision record — the detection sensitivity alignment between the alerting system and the SLA measurement granularity, so the team is aware of incidents before downtime accumulates silently against the SLA threshold.
  • Decisions never written down — the SLA commitment tier set by copying a competitor's marketing page and the maintenance exclusion policy copied from a contract template are canonical examples of decisions made without being labeled as decisions.
  • Capacity planning decision record — the reliability investment required to sustain an SLA commitment tier depends on the capacity model that predicts failure probability under peak load conditions.
  • Open-source extractor — find the SLA design discussions buried in your AI chat history: the enterprise sales preparation sessions, the post-renewal legal reviews, and the customer success conversations where measurement methodology disputes were resolved informally without being captured as decisions.