All field notes

How to Write an Uptime SLA Guarantee That Survives an Outage

99.99% uptime leaves 4 minutes 19 seconds of monthly downtime. Write an API SLA with exact measurement, exclusions, claims, and credits that pay on time.

On this page9 sections
  1. 01What service should an uptime SLA cover?
  2. 02How should an SLA define available and unavailable?
  3. 03Which uptime formula and sampling interval should you use?
  4. 04How high should the uptime commitment be?
  5. 05What should the service-credit schedule include?
  6. 06Which SLA exclusions are defensible?
  7. 07How should the SLA claim process work?
  8. 08What belongs in the final SLA checklist?
  9. 09Uptime SLA FAQ

The bad SLA usually looks fine on paper. Then launch day hits, the market-data feed starts wobbling, your support queue fills with confused brokers, and the provider still points to 99.9% uptime as if that settles the argument. It doesn't, because the number only means something if you know exactly how it was measured, sampled, and excluded.

That gap shows up fast in real estate platforms too. Listings, pricing, availability, and location lookups are time-sensitive, and a compliance report that averages over the wrong window can hide the outage that mattered most. The contract says one thing, the product tells a different story, and everyone starts debating the definition instead of the incident.

A usable uptime SLA settles those arguments before the outage. It names the covered service, defines a successful request, explains how downtime is counted, and says what the provider owes when the promise is missed.

This is an engineering and commercial guide, not legal advice. Counsel should review the final agreement for your jurisdiction and risk profile.

Key takeaways

  • Name the exact production service, customer, plan, and region.
  • Define failure using an executable customer-path check.
  • State the credit bands, claims process, and exclusions before an outage.

This article covers the contract. Use the companion API uptime operations guide to build the system that has to meet it.

What service should an uptime SLA cover?

An uptime SLA should cover one precisely named production service. "The platform" is too vague. List the API, region, plan, and customer account covered by the commitment.

For example:

The Covered Service is the production Orders API available at https://api.example.com/v1/orders for accounts subscribed to the High Availability plan. Sandbox endpoints, beta features, the dashboard, and customer-hosted components are not covered.

That sentence prevents a failure in an unrelated analytics page from becoming an API SLA claim. It also prevents the provider from quietly excluding a broken production endpoint after the fact.

If several endpoints have different business importance, give them separate commitments. A read-only catalog can tolerate more downtime than an order-creation route. Blending both into one average can make the headline number look healthy while the revenue path is down.

How should an SLA define available and unavailable?

Define availability using a condition that a monitor can execute from the customer path. Microsoft describes availability as uptime from the customer's perspective in its reliability target guidance. The SLA should turn that principle into one exact test.

An API interval might be available when at least 99% of valid requests during that minute finish within two seconds and return an expected non-5xx response. A critical write endpoint may also require a response body containing a persisted transaction identifier.

Answer these questions in the definition:

  • Which HTTP statuses count as provider failures?
  • Do timeouts and connection failures count?
  • What latency threshold makes a response unusable?
  • Are malformed and unauthorized customer requests removed from the denominator?
  • Is availability based on requests, minutes, regions, or customer accounts?
  • Does a partial failure count when one critical endpoint is down?

Google's AppSheet SLA is a useful public example because it defines downtime using server-side error rate and gives an explicit monthly formula. You do not need to copy its thresholds. You do need that level of specificity.

Which uptime formula and sampling interval should you use?

A minute-based formula is easy to audit because every interval has one result:

Monthly uptime percentage =
  (eligible minutes - unavailable minutes) / eligible minutes * 100

"Eligible minutes" must be defined. If scheduled maintenance is excluded, subtract only maintenance that met the notice and duration rules. If the service launched halfway through the month, say whether the earlier days are excluded rather than silently treating them as perfect uptime.

Sampling matters. If a probe runs every five minutes, a two-minute outage may never appear. If one probe from one region fails because its local network broke, the service may be healthy. Define probe frequency, locations, timeout, retry behavior, and how results are combined.

The cleanest setup is to use the same executable check for the contract and the monitoring system. PreMan's scheduled probes and endpoint metrics can supply that operating record. Configure the check to match the SLA's exact endpoint, authentication path, timeout, and expected result. The contract should not promise one test while the dashboard runs another.

For partial failures and excluded intervals, follow the separate service-credit calculation method. It keeps observed uptime distinct from credit-eligible uptime.

How high should the uptime commitment be?

Choose a commitment below the level the service can repeatedly deliver. In a 30-day month, 95% uptime permits 36 hours of downtime. Four nines permits about 4 minutes and 19 seconds. Five nines permits 25.9 seconds (Microsoft Azure Well-Architected Framework, accessed 2026).

Those targets require different systems and should carry different prices. One illustrative API pricing model could be:

Plan Monthly fee Commitment Monthly downtime budget
Standard $1,000 95% 36 hours
High availability $4,000 99.99% 4 minutes, 19 seconds
Mission critical $10,000 99.999% 25.9 seconds

These are example prices, not a survey of the market. The point is commercial: a customer buying five nines is also buying redundant infrastructure, rapid response, tighter change controls, and a larger financial remedy when you fail.

Do not sell five nines because it looks better in a proposal. Measure your last six to twelve months, include the failure modes you expect to cover, and leave room between normal performance and the contractual floor.

The API uptime cost guide explains why those three targets need different architecture, support, and pricing.

What should the service-credit schedule include?

A service-credit schedule needs four inputs: availability band, credit percentage, affected fee, and maximum liability. A reader should be able to calculate the amount without asking finance to interpret the clause.

For a 99.99% plan, a straightforward example is:

Measured monthly uptime Credit on the affected month's fee
At least 99.99% 0%
99.0% to less than 99.99% 10%
95.0% to less than 99.0% 25%
Less than 95.0% 100%

If the customer pays $4,000 per month and availability is 98.7%, the credit is $1,000. Say whether the credit is cash, a refund, or an offset against a future invoice. Cloud vendors often use future service credits and make them the exclusive remedy. AWS publishes that pattern in its service-level agreements.

For smaller customers, consider applying credits automatically. A claims maze saves money in the short term and tells customers that the guarantee was written to avoid payment.

Which SLA exclusions are defensible?

Defensible exclusions are narrow, named, and supported by evidence. They may include customer-caused configuration errors, force majeure events, and announced maintenance within a defined window. "Anything outside our control" invites a dispute over every dependency failure.

For each exclusion, define the evidence. If customer traffic exceeded a documented rate limit, retain the request counts and the limit active at the time. If scheduled maintenance is excluded, retain the notice and timestamps. If a third party failed, prove which requests depended on it and how the customer experienced the failure.

Do not exclude your entire hosting stack. You chose the cloud, database, DNS provider, and queue. The customer bought your API. Third party dependencies deserve their own allocation rules, not a blanket escape hatch.

Use the third party downtime guide to choose among pass-through, shared-risk, and customer-owned dependency treatment.

How should the SLA claim process work?

The claim process should state the submission window, accepted evidence, provider response deadline, and payment timing. Avoid requirements that only the provider can satisfy.

A fair process might allow a claim within 30 days of the affected month, accept the customer's request IDs and timestamps, require the provider to reconcile those against its monitor and incident log, and issue the credit on the next invoice. If the records disagree, name the escalation path.

PreMan can preserve the operational evidence: probe results, latency, incidents, alerts, and the repair handoff. That supports the guarantee, but it does not sign the contract or pay the credit. The API company remains responsible for the promise and compensation.

What belongs in the final SLA checklist?

Before signature, confirm that the SLA states every measurement and payment input:

  • the exact covered service and customers;
  • the availability formula and measurement window;
  • success, failure, timeout, and latency rules;
  • probe frequency, locations, and retries;
  • maintenance and third party treatment;
  • credit bands, credit base, cap, and payment method;
  • claim evidence and deadlines;
  • the contract version and effective date.

Then configure the monitor from the signed definition and run a controlled failure test. If the test cannot produce an incident record and a correct sample credit calculation, the SLA is not ready to sell.

Uptime SLA FAQ

What is the difference between an SLA, SLO, and SLI?

An SLI is the measured indicator, such as successful-request percentage. An SLO is the internal target for that indicator. An SLA is the customer contract and carries financial or legal consequences when missed. Microsoft separates these terms in its reliability target guidance.

Should latency count as downtime?

Yes, when a slow response is unusable for the covered operation. Define the latency threshold, percentile, measurement point, and interval in the SLA. A response that arrives after a trading or checkout deadline should not count as available merely because it eventually returns 200 OK.

Should service credits be automatic?

Automatic credits create the clearest guarantee because the provider already owns the monitor and billing record. If claims are required, give customers at least one practical evidence path and a clear deadline. AWS's Budgets SLA shows the more traditional claim-based approach and its required timestamps and logs.

→ Build the evidence behind your SLA with PreMan

Bring the loop to your API

Catch the regression. Open a verified fix.

Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.