All field notes

How to Calculate API Uptime and Pay SLA Service Credits

12 failed minutes equal 99.9722% uptime in a 30-day month. Calculate exclusions, credit bands, and the exact SLA amount owed with a clear monthly receipt.

On this page10 sections
  1. 01Should uptime use requests or time intervals?
  2. 02What belongs in the uptime denominator?
  3. 03What makes an API interval successful?
  4. 04How should partial API failures be counted?
  5. 05How should SLA exclusions be recorded?
  6. 06How do you convert missed uptime into money?
  7. 07How should the plan price reflect the guarantee?
  8. 08What should the SLA system automate?
  9. 09What should a monthly SLA receipt include?
  10. 10SLA credit calculation FAQ

The outage is over. Engineering says it lasted 18 minutes. Support says customers were failing for 31. The status page shows 12 because that is when someone updated it. Finance cannot calculate the credit until somebody decides which clock counts.

This is why uptime guarantees fail after the incident. The contract contains a percentage, but the company never built the calculation that turns raw requests into a monthly result and a dollar amount.

The fix is to make the calculation mechanical. Define the unit, collect the evidence, apply exclusions without deleting history, map the result to a credit band, and retain a receipt the customer can understand.

Key takeaways

  • Choose time intervals or valid requests as the contractual denominator.
  • Keep observed uptime separate from credit-eligible uptime.
  • Multiply the affected fee by the published credit percentage.
  • Send customers a receipt that shows incidents, exclusions, and credit math.

The calculation begins with the definitions in your uptime SLA guarantee. Do not let the monthly report introduce rules that are absent from the signed contract.

Should uptime use requests or time intervals?

Time-based uptime gives every eligible interval one result, while request-based uptime gives every valid request one result. Choose the method that best represents the covered customer operation and record it in the contract.

Monthly uptime percentage =
  (eligible intervals - unavailable intervals)
  / eligible intervals
  * 100

If 12 of 43,200 one-minute intervals are unavailable, monthly uptime is approximately 99.9722%.

Request-based uptime uses valid requests:

Successful request percentage =
  successful valid requests
  / all valid requests
  * 100

If 80,000 of 10 million valid requests fail, success is 99.2%.

The two methods answer different questions. Time-based measurement gives a low-traffic customer credit for a real outage even if they happened not to send a request. Request-based measurement captures partial failures at scale but can let high-volume healthy routes dilute a complete failure on a critical low-volume route.

Many API providers should use time-based synthetic checks for the contractual commitment and request telemetry as supporting evidence. Whatever you choose, name it in the contract.

Microsoft's SLA monitoring guidance recommends endpoint monitoring, request tracing, dependency monitoring, timestamps, and aggregated reporting. Those records support either method.

What belongs in the uptime denominator?

The denominator should include every interval or request covered by the contract and remove only proven exclusions. Most disputes hide here. Decide whether it includes:

  • the full calendar month or only the period after activation;
  • scheduled maintenance;
  • invalid, unauthorized, or rate-limited requests;
  • test and sandbox traffic;
  • regions or endpoints the customer did not buy;
  • customer-requested suspension;
  • intervals with no traffic;
  • documented third party exclusions.

Never remove a period without retaining the reason and evidence. Keep two values: observed availability and credit-eligible availability. The first says what customers experienced. The second applies the contract.

AWS's Budgets SLA shows how specific these rules can be. It defines a one-minute interval, what "Unavailable" means, how a partial month is treated, which exclusions apply, what evidence a claim needs, and how credits map to availability bands.

What makes an API interval successful?

A successful interval should prove that the covered customer operation finished within its contractual limit. 200 OK is not enough for many APIs. A useful check can require:

  • a connection before the timeout;
  • an allowed status code;
  • a total response time below the contractual limit;
  • required fields in the response;
  • a safe business assertion, such as a known resource being readable;
  • successful authentication through the normal customer path.

Write endpoints need extra care. Do not create a real payment or order every minute. Use an idempotent synthetic transaction in a production-safe test account, provide a read-only verification route, or monitor the write path through controlled canaries.

If several probes disagree, define quorum. For example, an interval may be unavailable when probes from at least two of three regions fail twice within the minute. The second attempt reduces noise; it should not erase the first failure or stretch beyond the promised detection window.

For a practical monitoring and response loop, use the API uptime maintenance guide.

How should partial API failures be counted?

Calculate partial failures at the narrowest contractual boundary, usually the customer, covered service, and region. A platform-wide average can bury the exact endpoint failure a customer paid you to avoid.

Calculate availability at the narrowest contractual boundary: usually customer account, covered service, and region. For a bundle of critical endpoints, either require every critical operation to pass or assign explicit weights before the month begins.

Avoid subjective incident weighting. "Checkout was only partly down" is not a calculation. "Thirty-eight percent of valid checkout requests returned 5xx for 14 one-minute intervals" is.

If the contract uses an affected-customer ratio, define how unique customers are counted without exposing personal data. Cloudflare's Business SLA is one public example that incorporates affected visitors into its credit formula.

How should SLA exclusions be recorded?

Record exclusions in a separate ledger instead of deleting them from incident history. Suppose the monitor records 72 unavailable minutes:

  • 30 minutes from an application deploy;
  • 25 minutes from a named customer-owned identity system;
  • 12 minutes of announced maintenance that met the notice requirement;
  • 5 minutes from an unexplained network failure.

The observed availability is based on all 72 minutes. The credit-eligible calculation may remove the 25 customer-owned minutes and 12 maintenance minutes, leaving 35.

Do not exclude the unexplained five minutes. The provider bears the burden of proving an exclusion. "Probably the internet" is not evidence.

The monthly receipt should show each removed interval, clause, cause, and evidence reference. That makes an audit possible and discourages creative accounting.

The third party downtime framework explains how to classify customer-owned systems, required vendors, and shared-risk failures before applying this ledger.

How do you convert missed uptime into money?

Map credit-eligible uptime to the published band, then multiply the affected monthly fee by that percentage. For a plan promising 99.99%, one example is:

Credit-eligible monthly uptime Service credit
At least 99.99% 0%
99.0% to less than 99.99% 10%
95.0% to less than 99.0% 25%
Less than 95.0% 100%

If the affected monthly service fee is $4,000 and credit-eligible uptime is 99.919%, the customer receives a 10% credit:

$4,000 * 0.10 = $400

State what "affected monthly service fee" means. If the customer buys API access, implementation services, and a data package, the credit may apply only to the API line item. If only one region failed, say whether the base is the regional fee or the full plan fee.

Also state the cap, minimum credit, tax treatment, currency, and whether the remedy is a future invoice credit or a refund. Finance should be able to run the formula without asking engineering to interpret the contract.

How should the plan price reflect the guarantee?

Different guarantees justify different prices because they change delivery cost and credit exposure. An illustrative model might charge $1,000 per month for 95% uptime, $4,000 for 99.99%, and $10,000 for 99.999%.

Those are worked-example prices, not universal rates. They expose the shape of the promise:

Plan Price Downtime budget in 30 days Maximum credit in this example
95% $1,000 36 hours $250
99.99% $4,000 4 minutes, 19 seconds $4,000
99.999% $10,000 25.9 seconds $10,000

The premium pays for a smaller failure budget, stronger engineering, faster response, and more provider revenue at risk. A customer should not pay $10,000 for five nines if the contract excludes every required dependency and caps the remedy at $50.

See the full API uptime pricing model for architecture, support, and correlated-risk costs behind these example tiers.

What should the SLA system automate?

Automate the factual record: probe execution, timestamps, response validation, incident opening, recovery checks, and calculation inputs. PreMan can run scheduled endpoint probes, store recent results, calculate uptime and latency windows, manage alerts, and turn an incident into a coding-agent handoff.

Configure the probe from the signed SLA definition. Match the endpoint, timeout, expected status, and response assertion. Retain the contract version used for each month so a later edit does not rewrite old results.

Human review still matters for exclusions, force majeure, customer-caused events, and ambiguous partial failures. The goal is not to let software invent a legal outcome. It is to make every factual input reproducible.

PreMan helps an API provider stand behind an uptime guarantee. It does not pay the provider's credits or promise that upstream systems will never fail. The provider owns the financial commitment.

What should a monthly SLA receipt include?

A monthly receipt should let the customer reproduce the result without access to internal systems. Include:

  • covered service, account, plan, region, and month;
  • committed uptime and price;
  • total eligible intervals or requests;
  • unavailable intervals with incident IDs;
  • observed uptime;
  • every exclusion and its evidence;
  • credit-eligible uptime;
  • applicable credit band and calculation;
  • credit amount and invoice date;
  • dispute contact and deadline.

Send it automatically for premium plans, including months with no credit. Customers should not have to reverse-engineer your status page to learn whether the guarantee was met.

The receipt also improves pricing. If the $4,000 plan issues credits every quarter, it may be underpriced or overpromised. If the five-nines plan never approaches its budget but costs little more to operate, the team may have room to refine the offer. Evidence turns uptime from a sales adjective into a product with known cost and liability.

SLA credit calculation FAQ

How many minutes of downtime does 99.99% allow?

It allows 4.32 minutes in a 30-day month, or about 4 minutes 19 seconds. Five nines allows only 25.90 seconds. Microsoft publishes these values in its reliability target table, which also shows weekly and annual equivalents.

Can a status page calculate the service credit?

Not by itself. A status page communicates incidents, but the credit calculation also needs the covered account, eligible denominator, excluded intervals, affected fee, and credit band. Use status events as evidence inputs rather than treating the public incident duration as the complete invoice calculation.

Should excluded downtime disappear from the report?

No. Show observed uptime first, then list each exclusion and calculate credit-eligible uptime separately. This keeps the customer record honest while applying the contract. Every removed interval should name the clause, cause, timestamps, and evidence used to approve the exclusion.

→ Create an auditable uptime record with PreMan

Bring the loop to your API

Catch the regression. Open a verified fix.

Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.