All field notes

How to Maintain API Uptime When Every Minute Has a Price

99.99% uptime allows 4 minutes 19 seconds of downtime each month. Learn how to measure failures, protect the budget, and price an API SLA with evidence.

On this page7 sections
  1. 01How should you define API availability?
  2. 02How much downtime does each uptime target allow?
  3. 03Which failures should you remove first?
  4. 04How should incident response protect uptime?
  5. 05Why should an uptime guarantee include compensation?
  6. 06How should you review uptime each month?
  7. 07API uptime FAQ

The bad uptime plan usually looks fine in a dashboard. Then launch day hits, a database connection pool fills up, requests start timing out, and the status page remains green because the health check can still return 200 OK.

That is the uncomfortable part of API reliability: the server can be running while the product is unusable. A login endpoint that answers in 18 seconds is technically available. A checkout endpoint that returns a successful HTTP response before dropping the order is technically available. Your customer will not care about either technicality.

Maintaining API uptime starts with a definition that matches what the customer is trying to do. It ends with money. If you sell an uptime guarantee and miss it, the customer should receive a service credit without having to argue with your support team over whose graph is correct.

Key takeaways

  • Measure the customer operation, not whether a process is running.
  • A 99.99% monthly SLA leaves about 4 minutes 19 seconds of downtime.
  • Higher uptime tiers should fund redundancy, faster response, and service credits.

If the terminology is still fuzzy, start with what an API contract actually covers. An uptime promise is one measurable part of that broader contract.

How should you define API availability?

Define API availability from the customer's path. An internal process check answers, "Is the application alive?" An uptime check should answer, "Can a customer complete the promised operation?" Microsoft's SLA monitoring guidance likewise recommends endpoint monitoring, timestamped request traces, dependency checks, and aggregated availability reporting.

For a payments API, that may mean sending an authenticated request to a safe test endpoint and checking the response body. For a listings API, it may mean fetching one known listing and confirming that price, availability, and location fields are present. For an agent tool, it may mean completing the MCP handshake, listing tools, and calling a read-only fixture.

A useful check has five parts:

  1. It runs from outside the production application.
  2. It exercises the same network and authentication path a customer uses.
  3. It has a strict timeout.
  4. It validates the response, not only the status code.
  5. It records the exact start time, duration, result, and failure reason.

Do not quietly change that definition after an incident. The measurement rule belongs in the SLA and in the monitor configuration.

How much downtime does each uptime target allow?

Convert the percentage into a failure budget before selling it. A 30-day month contains 43,200 minutes, so the monthly budgets look like this:

Uptime commitment Maximum downtime in 30 days
95% 36 hours
99.9% 43 minutes, 12 seconds
99.99% 4 minutes, 19 seconds
99.999% 25.9 seconds

The difference between 99.99% and 99.999% is not a decorative nine. Five nines gives you less than half a minute. One slow failover can consume the month. Microsoft's reliability target guidance publishes the same monthly figures: 4.32 minutes at 99.99% and 25.90 seconds at 99.999%.

Track the remaining budget throughout the billing period. Google SRE's example error-budget policy treats the budget as a control on release pace, not a punishment. If a team has used 80% by the tenth day, freeze risky releases and repair the recurring failure.

Which failures should you remove first?

Remove failures that can take down the entire customer operation. APIs usually fail in ordinary ways: one availability zone disappears, a certificate expires, or a retry storm makes a small incident larger. The AWS Well-Architected Reliability Pillar centers the same themes: resilient architecture, controlled change, and proven recovery.

Start with the boring controls:

  • Run more than one application instance and prove that traffic moves when one disappears.
  • Put hard timeouts around database and third party calls.
  • Use bounded retries with jitter. Never retry an unsafe write unless the request is idempotent.
  • Keep connection pools below the real limits of the database and downstream services.
  • Roll out changes gradually and retain a tested rollback path.
  • Separate liveness from readiness so a sick instance stops receiving traffic without entering a restart loop.
  • Test certificate, DNS, secret, and token rotation before the old value expires.

Every control needs a failure test. A failover plan that has never survived a failed node is a diagram, not a capability.

For the contractual definition behind this operating checklist, use the separate SLA drafting guide.

How should incident response protect uptime?

Incident response should minimize detection and recovery time while preserving evidence. The operating loop must be short enough that nobody reconstructs context in the middle of an outage:

  1. A synthetic check fails.
  2. The monitor repeats the check to rule out a one-off network error.
  3. A sustained failure opens one incident instead of sending fifty duplicate alerts.
  4. The alert includes the endpoint, response, latency, recent changes, and prior failures.
  5. An owner mitigates first, then investigates.
  6. Recovery closes the incident only after successful checks from the customer path.

PreMan provides this measurement and response layer for saved API endpoints. It runs scheduled probes, calculates uptime and latency over the selected window, opens and resolves incidents, and can package a fired alert into a coding-agent fix task. That gives the team one chain of evidence from failed check to repair.

PreMan cannot make an unhealthy database healthy, and monitoring software is not an insurance policy. The contractual guarantee remains yours. PreMan makes the guarantee measurable and gives your team a faster path back to green.

Read what PreMan does before treating it as part of the contract. It supplies monitoring and response evidence; it does not assume the provider's financial liability.

Why should an uptime guarantee include compensation?

An SLA without a remedy is an aspiration. Financial compensation makes the promise enforceable and gives the provider a reason to price reliability risk. The remedy does not have to cover every dollar of customer loss, but failure should cost the provider something.

Service credits are the usual mechanism. AWS, for example, publishes SLAs in which credits rise as monthly availability falls; its AWS Budgets SLA moves from a 10% credit to 25%, then 100% below 95% availability. The exact schedule is less important than having one that is automatic, understandable, and tied to the affected service fees.

Here is a worked pricing model for an API company. These are illustrative commercial tiers, not market averages:

Plan Monthly price Uptime commitment Architecture and support expectation
Standard $1,000 95% Single-region service, business-hours response
High availability $4,000 99.99% Redundant capacity, on-call response, tested failover
Mission critical $10,000 99.999% Multi-region design, continuous response, strict change controls

The higher price pays for reserve capacity, engineering time, safer releases, support coverage, and the credit risk you take on. If the five-nines plan uses the same infrastructure and response process as the 95% plan, the extra $9,000 is not a reliability product. It is wishful pricing.

The worked API uptime pricing model shows how to cost those differences. The service-credit calculation guide turns a missed target into a dollar amount.

How should you review uptime each month?

Reconcile three records every month: monitor results, the incident log, and credits issued. Look for disagreements before a customer finds them.

Ask which endpoints consumed the failure budget, how long detection and recovery took, whether exclusions were applied, and whether the same failure has happened before. Then price the next contract using that evidence. A team that reliably delivers four nines can charge for four nines. A team that cannot should sell a lower commitment until the architecture catches up.

The guarantee becomes credible when measurement, response, and compensation agree. That is the job: know when the API failed, fix it quickly, and pay what the contract says when you miss.

API uptime FAQ

Can an API provider promise 100% uptime?

Yes, but the promise is a financial guarantee rather than proof that outages are impossible. Cloudflare's Business SLA promises 100% uptime and defines credits when that level is missed. A provider making the same promise needs an exact measurement rule, a credit formula, and enough financial capacity to honor it.

Does scheduled maintenance count as downtime?

Only the contract can answer that. If maintenance is excluded, define the notice period, maximum duration, affected services, and evidence required. Do not label emergency repair as scheduled after it begins. Keep observed availability visible even when a valid maintenance interval is removed from the credit calculation.

How often should an API uptime check run?

The interval must be shorter than the outage the SLA is meant to detect. A five-minute probe can miss a two-minute failure, while 99.999% allows only 25.9 seconds per 30-day month. Match probe frequency, timeout, and retry behavior to the contract instead of choosing a convenient dashboard default.

→ Monitor your API uptime with PreMan

Bring the loop to your API

Catch the regression. Open a verified fix.

Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.