How to Account for Third Party Downtime in an SLA
Two 99.99% serial dependencies yield about 99.98% availability. Learn how to attribute vendor downtime and write fair SLA credit rules without hiding outages.
On this page8 sections
- 01Which customer operation should the SLA measure?
- 02Which third party downtime policy should you choose?
- 03Why can't you inherit a vendor's uptime percentage?
- 04What evidence should attribute a dependency outage?
- 05How should dependency risk change the price?
- 06How do you calculate credits without hiding the outage?
- 07When should the API degrade instead of fail?
- 08Third party downtime FAQ
Your identity provider goes down for 47 minutes. The API is still accepting TCP connections, but every authenticated request fails. Your customer calls it downtime. Your infrastructure vendor calls it a third party incident. Your contract says dependencies are excluded.
That exclusion may protect the invoice. It will not protect the relationship.
Customers buy the service you assembled, not the individual uptime percentages of your vendors. If login, checkout, search, or data delivery stops working, they experience your product as unavailable. A sensible SLA can distinguish failures you control from failures you do not, but it should not turn the dependency graph into a list of excuses.
Key takeaways
- Measure whether the customer's operation worked before assigning blame.
- Two required 99.99% serial dependencies yield roughly 99.98% availability.
- Keep customer-experienced and credit-eligible uptime as separate figures.
Start with the exact availability definition in your uptime SLA guarantee. Dependency rules cannot repair a vague covered-service clause.
Which customer operation should the SLA measure?
Measure the complete customer operation before diagnosing its components. An order request may cross a CDN, API gateway, identity provider, application, database, payment processor, and webhook provider. A failure in any required link can break the transaction.
Label dependencies as one of three types:
- Required: the operation cannot succeed without it.
- Degradable: the operation can return a reduced but useful result.
- Asynchronous: the initial operation can succeed while later delivery waits.
This classification changes the SLA. If recommendations fail but checkout works, the commerce API may remain available. If the payment processor fails and there is no alternate path, order placement is down even if every server you own is healthy.
Measure the complete transaction from outside your stack. Internal component metrics help diagnose the cause, but the SLA should begin with the customer's result.
Which third party downtime policy should you choose?
Choose and document the allocation policy before an incident. There are four workable approaches.
Count all required dependencies
Any failure that breaks the covered operation counts against your SLA. This is the clearest customer promise and the hardest one to operate. It gives the provider a strong reason to buy redundant vendors, build fallback modes, and negotiate useful upstream remedies.
This approach fits premium plans. If a customer pays $10,000 per month for 99.999% availability, "our identity vendor was down" is unlikely to feel like a mission-critical service.
Exclude named dependencies
The SLA lists the specific vendor services that are outside the uptime calculation. This can work when the customer selected or directly contracts with the vendor, such as a customer-owned cloud account or identity tenant.
Keep the list narrow. "All third parties" could exclude nearly the entire production stack. Name the service, covered integration, evidence, and fallback behavior.
Share the failure
A contract can split the remedy when both parties accepted the dependency risk. For example, a direct platform outage may trigger a 25% credit while a named payment-processor outage triggers 10%.
The service was still unavailable, so the event remains in the availability report. Only the credit treatment changes. This is more honest than making a visible outage disappear from the denominator.
Offer two availability figures
Report customer-experienced availability and provider-controlled availability side by side.
Customer-experienced availability answers whether the operation worked. Provider-controlled availability removes contractually excluded causes. Credits use the second figure, but the first stays visible. This preserves the engineering truth while honoring the allocation negotiated in the contract.
Why can't you inherit a vendor's uptime percentage?
An upstream 99.99% commitment does not give your API 99.99% uptime. Microsoft notes that a composite SLO is the product of its contributing factors in its reliability target guidance. If two required independent services each deliver 99.99%, their combined theoretical availability is:
0.9999 * 0.9999 = 0.99980001, or about 99.98%
Add more serial dependencies and the number falls. The vendor's definition may also differ from yours. It may measure five-minute intervals, exclude regional failures, require a claim, or count only server-side errors. The provider may offer a credit even though your customer lost far more revenue.
Google's AppSheet SLA, for example, excludes performance issues caused by customer or third party equipment outside Google's primary control. Cloudflare's Business SLA has its own affected-customer formula, exclusions, and claim rules. These are useful references, not terms you automatically pass through to your customers.
Review each dependency for:
- its actual commitment and measurement method;
- service-credit schedule and cap;
- claim deadline and required evidence;
- regional and product exclusions;
- incident-history access;
- exit, fallback, or second-vendor options.
The upstream credit is a recovery source for you. It does not have to equal the credit you owe your customer.
The practical response is architectural as well as contractual. The API uptime maintenance guide covers timeouts, bounded retries, redundant capacity, and failure testing.
What evidence should attribute a dependency outage?
Attribution requires an external transaction result plus correlated dependency evidence. During a partial outage, an identity provider may return intermittent 503s while your retry logic turns them into timeouts. The original cause and customer-visible duration are not always the same.
The SLA should say how attribution works:
- The external transaction check establishes whether the service was unavailable.
- Correlated request IDs, response codes, and dependency timing identify the initiating failure.
- The excluded period ends when the dependency and your service can complete the transaction again.
- Extra downtime caused by your slow recovery, bad retry policy, or stale circuit breaker remains your responsibility.
Do not rely only on the vendor's status page. Status pages can lag, group incidents broadly, or use a different service boundary. Retain your own timestamped request evidence.
PreMan's scheduled endpoint probes, result history, latency metrics, alerts, and incidents can provide the outside-in record. Pair those checks with application logs and dependency telemetry for attribution. PreMan can show when the API failed and package the incident for repair; it cannot decide the legal allocation on its own.
How should dependency risk change the price?
Higher uptime should cost more because someone pays for redundancy and assumes the service-credit risk. Consider this illustrative pricing model:
| Plan | Price per month | Commitment | Dependency treatment |
|---|---|---|---|
| Standard | $1,000 | 95% | Named external providers excluded from credits |
| High availability | $4,000 | 99.99% | Required dependencies count, except customer-selected systems |
| Mission critical | $10,000 | 99.999% | Required dependencies count; alternate providers or degraded modes required |
These figures are an example, not fixed market pricing. The commercial logic matters. A five-nines plan cannot depend on one vendor with no fallback and then exclude that vendor from the contract. The customer would be paying for a number that disappears at the first interesting failure.
For a required vendor, price the cost of secondary capacity, data replication, failover testing, support, and credits into the premium plan. If redundancy is impractical, lower the commitment or negotiate a specific shared-risk clause.
For a full cost breakdown, see how to price 95%, 99.99%, and 99.999% API uptime.
How do you calculate credits without hiding the outage?
Keep observed uptime and credit-eligible uptime in the same receipt. Suppose the High Availability plan costs $4,000 per month and promises 99.99%. The API has 20 direct downtime minutes and 40 caused by a required identity provider. The total produces roughly 99.861% uptime in a 30-day month.
If required dependencies count, the full hour goes into the credit band. Under a sample 10% band for availability below 99.99% but at least 99%, the customer receives $400.
If the identity provider is a named exclusion, publish both numbers:
- Customer-experienced uptime: approximately 99.861%.
- Credit-eligible uptime: approximately 99.954% after removing the 40 excluded minutes.
The second number still misses 99.99%, so the customer receives the same credit in this example. If the remaining 20 minutes happened to fit within the contracted budget, the report would show no credit due while still acknowledging the hour customers experienced.
The SLA service-credit calculation guide provides the full denominator, exclusion ledger, and invoice calculation.
When should the API degrade instead of fail?
A degraded mode is useful when it preserves the customer's essential operation without creating a larger correctness or safety risk. Cache safe reads when a catalog provider fails. Queue non-urgent webhooks. Let existing sessions continue briefly during an identity outage when the risk model allows it.
Each degraded mode needs a product decision and a test. Stale data may be acceptable for a restaurant menu and dangerous for a brokerage balance. Queuing a notification is safe; queuing a time-sensitive trade may not be.
The SLA should describe whether the degraded operation counts as available. The monitor should exercise that definition. Once those agree, third party downtime stops being an argument about blame and becomes a known risk with a price, an owner, and a response plan.
Third party downtime FAQ
Who pays the customer when a cloud provider goes down?
Your customer contract decides who pays. If your SLA counts the cloud dependency, you owe the stated credit even if the vendor rejects your upstream claim. Treat any vendor credit as a separate recovery source. Do not promise a pass-through remedy unless the timing, amount, and exclusions truly match.
Do multiple vendor SLAs compound?
Yes, for required serial dependencies. Multiply the component availability figures rather than taking the lowest one. Two independent services at 99.99% produce about 99.98% composite availability. Microsoft documents this product-distribution method in its reliability target guidance.
Can customer-owned systems be excluded?
Yes, if the contract names the customer-owned account or system and defines the evidence required. Keep the customer-visible outage in the incident report, then remove only the proven excluded interval from the credit calculation. This prevents the financial rule from rewriting what users actually experienced.
Bring the loop to your API
Catch the regression. Open a verified fix.
Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.