The cost of downtime: statistics, and how to calculate your own number
By the Erzon engineering team · Last updated July 13, 2026
TL;DR
- The most-quoted number in the industry is Gartner’s $5,600 per minute (roughly $336,000 per hour), published in 2014 and still cited everywhere. Gartner’s own analyst framed it as an average across a wide range, from about $140,000 to $540,000 per hour depending on the business.
- ITIC’s recurring hourly-cost surveys have consistently found that over nine in ten mid-size and large enterprises put a single hour of downtime above $300,000, with roughly four in ten reporting $1 million or more.
- The Uptime Institute’s 2022 Annual Outage Analysis found over 60 percent of outages cost more than $100,000 in total, up from 39 percent in 2019, and the share costing over $1 million rose to about 15 percent.
- Recovery speed varies enormously: DORA’s State of DevOps research puts elite teams at under an hour to restore service and low performers at between a week and a month.
- These are enterprise-weighted figures. For an SMB, the only honest number is the one you compute yourself, and this page walks through exactly how, with the formula and worked arithmetic by company size.
A note on the numbers
Downtime statistics are famously slippery. Vendor studies skew high (they sell reliability), self-reported surveys skew toward memorable incidents, and per-hour averages blend a Fortune 500 payment outage with a small SaaS blip. This page limits itself to figures that are real, widely published, and attributable: Gartner’s 2014 estimate, ITIC’s annual hourly-cost surveys, the Uptime Institute’s annual outage analyses, and Google’s SRE and DORA research. Where sources disagree, we give the range. Where a precise number would have to be invented, we describe the pattern instead. And because every one of these figures skews enterprise, the second half of the page is a framework for computing your own.
The headline figures
Gartner: $5,600 per minute. The single most-cited downtime statistic comes from a 2014 Gartner analysis pegging the average cost of IT downtime at $5,600 per minute, which works out to roughly $336,000 per hour. The same analysis stressed the spread: from about $140,000 to $540,000 per hour depending on industry, business model, and which system is down. The number is twelve years old and was an average even then. Treat it as an order-of-magnitude anchor for enterprises, not a prediction for your business.
ITIC: over $300,000 per hour for most enterprises. ITIC has run an hourly-cost-of-downtime survey for years, and the results are remarkably stable: in recent editions, over 90 percent of surveyed mid-size and large enterprises said one hour of downtime costs them more than $300,000, and roughly 40 percent put the figure at $1 million to $5 million or beyond. These are self-reported estimates, but the sample is large and the pattern has held across many annual editions.
Uptime Institute: the expensive tail is growing. The Uptime Institute’s Annual Outage Analysis is the best recurring source on outage frequency and severity. Its 2022 edition found that over 60 percent of outages resulted in total losses above $100,000, up sharply from 39 percent in 2019, and that the share of outages costing over $1 million rose to about 15 percent. Its surveys also consistently find that a large majority of operators, around four in five, experienced at least one outage in the previous three years. The direction across editions is clear: outages per site are slowly becoming less frequent, but each serious one costs more.
For SMBs, published averages mislead in both directions. Small-business downtime numbers floating around the web are mostly unattributed or recycled vendor marketing, so we will not repeat them. The honest approach is the calculation below. What we can say from the arithmetic itself: direct lost revenue is usually the smallest line item for a small company, and recovery labor plus churn usually dominate.
How long recovery actually takes
Cost per hour only matters multiplied by hours down, and that multiplier varies more between teams than the hourly cost does.
DORA’s benchmark bands. Google’s DORA State of DevOps research, the largest ongoing study of software delivery performance, groups teams by time to restore service: elite performers restore in under an hour, high performers within a day, and low performers take between a week and a month. The gap between the best and worst is not 2x, it is several hundred x.
Diagnosis is the long pole. In practice most of an incident is spent working out what is wrong, not fixing it. The eventual fix is very often small: a rollback, a config revert, a restarted process, one query killed. This is why the same outage costs one team 40 minutes and another team two days, and why observability and a rehearsed diagnostic process pay for themselves in the first real incident.
The clock has phases. Time to detect (did monitoring fire, or did a customer email?), time to engage (is someone on call and able to act?), time to diagnose, time to fix, time to verify. Teams that measure these separately usually discover that detection and engagement, the cheap parts to improve, are eating a surprising share of total downtime.
What actually causes outages
Change is the leading trigger. Google’s SRE literature famously reports that roughly 70 percent of outages are due to changes in a live system: deploys, configuration changes, migrations. Industry post-incident analyses agree. Your deploy pipeline is your biggest single reliability lever, and health-gated rollouts plus rehearsed rollbacks buy more uptime than exotic redundancy.
Human error is implicated in the majority. The Uptime Institute has repeatedly attributed the majority of outages, in whole or in part, to human error, which on inspection usually means process failure: the checklist that did not exist, the runbook nobody updated, the one person who knew being on a plane. Blaming the person changes nothing. The fix is reviews on risky changes, automation of the repetitive steps, and runbooks written before the incident.
Third-party and cloud dependencies are a rising cause. Recent Uptime Institute editions note a growing share of outages originating with external providers: cloud platforms, SaaS dependencies, CDNs, DNS. Your uptime ceiling is the uptime of your dependencies. Know which vendor failures take you down and which you can degrade around.
Capacity and housekeeping failures remain stubbornly common. Full disks, expired certificates, exhausted connection pools, and unmonitored queues appear in incident analyses year after year, precisely because they are nobody’s job until they fire. Trend alerts on disk, certificate expiry, and connection counts are the cheapest reliability wins available.
Calculate your own number
The formula is simple. Running it once, in calm conditions, is the single most useful thing you can take from this page.
Downtime cost = (hourly revenue through the affected system x hours down x impact fraction) + recovery labor + SLA credits and refunds + churn
Line by line:
1. Direct revenue. Take annual revenue that flows through the affected system and divide by the hours in which that revenue actually arrives (8,760 for an always-on SaaS, closer to 4,000 for a business-hours B2B tool, and peak-weighted for ecommerce, where a Black Friday hour can carry 10x an average hour). Multiply by hours down and by the fraction of users actually blocked: a full outage is 1.0, a degraded checkout might be 0.3. Then apply an honest recovery discount: some ecommerce purchases are delayed rather than lost, while missed signups and abandoned carts mostly do not come back.
2. Recovery labor. Count everyone involved, not just the engineer typing: the two others pulled in to help, the manager coordinating, the support staff answering tickets. Multiply people x hours x a loaded rate (commonly $75 to $150 per engineer-hour for salaried staff, more for contractors or after-hours emergencies). Then multiply the total by roughly two, because the incident does not end when the site comes back: there is cleanup, data repair, the post-incident review, and the follow-up fixes.
3. Contractual and marketing costs. SLA credits owed to customers, refunds and goodwill gestures, and one line item almost everyone forgets: paid advertising that kept running while the destination was down. That traffic was bought and lost.
4. Churn and pipeline. The hardest to measure and often the largest. Customers rarely churn over one outage; they churn over the second and third. A conservative approach: estimate the number of customers or active deals plausibly lost, and value each at lifetime value, not one month’s fee. If you sell to enterprises, add the deals where your outage becomes a procurement objection six months later.
Worked examples by company size
These are illustrative arithmetic, not case studies. Substitute your own inputs.
Small SaaS, $1M ARR, 4-hour full outage. Revenue runs about $114 per hour, so direct revenue at risk is roughly $460, and since subscriptions are not refunded by the hour, the realized revenue loss may be near zero. But: two engineers for four hours plus a founder coordinating, then doubled for follow-up, is around 20 loaded hours, call it $2,000 to $3,000. Add a handful of goodwill credits. Then churn: losing just two customers paying $200 per month, valued at a two-year lifetime, is $9,600. Total: roughly $12,000 to $15,000, of which direct revenue was almost nothing. For small SaaS, downtime is a trust problem wearing a technical costume.
Mid-market ecommerce, $20M annual revenue, 4-hour outage during a busy evening. Average revenue is about $2,300 per hour, but a busy evening runs perhaps 2x average, so $18,000 or so of orders were blocked. Assume half of those buyers return later: $9,000 lost outright. Recovery labor: five people for four hours plus follow-up, around $6,000 loaded. Paid ads that kept sending traffic to a dead checkout: $1,500 wasted. Refunds and support overtime: $2,000. Total: roughly $18,000 to $20,000 for one evening, before counting any repeat-purchase erosion. If the same outage had landed on a peak sale day, the revenue line alone would have been several times larger.
Enterprise. At this scale the survey figures stop looking inflated. An hour of a down order-processing or payments system idles hundreds or thousands of employees (all still being paid), triggers contractual penalties, and can create regulatory or audit exposure. That is how ITIC’s respondents get to $300,000-plus per hour without exaggerating: the labor line alone, thousands of loaded employee-hours, can carry the figure before a dollar of lost revenue is counted.
The figures at a glance
| Finding | Figure | Source |
|---|---|---|
| Average IT downtime cost | $5,600 per minute (about $336,000 per hour), range roughly $140,000 to $540,000 per hour | Gartner, 2014, still the most-cited estimate |
| Enterprises above $300,000 per hour | Over 90 percent of mid-size and large firms surveyed | ITIC hourly-cost-of-downtime surveys, recurring |
| Enterprises at $1 million-plus per hour | Roughly 40 percent | ITIC surveys |
| Outages costing over $100,000 total | Over 60 percent in 2022, up from 39 percent in 2019 | Uptime Institute Annual Outage Analysis |
| Outages costing over $1 million total | About 15 percent | Uptime Institute Annual Outage Analysis 2022 |
| Operators with an outage in the last 3 years | Around four in five | Uptime Institute surveys |
| Outages caused by changes to a live system | Roughly 70 percent | Google SRE published research |
| Outages involving human error | Majority, when process failure is included | Uptime Institute, post-incident analyses |
| Time to restore service, elite vs low performers | Under 1 hour vs between a week and a month | DORA State of DevOps research |
| Engineering time lost to unplanned work | Commonly reported in the 20 to 40 percent range | DevOps and SRE industry surveys |
Caveat: every figure above varies by survey year, sample, and methodology, and the enterprise-weighted ones do not predict SMB costs. Treat them as directionally reliable patterns, use the calculation above for your own number, and be suspicious of any page that quotes downtime costs to the dollar without saying whose dollar.
The quieter costs
Two impacts rarely make the survey headlines but show up in every real incident. Engineer drain: industry surveys of on-call practice repeatedly find teams losing a double-digit percentage of engineering capacity to firefighting and unplanned work, which is roadmap time silently converted to recovery time, and it compounds, because the fixes not shipped this quarter become the incidents of next quarter. Trust decay: status-page history is now standard vendor due diligence, and prospects read your last twelve months of incidents before they read your feature list. Reliability compounds in both directions.
Reading the trend
The consistent multi-year story is this: outages per site are slowly declining as tooling and cloud platforms mature, while the cost of each serious outage rises, because more revenue is online, architectures are more interconnected, and customer tolerance is lower. The rational response is not fear, it is arithmetic: run the calculation above once, invest in the boring preventions (deploy safety, capacity alerts, dependency awareness), and have a plan for the day the numbers say will still come.
When the number stops being a statistic
Every figure on this page is an abstraction until production is down and it’s yours. Erzon exists for that moment: we triage for free, respond within one business hour, quote flat before work starts, and every fix ships with a written root-cause report and a prevention list, so the same hour doesn’t cost you twice.
Questions on this
How much does one hour of downtime cost a small business?
There is no single defensible number. The famous survey figures (Gartner's often-quoted $5,600 per minute, ITIC's $300,000-plus per hour) describe mid-size and large enterprises, not SMBs. For a small business the honest method is to compute your own: hourly revenue through the affected system, times hours down, plus the loaded cost of every person pulled in, plus SLA credits and refunds, plus an estimate for churn. Even modest inputs usually land in the low thousands per hour, and for small SaaS companies the churn and trust component typically outweighs the direct lost revenue.
How do I calculate the cost of downtime for my business?
Use four line items. One, direct revenue: hourly revenue through the affected system, times hours down, times the fraction of traffic actually blocked. Two, recovery labor: everyone involved, times hours, times a loaded rate (commonly $75 to $150 per engineer-hour), then multiplied by roughly two to cover post-incident cleanup, review, and follow-up fixes. Three, contractual costs: SLA credits, refunds, and any paid advertising that kept running while the site was down. Four, churn: a conservative estimate of customers or deals lost, valued at lifetime value, not one month's fee.
What causes most production outages?
Across sources such as the Uptime Institute's annual outage analyses and published post-incident reports, the leading causes are consistently unglamorous: change-related failures (bad deploys and configuration mistakes), power and network infrastructure issues, capacity exhaustion, and third-party or cloud provider dependencies. Google's SRE literature puts roughly 70 percent of outages down to changes in a live system, and the Uptime Institute has repeatedly found human error implicated in the majority of incidents once process failure is included.
What is a realistic benchmark for time to restore service?
Google's DORA State of DevOps research is the standard reference. Its published performance bands put elite teams at under one hour to restore service, high performers at under a day, and low performers at somewhere between a week and a month. Most teams sit in the middle, and in practice the bulk of any incident is spent diagnosing, not fixing: the fix is often a one-line change found after hours of narrowing down.
Are outages getting more or less frequent?
Industry tracking, notably the Uptime Institute's annual reports, suggests the rate of outages per site has gradually improved, but the cost of each serious outage keeps rising because more revenue flows through digital systems and architectures are more interconnected. Fewer outages, higher stakes per outage, is the fair one-line summary.