Female engineer typing on her laptop during work hours at engineering facility

Photo by ThisisEngineering / Unsplash

Diminishing Marginal Returns in Fault-Tolerance Investment

The Productivity Cliff of Resilience Engineering

As a system becomes increasingly protected against failure (for every pound you spend, you're getting less protection than the last pound), the pound spent will yield less protection than the pound spent prior to it. This is not merely a gut feeling: it is a legitimately measured economic trend, which is now forcing technology companies and operators to rethink their capital allocation for fault-tolerance spending. Simply put, the principle of diminishing marginal returns, found in introductory microeconomics, translates to fault-tolerance investment with a high degree of precision. Data from the past decade of cloud computing spending demonstrates this trend to be irrefutable.

For example, a system with 99% uptime will have approximately 1.68 hours of downtime per week. But at 99.9% uptime, unplanned downtime will be reduced to approximately 8.7 hours per year — an impressive improvement. However, if your system is going to be operational for 99.99% of the time, you will need to implement redundant architectures, cross-region failover, and automated recovery tools, which will typically incur three to five times the cost of the infrastructure required to support 99.9% uptime. The next step from 99.99% to 99.999% uptime (which equates to approximately 5.3 minutes of downtime per year) can incur an additional order of magnitude in Costs. As Gartner estimates, the average cost of IT downtime ranges from £4,000 to £9,000 per minute; however, the incremental business value that can be gained from shortening the annual downtime remains far below the capital required to eliminate it.

Why costs accumulate faster than benefits

The dynamics contributing to the asymmetry are several. The first dynamic: the cheapest fault modes are eliminated first. Redundant power supplies, basic load balancing, and geographical replication are among the most inexpensive and simplest fault modes to address. As we continue to eliminate common and easily tractable failure types, we will eventually be left with rare, correlated, or cascading failures which require expensive design and engineering solutions — chaos engineering (simulations) and bespoke circuit-breaker logic.

As complexity increases with each additional layer of fault-tolerant logic, the risk of failure becomes greater still. Chaos engineering provides insight into how the very design elements such as health checks, retry logic, and circuit breakers can fail spectacularly due to unexpected load; and the operational waste created by retry storms is illustrative: systems designed to be fault-tolerant amplify loads on already degraded nodes and create platform-wide outages from single, localised faults. Each layer of fault-tolerant logic added to eliminate these failures further increases the cost curve while decreasing reliability.

A study in the economic impact of fault-tolerance investment

In Google's publicly available internal documents on site reliability engineering (SRE), available as an SRE book and in associated blogs, the concept of "error budgets" was developed as a formal recognition that, while striving for 100% uptime may seem advisable, it is neither economically feasible nor rational. Google's establishment of an upper limit on engineering effort spent on uptime — defined by error budgets and aligned to SLOs (Service Level Objectives) — has demonstrated a Diminishing Marginal returns curve beyond that threshold.

In addition to Google's example, the financial services sector experienced a significant increase in pre-deployment fault detection and rollback automation following the Knight Capital trading incident in August 2012. The Knight Capital incident saw a software error cause a loss of approximately $440 million (£340 million) to the firm within less than 45 minutes. By 2016, major broker-dealers in the United States increased the proportion of their technology expenditures allocated toward building resilience infrastructure from approximately 12% of all technology expenditures in 2010 to between 20% and 25% of the total technology expenditures of broker-dealers by 2016 (as noted in the research completed by Celent). However, during the same timeframe, the overall rate of occurrence of significant technology incidents within the financial services sector decreased only by approximately 15%, which demonstrates a significant disconnect between the growth of input, in terms of pounds spent, versus output improvement, in terms of incidents avoided.

Economic Disconnect Between Fault-Tolerance Investments and Efficiency

The disconnect between the investment in fault-tolerance and the resultant efficiencies does not indicate that the return on investment in fault-tolerance is zero, but represents a much greater impact from each individual marginal investment decision on whether or not to invest in a given area compared to the combined impact of all marginal investment decisions in the areas identified above. Firms that utilise their capital to invest in resilience as merely a compliance-checking exercise are creating disproportionately high volumes of investment in areas providing limited returns (such as maintaining redundant warm-standby operations) while creating disproportionately low volumes of investment in observability tools that have considerable diagnostic capability at the very beginning of the reliability curve.

Given the increasing interdependence of distributed systems, the slope of the diminishing returns curve will continue to steepen, likely to an even greater degree.