A climber’s gloved hands check two independent safety ropes attached to separate anchors on a mountain rock face.

The Last Five Percent Is Not Waste: The Leadership Case for Strategic Slack

Leadership | Operational Resilience | Corporate Governance

Relentless efficiency can make an organization look exceptional—until one supplier fails, one system freezes or one team reaches its limit. The leader’s task is not to choose between discipline and resilience. It is to know which reserves must never be optimized away.

By Frank Farnel | Responsible Public Affairs | September 15, 2026

A climber’s gloved hands check two independent safety ropes attached to separate anchors on a mountain rock face.
Redundancy looks inefficient only until the primary line fails. Strategic slack is the capacity an organization deliberately preserves so that one shock does not become a collapse. Illustrative image.

Executive Summary

  • Efficiency and resilience are not opposites, but they optimize for different conditions. Efficiency performs best inside expected variation; resilience preserves acceptable performance when assumptions fail.
  • Strategic slack is not unallocated spending. It is a governed portfolio of reserves—capacity, time, options and decision authority—attached to critical outcomes and tested against credible disruptions.
  • Toyota’s post-2011 supply-chain mapping, TSB’s 2018 banking-system migration, Maersk’s NotPetya experience and Netflix’s controlled failure experiments show four different truths: visibility matters, transition options must be preserved, hidden dependencies can turn local failure into system failure, and redundancy must be exercised.
  • The board should not ask whether the company is “resilient” in the abstract. It should ask which services must continue, how long they can be interrupted, what dependency could defeat the plan and who may override normal rules during disruption.
  • The leadership challenge is to protect the right slack while removing the wrong slack. Untested backups, duplicated bureaucracy and excess inventory without a defined risk purpose are costs. A trained alternate team, recovery time, diversified supply or a rehearsed manual process can be strategic assets.

The Efficiency Trap

Few executives are promoted for defending an empty production slot, a second supplier that costs slightly more or a team that appears to have time to think. They are promoted for improving utilization, shortening cycle time and removing expense. Those disciplines matter. A company that calls every inefficiency “resilience” will eventually become slow, expensive and complacent.

Yet the opposite mistake is now common. Organizations remove buffers until ordinary work consumes nearly all available capacity. They consolidate suppliers, centralize decisions, automate exceptions and reduce headcount around the assumption that tomorrow will resemble the average of yesterday. The resulting operation can be impressively lean and surprisingly brittle.

The 2026 Business Continuity Institute report describes a field moving from framework design to delivery. That transition is important. A policy, a risk register and a crisis committee do not themselves create resilience. Continuity depends on whether people, systems, suppliers and decision rights still work when the organization is already under strain.

This is a leadership issue before it is a technical one. Specialists can identify recovery-time objectives, test systems and model supply interruptions. Only senior management and the board can decide which level of interruption is unacceptable, which reserve deserves funding and which short-term efficiency target must give way to long-term continuity.

The hard question is therefore not, “How much slack can we afford?” It is, “Where would the absence of slack make one failure irreversible?”

What Strategic Slack Actually Means

Organizational slack has often been treated as the difference between resources available and resources strictly required for current operations. That definition is useful but incomplete. It can include cash, inventory, spare capacity, time, talent or managerial attention. Some of it supports adaptation. Some merely conceals weak execution.

Recent empirical work helps sharpen the distinction. A 2022 study by Dominic Essuman and colleagues found that resource slack supports operational resilience through organizational attention. Reserves create value when they allow the organization to notice, interpret and respond to disruption. Resources sitting outside any process of attention or action may never become resilience at all.

That finding connects with research on high-reliability organizations. Karl Weick, Kathleen Sutcliffe and David Obstfeld argued that reliability depends on collective mindfulness: sustained attention to failure, operational detail, complexity, recovery and expertise. In practical terms, a spare resource matters only if the organization can see when it is needed and move it to the point of failure.

Strategic slack can therefore be defined more precisely:

Strategic slack is a deliberately funded, clearly owned and regularly tested reserve that protects a critical outcome when normal assumptions no longer hold.

Four reserves make that definition operational.

1. Capacity Reserve

This is the ability to absorb additional demand or replace lost capability: backup infrastructure, cross-trained personnel, spare production, liquidity, inventories of genuinely critical components and surge contracts. Capacity reserve is visible on a balance sheet or staffing plan, which makes it an easy target for cost reduction. Its value, however, lies in the loss it prevents rather than the revenue it produces in normal periods.

2. Time Reserve

Time is often the first buffer leaders surrender. Schedules leave no room for verification, maintenance, recovery or deliberation. A time reserve may be a recovery window before a regulatory deadline, maintenance performed before equipment reaches its limit, protected management time for weak signals or a product-release sequence that allows an early defect to be contained. Time becomes strategic when it prevents urgency from destroying judgment.

3. Option Reserve

An option reserve preserves more than one credible path. It can take the form of a second supplier, interoperable technology, transferable data, alternative transport routes, a manual workaround or a contract that allows volume to shift. Optionality usually costs more than dependence on the cheapest single path. The relevant comparison is not unit price against unit price; it is total cost against the probable impact of losing the sole path.

4. Decision Reserve

Organizations also need slack in authority. A crisis slows when every exception must climb the hierarchy, yet decentralization without boundaries can multiply damage. Decision reserve means designated people have information, competence and permission to depart from normal procedures within predefined limits. This is the human counterpart to technical redundancy.

These reserves reinforce one another. Spare inventory without supplier visibility may be the wrong inventory. An alternate system without trained operators is a prop. A crisis team without delegated authority is a conference call. A decentralized decision without reliable information is a guess.

Case One: Toyota Turned a Blind Spot Into Visibility

The Great East Japan Earthquake of March 2011 exposed the depth of modern automotive supply chains. Toyota later reported that procurement disruption extended to approximately 1,260 items and could affect up to 80 percent of its global vehicle production. The problem was not simply insufficient inventory. It was limited visibility several tiers below direct suppliers.

Toyota’s response was not to abandon lean production. It built a more discriminating form of resilience around critical dependencies. By 2016, the company said its RESCUE supply-chain database contained information on roughly 6,800 items. It also described regular training with suppliers and the integration of business-continuity practices into routine work.

The distinction matters because “more inventory” is a blunt response. If leaders do not know which sub-tier plant, material or tool creates a bottleneck, they may hold large stocks of easily replaceable parts while remaining exposed to one obscure component. Visibility directs slack toward the dependency that can actually stop production.

Toyota’s system should not be romanticized. No database eliminates earthquakes, pandemics or semiconductor shortages, and Toyota itself later faced production cuts during the global chip crisis. The more credible lesson is narrower: lean operations and resilience can coexist when reserves are selective, supplier information is maintained and continuity work is treated as part of operations rather than an annual compliance exercise.

Analysis: Toyota’s leadership choice was to preserve the productivity logic of its operating system while changing what counted as waste. Unmapped dependency became waste. Unrehearsed supplier communication became waste. Selective buffers around components with long recovery times became protection rather than excess.

Case Two: TSB Removed the Exit Ramp From a Critical Migration

In April 2018, TSB moved customer and corporate-service data to a new banking platform. The data migration itself succeeded, but the platform immediately experienced technical failures. According to the joint findings of the Financial Conduct Authority and Prudential Regulation Authority, every TSB branch and a significant proportion of the bank’s 5.2 million customers were affected by the initial problems. Disruption reached branch, telephone, online and mobile banking, and some issues continued until the bank returned to business as usual in December 2018.

The regulators imposed a combined £48.65 million penalty in December 2022. They found that TSB had not organized and controlled the migration program adequately and had failed to manage the operational risks arising from its outsourcing arrangements with a critical third-party supplier. TSB had also paid £32.7 million in customer redress. The figures do not capture every commercial or reputational consequence, but they establish that the transition failure produced material harm long after the technical cutover weekend.

TSB’s board later commissioned an independent review by Slaughter and May. When publishing it in November 2019, the board said the purpose was to identify lessons for TSB and the wider industry, while noting areas where the bank did not agree with the report. A separate UK Parliament Treasury Committee report treated the episode as the most prominent example in a broader examination of recurring IT failures in financial services.

The central leadership lesson is not that complex migrations should never happen. Legacy systems also create risk, and indefinite delay can be its own form of fragility. The issue is whether a transformation preserves enough time and optionality to stop, contain, reroute or roll back when evidence contradicts the plan. A successful data transfer is not the same as a service that customers can use under real demand.

Analysis: TSB illustrates transition risk. When a critical change concentrates customers, channels, suppliers and deadlines into one cutover, the organization spends several reserves at once. Leaders need staged exposure where feasible, explicit stop criteria, independently tested fallback arrangements and sufficient customer-service capacity for the possibility that the technical plan is wrong. Without those safeguards, commitment becomes irreversibility.

Case Three: Maersk Recovered Quickly—but Paid for Hidden Concentration

On June 27, 2017, the destructive NotPetya malware struck companies around the world. Maersk confirmed the next day that IT systems were down across multiple sites and selected business units, while APM Terminals was affected at a number of ports. The company said it had contained the issue and was implementing business-continuity plans.

Maersk’s 2017 annual report estimated the financial effect at $250 million to $300 million, including lost revenue, restoration costs and extraordinary operating costs. It also said the company had launched immediate and long-term initiatives to improve cyber resilience and reinforce continuity plans.

The case is often remembered for the speed and determination of the recovery. That deserves recognition. But rapid recovery should not obscure the initial vulnerability. A globally integrated enterprise had allowed one digital event to affect operations across business units and geographies. Integration created efficiency; insufficient containment allowed the disruption to travel.

Cyber resilience makes the logic of slack unusually clear. A backup connected to the same compromised environment may not be a backup in any useful sense. A recovery plan that depends on the same identity system, vendor or communications channel as normal operations can reproduce the original point of failure. Genuine redundancy requires separation, not duplication alone.

Analysis: Maersk demonstrates the difference between recovery capability and continuity capability. The organization mobilized impressively after the attack, but the economic loss shows how expensive it is to discover concentration during an event. Leadership must examine whether supposed alternatives share the same underlying dependency.

Case Four: Netflix Makes Failure Part of Normal Work

Most organizations test resilience in a scheduled exercise that everyone knows is an exercise. Netflix took a more demanding approach. Its Chaos Monkey tool randomly terminates virtual-machine instances in the production environment so engineers design services to tolerate instance failure. Netflix’s public documentation says the tool exists to ensure that services are resilient to such failures.

The idea developed into the broader discipline of chaos engineering: running controlled experiments to build confidence that a distributed system can withstand turbulent conditions. In their 2017 paper, Netflix practitioners described how Chaos Monkey was restricted to normal working hours so engineers could respond if an experiment revealed a weakness. That detail is important. The organization did not confuse resilience testing with recklessness; it created a bounded environment for learning.

Netflix’s example broadens the meaning of slack. Redundant computing capacity is necessary, but insufficient. The organization also preserves learning capacity: engineers have the authority, observability and time to detect what happened and improve the design. A backup that is never tested produces reassurance, not evidence.

Analysis: Netflix treats resilience as a practiced capability. The leadership contribution is cultural as much as technical: failure is surfaced deliberately, during a period when the organization can respond, so that hidden weakness becomes visible before an uncontrolled event finds it.

The Four-Reserve Leadership Test

ReserveBoard-level questionEvidence it is realFalse comfort
CapacityWhat critical service can absorb a sudden loss or surge, and for how long?Named capacity, trained people, funded inventory or infrastructure, and a measured recovery windowA generic contingency budget or backup resource already committed elsewhere
TimeWhere does the operating plan allow verification, maintenance and recovery before harm becomes irreversible?Protected windows, realistic deadlines, staged deployment and explicit stop criteriaSchedules that assume every handoff occurs on time
OptionsWhich dependency has only one viable path, even if several vendors appear on paper?Tested alternatives with different underlying infrastructure, geography or ownershipMultiple suppliers that share the same sub-tier source, platform or route
Decision authorityWho may change the plan when the plan becomes unsafe or impossible?Predelegated limits, accessible information, practiced escalation and deference to relevant expertiseA crisis committee that still requires normal approvals for every exception

When Slack Becomes Waste

A defense of strategic slack can easily become a defense of everything an organization already owns. That is not the argument. Slack loses its strategic character in four circumstances.

First, it has no critical outcome attached to it. “Extra capacity” is not a strategy unless leaders can state what service it protects and against which disruption. Second, it is not maintained. A dormant system with expired credentials, obsolete data or untrained users is inventory theater. Third, it shares the same failure mode as the primary system. Two suppliers in the same flood zone or two applications on the same identity platform may be one option wearing two names. Fourth, it is never exercised. Untested redundancy accumulates assumptions faster than capability.

Leaders should also resist the claim that every shock justifies permanent duplication. Some interruptions are tolerable. Some capabilities can be restored rather than continuously maintained. Some low-probability risks are cheaper to insure, contract or accept. The correct level of slack depends on impact, recovery time, substitutability and the organization’s obligations to customers, employees, regulators and society.

Hypothesis: As firms automate more decisions and concentrate more activity in shared digital platforms, the most valuable slack may shift away from physical inventory toward human override capacity, recoverable data, independent communications and time to verify automated action. This is a forward-looking judgment, not an established empirical finding.

What Leaders Should Do Now

  1. Name the outcomes that cannot fail. Begin with services and obligations, not assets. A server is replaceable; a payment, a safety function or a public service may have a strict maximum interruption.
  2. Set an explicit tolerance. Define how much degradation and downtime the organization will accept. Without a tolerance, resilience spending becomes either unlimited or perpetually deferrable.
  3. Map the dependency beneath the vendor. Identify shared cloud platforms, logistics routes, sub-tier manufacturers, identity systems, data sources and key individuals. Apparent diversification often hides common concentration.
  4. Assign every reserve an owner and trigger. Someone must know when backup capacity is activated, who pays for it and when normal operating rules may be suspended.
  5. Test under bounded pressure. Tabletop exercises are useful, but operational tests reveal different weaknesses. Remove a system, supplier, site or decision-maker in a controlled window and observe what the organization actually does.
  6. Measure recovery, not paperwork. Track time to detect, decide, switch, communicate and restore. Count the assumptions that failed during the exercise. A completed plan is not a performance metric.
  7. Protect the reserve during budget season. Require any proposal to remove critical slack to state the new failure tolerance, compensating control and accountable executive. Efficiency savings should never make risk disappear from the decision record.

The most useful board conversation may be a simple one: “Show us the capacity we pay for but hope not to use.” If management cannot identify it, the organization may be relying on luck. If it can identify it but cannot demonstrate when it was last tested, the organization may be paying for reassurance. If it can identify, explain and exercise it, that apparently idle capacity is doing leadership work.

Conclusion: Readiness Has a Carrying Cost

Every reserve looks expensive before a disruption and insufficient after one. That asymmetry creates a predictable leadership bias. Quarterly results display the cost of spare capacity immediately; they rarely display the catastrophe that did not happen because the reserve existed.

The answer is not to celebrate inefficiency. It is to make resilience governable. Critical reserves should have a purpose, an owner, a trigger, a test and an expiration or review date. Leaders should be able to explain why a second line, supplier, system or team is worth maintaining—and just as importantly, why another one is not.

The last five percent is not automatically strategic. But eliminating it automatically is not discipline. It is a decision to make performance depend on every assumption remaining true at once. In a world of interdependent operations, geopolitical shocks, extreme weather and digital concentration, that is not efficiency. It is an unpriced bet.

Key Evidence

  • Toyota said its post-earthquake RESCUE database stored supply-chain information on approximately 6,800 items by 2016. Source: Toyota.
  • UK regulators said every TSB branch and a significant proportion of its 5.2 million customers were affected by the initial problems following the April 2018 migration. Source: FCA and PRA.
  • The FCA and PRA imposed a combined £48.65 million penalty for operational risk management and governance failures connected with the migration program. Source: FCA and PRA.
  • TSB had paid £32.7 million in customer redress, and the regulators said the bank did not return to business as usual until December 2018. Source: FCA and PRA.
  • Maersk estimated the 2017 NotPetya cyberattack’s financial effect at $250 million to $300 millionSource: Maersk 2017 Annual Report.

Glossary

Business continuityThe capability to continue delivering prioritized products or services at an acceptable level during disruption.Chaos engineeringControlled experimentation on a system to discover weaknesses and build confidence that it can withstand failure.High-reliability organizationAn organization that performs complex, high-risk work with sustained attention to weak signals, operational detail, expertise and recovery.Operational resilienceThe ability to prevent, absorb, adapt to, recover from and learn from disruption while protecting critical outcomes.Single point of failureA component, person, supplier or process whose loss can stop the wider system because no independent alternative exists.Strategic slackA deliberately funded, owned and tested reserve of capacity, time, options or authority linked to a critical outcome.

References and Further Reading

Standards and Current Resilience Research

  1. Business Continuity Institute. BCI Operational Resilience Report 2026. BCI, May 13, 2026.
  2. International Organization for Standardization. ISO 22301:2019—Security and Resilience: Business Continuity Management Systems—Requirements. ISO, 2019.
  3. International Organization for Standardization. ISO 22316:2017—Security and Resilience: Organizational Resilience—Principles and Attributes. ISO, 2017.
  4. Dominic Essuman, Patience Aku Bruce, Henry Ataburo, Felicity Asiedu-Appiah and Nathaniel Boso. Linking Resource Slack to Operational Resilience: Integration of Resource-Based and Attention-Based PerspectivesInternational Journal of Production Economics, Vol. 254, 2022, Article 108652.
  5. Karl E. Weick, Kathleen M. Sutcliffe and David Obstfeld. Organizing for High Reliability: Processes of Collective MindfulnessResearch in Organizational Behavior, Vol. 21, 1999, pp. 81–123.

Case-Study and Primary Sources

  1. Toyota Motor Corporation. Five Years On: Toyota’s Efforts to Build a Disaster-Resilient Supply Chain. March 11, 2016.
  2. Financial Conduct Authority and Prudential Regulation Authority. TSB Fined £48.65m for Operational Resilience Failings. Bank of England, December 20, 2022.
  3. TSB Bank plc. TSB Board Publishes Independent Review of 2018 IT Migration. November 19, 2019.
  4. House of Commons Treasury Committee. IT Failures in the Financial Services Sector. Second Report of Session 2019, October 28, 2019.
  5. A.P. Møller–Mærsk A/S. Cyber Attack Update. June 28, 2017.
  6. A.P. Møller–Mærsk A/S. Annual Report 2017. February 9, 2018.
  7. Netflix. Chaos Monkey Documentation. Accessed September 15, 2026.
  8. Casey Rosenthal, Lorin Hochstein, Aaron Blohowiak, Nora Jones and Ali Basiri. Chaos Engineering. IEEE Software / arXiv preprint, 2017.

Source and Methodology Note

Research was completed on September 15, 2026. The article prioritizes official corporate disclosures, regulatory findings, standards bodies and peer-reviewed or scholarly research. Corporate descriptions of their own resilience programs are treated as primary evidence of actions and stated intent, not independent proof of effectiveness. The Toyota and Netflix cases illustrate designed resilience practices; they do not establish that either organization is immune to disruption. The TSB and Maersk figures describe specific historical events and should not be generalized into current performance assessments. The proposed “four reserves” model and the interpretation of the cases are the author’s analysis. The paragraph explicitly labeled “Hypothesis” is forward-looking and has not been presented as a verified empirical conclusion.

Suggested Internal Links

Hashtags: #Leadership #OperationalResilience #CorporateGovernance


Discover more from Responsible Public Affairs

Subscribe to get the latest posts sent to your email.

Share This :
Facebook
X
LinkedIn
Print
Email
WhatsApp

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from Responsible Public Affairs

Subscribe now to keep reading and get access to the full archive.

Continue reading