---
title: "The Quiet Contraction of the Failure Interval: Accounting for Reliability Erosion"
description: "MTBF decline is the shortening of average operating time between equipment failures, and what makes it institutionally dangerous is that it is experienced as a series of discrete incidents while never being measured as a trend. The erosion appears first in production schedule buffer, then in spare parts turnover, and only last on the balance sheet. The neutralizing mechanism is not individual vigilance but an asset-level failure interval record reviewed on a fixed cadence."
url: https://www.beirek.com/en/blog/mtbf-decline-asset-reliability-erosion
canonical: https://www.beirek.com/en/blog/mtbf-decline-asset-reliability-erosion
published: 2026-01-10
modified: 2026-01-10
category: "Operations & Supply Chain"
category_url: https://www.beirek.com/en/blog/category/operations-supply-chain
language: en-US
reading_time_minutes: 8
publisher: BEIREK LLC
publisher_url: https://www.beirek.com
license: "© BEIREK LLC — citation with attribution and link permitted"
keywords: ["MTBF decline","reliability erosion","maintenance capex","asset-level failure interval","technical due diligence"]
topics: ["Reliability engineering and asset management","Operations and maintenance decision architecture","Technical due diligence and valuation impact","Supply chain and spare parts exposure"]
alternate_language_url: https://www.beirek.com/tr/blog/mtbf-decline-asset-reliability-erosion
---

# The Quiet Contraction of the Failure Interval: Accounting for Reliability Erosion

> **In short:** MTBF decline is the shortening of average operating time between equipment failures, and what makes it institutionally dangerous is that it is experienced as a series of discrete incidents while never being measured as a trend. The erosion appears first in production schedule buffer, then in spare parts turnover, and only last on the balance sheet. The neutralizing mechanism is not individual vigilance but an asset-level failure interval record reviewed on a fixed cadence.

*As the average interval between failures at a plant shortens over successive years, the contraction rarely surfaces in the maintenance budget; it hides instead in the buffer that planning teams quietly add to the production schedule. Erosion advances not as an event but as a normalized curve, and it typically reaches the valuation table only when carried there by the acquiring party.*

---

In the weekly operations meeting of a manufacturing plant, downtime items are read out one by one: a bearing replacement on a filling line, a valve failure in a compressor, a drive motor renewal on a conveyor. Each item carries a plausible explanation of its own, each has a root cause written against it, each has been closed. The question never asked in that meeting is how many days elapsed between this failure and the previous failure of the same equipment, and whether that number differs from the same measurement taken three years earlier. To the extent that downtime records are kept on an incident basis, every failure presents itself as an independent case; sorted by asset and arrayed along a time axis, however, the same records tend to reveal a steadily narrowing interval in most facilities.

The same pattern asserts itself on the procurement side. Spare parts inventory levels climb year over year, the critical parts list widens, requests for expedited shipment from suppliers grow more frequent — and each of these is approved on its own reasoning: lead times have lengthened, a supplier price increase is coming, waiting proved expensive the last time. All of that reasoning is accurate. The only thing that is not accurate is treating these items as unrelated, since accelerating spare parts turnover is the first shadow that failure frequency casts on the balance sheet, and it typically surfaces in working capital before it surfaces in the maintenance budget.

The name for this behavioral pattern is **MTBF decline** — the shortening, over time, of mean time between failures. Its mechanism operates on two layers. The physical layer is comparatively well understood: wear, fatigue, tolerance drift, degradation of the lubrication regime and accumulated thermal cycling raise an asset's hazard rate non-linearly across its service life. This is expected engineering behavior and is not, in itself, a management problem. What produces the management problem is the second layer, namely that repair itself degrades reliability. Every intervention introduces a small deviation into the system through disassembly and reassembly tolerances, temporary component substitutions and non-original parts, and these deviations accumulate in a way that brings the next failure forward.

When the two layers combine, the resulting curve takes a shape that is thoroughly inconvenient from a decision-maker's standpoint: a contraction that is very slow at first and then accelerates. In the early phase, when the interval falls from, say, twelve months to ten, perceiving that as a deviation is close to impossible; seasonal load variation, raw material quality or operator turnover could each account for the difference on its own, and frequently do. In the late phase the contraction is rapid enough that what is required is no longer diagnosis but crisis management. The decision-maker receives the least signal precisely in the region of the curve where the cost of intervention is lowest.

Layered on top of this is the fact that the prevailing preferences are rational in the short run. Repair is invariably cheaper than replacement; it is expensed rather than capitalized, it clears a low approval threshold, it asks the production schedule for a one-week window, and its result is immediately visible in that the machine runs again. Replacement is capital expenditure, travels to the investment committee, requires justification built on the estimated cost of a failure that has not yet occurred, and asks production for a window measured in months. Under this asymmetry, choosing repair is not merely comprehensible from a plant manager's position but correct. The difficulty lies not in the choice but in the choice remaining fixed after the conditions have shifted: even once the interval has begun to contract, the same approval architecture continues to produce the same decision, given that the architecture never measures the interval at all.

The first visible surface of the institutional cost is not maintenance expense. Maintenance expense can remain comparatively stable for an extended period, precisely because of the architecture described above — small repairs stay within budget, and overruns are absorbed at year-end from other line items. The cost accumulates first in the **buffer allowance of the production schedule**. As the planning team observes that the line cannot hold its committed cycle time, it quietly adds slack to the program; delivery commitments stretch by a week, then by two. That slack appears nowhere as a line item, because the plan has already been updated and actuals conform to the plan. Capacity utilization stays high, delivery performance looks sound, and the plant has effectively concealed capacity it has lost by lowering its own target.

The second surface is the hardening of the supply chain. As the failure interval shortens, the proportion of expedited orders rises, and expedited ordering erodes bargaining position against suppliers directly. Leverage over price negotiation, lead time commitments and quality specification does not remain with a buyer requesting one part at a time. The consequence is a quiet increase in supplier concentration — the single source able to deliver quickly becomes, over time, the default source — and single-source exposure accumulates as a dependency manufactured by the plant's own maintenance decisions. A comparable tightening is observable on the insurance side, where rising claim frequency under machinery breakdown policies returns at renewal as higher deductibles and narrowed coverage.

The third and most expensive surface opens when the business changes hands. In the sale of an asset or a facility, the buy-side technical diligence team reads maintenance records not as an incident log but as an asset-level time series, and the narrowing interval is the first finding that series yields, typically within hours. Its effect on pricing generally arrives not as an EBITDA adjustment but through two other items: an upward revision of the maintenance capex assumption for the coming three years, and an escrow securing representations and warranties concerning post-closing equipment performance. What is painful for the seller here is that the management of the decline through repair becomes a finding in its own right, since the records demonstrate that the problem was known while the replacement decision was deferred — and that deferred capex is settled out of the seller's price rather than the buyer's.

This tendency cannot be managed through individual vigilance, because the difficulty is not that anyone is inattentive but that no one is looking at the correct time scale. The neutralizing mechanism is structural and separates into four components. The first is the unit of measurement: the failure interval is tracked not at plant aggregate but per critical asset, on a rolling twelve-month window, since an aggregate average carries little information to the extent that it permits an improving asset to mask a deteriorating one. The second is the threshold: the rate of contraction that triggers a replacement decision is defined before any failure occurs and while the asset is being commissioned, because a threshold defined afterwards is invariably calibrated to the prevailing condition. The third is the approval architecture: once the threshold is breached, the investment request is not left to the plant manager's initiative but reaches the committee agenda automatically. The fourth is record discipline, under which each repair documents whether the replaced part was original, whether an applied workaround has become permanent, and whether the intervention altered the next maintenance interval.

The intervention BEIREK establishes on capital-intensive facility and portfolio mandates binds these four components into a single operating cadence. On acquisition, commissioning or performance improvement mandates, the first construct put in place is a critical equipment list together with a failure interval time series for every asset on that list; the record is rebuilt retrospectively from existing maintenance logs, since most facilities hold the data without ever having compiled it as a series. On top of that sits a decision record defining replacement thresholds and the approval path activated when a threshold is breached, its function being to fix the rationale for a decision at the moment of proposal rather than the moment of approval, given that why a replacement request was declined is institutional knowledge of the same value as why one was accepted.

This structure cannot operate without a review held on a fixed cadence. The arrangement applied places the reliability curve not in the monthly operations meeting but in the quarterly asset review, on the reasoning that a monthly rhythm carries the noise of individual failures while a quarterly rhythm renders the trend legible. In the same session, production schedule buffer, spare parts turnover and expedited order ratio are read alongside one another — when those three indicators move together, it becomes reasonable to state that erosion has begun even where the failure interval has not yet contracted by a statistically meaningful margin. The three indicators mean different things to different parties: forward-period capex for the sponsor, pressure on DSCR for the lender, and shift-planning flexibility for the operator.

A shortening interval between failures does not say that a facility is aging; it says how the aging is being managed. The same curve becomes an input to the replacement calendar in a business where thresholds are pre-defined and records are kept, while in a business without those records it functions solely as a finding that enters the acquirer's price negotiation. The difference between the two is architectural rather than technical — and that architecture is established while the asset is being commissioned, not after the failure has occurred.

The single question a management team might reasonably put to itself is this: can the failure intervals of the five most critical assets over the past three years be read today, side by side, in one table?

## Key Points

- The shortening of the failure interval is a curve rather than an event, and incident-based maintenance records render that curve structurally invisible.
- The first institutional trace of erosion is not maintenance expense but the buffer time silently added to the production schedule and the rising frequency of expedited orders.
- Repair is always cheaper than replacement in the short run; the problem lies not in that preference but in its persistence after the underlying conditions have changed.
- At the diligence table, reliability erosion is priced not as an EBITDA adjustment but through a revised maintenance capex assumption and an escrow tied to equipment performance warranties.
- The neutralizing mechanism is an asset-level time series of failure intervals combined with a replacement threshold defined before, rather than after, the failure occurs.

## Questions

### What is MTBF decline, and why is it difficult to detect?

MTBF decline is the shortening, over time, of the average operating period between an asset's failures. It resists detection because maintenance records are usually kept on an incident basis, with each failure closed against its own root cause. The contraction becomes visible only when failures of the same asset are arrayed along a time axis; in the early phase the difference is small enough to be explained by seasonality, raw material quality or operator turnover.

### Which early indicators signal that the failure interval is contracting?

Three indicators, read together, provide early warning: buffer time quietly added to the production schedule, rising spare parts turnover, and an increasing proportion of expedited orders. These begin moving before the failure interval has contracted by any statistically meaningful margin. Maintenance expense, by contrast, is a lagging indicator, since small repairs remain within budget long enough to mask the underlying erosion.

### How does reliability erosion affect valuation in a company sale?

When the buy-side technical team reads maintenance records as an asset-level time series, the narrowing interval emerges quickly. The effect is generally priced not as an EBITDA adjustment but through an upward revision of the forward maintenance capex assumption and an escrow securing representations and warranties on equipment performance. In practice, the cost of deferred replacement is settled out of the seller's price.

### When should a replacement decision be taken in place of repair?

Left to the conditions of the moment, repair wins almost every time, since it is expensed, clears a low approval threshold and asks production for only a short window. The workable approach defines the contraction threshold triggering replacement while the asset is being commissioned, and routes the request automatically to the committee agenda once that threshold is breached. A threshold set afterwards is calibrated to the prevailing condition and loses its function.

---

Source: https://www.beirek.com/en/blog/mtbf-decline-asset-reliability-erosion
Publisher: BEIREK LLC — https://www.beirek.com
