In the weekly operations review, shipment performance is typically reported on two separate lines: the share of orders that left on the promised date, and the share of requested quantity that was actually filled. Read in isolation, each figure sits within an acceptable band, and the agenda moves briskly to the next heading; yet the composition of the calls reaching customer service that same week, taken together with the tone the sales team brings back from the field, does not reconcile with that picture. The gap is usually attributed to measurement lag, to a handful of individual accounts, or to seasonal variation. What the practice of reporting two ratios separately actually produces, however, is a structure in which the single performance the customer experiences appears in no document at all, because the customer does not live through two events — the customer lives through one delivery.

A second observation sits in how a partial shipment is recorded inside the system. When an order line ships incompletely, the residual quantity is typically carried into a new delivery line while the original line is treated as closed, and that closed line enters the population counted as delivered on time. In parallel, moving a delivery date by mutual agreement with the customer quietly repositions the reference point against which performance is assessed, since the comparison is no longer made against the first commitment given but against the most recently revised one. Neither recording practice is adopted in bad faith; both follow naturally from standard enterprise resource planning logic, having been designed to keep order books clean rather than to keep promises honest, and precisely for that reason neither is questioned anywhere along the reporting chain.

The name for this pattern is OTIF failure — a delivery that is not on time, not in full, or that satisfies neither condition. What distinguishes the measure is that its two components are assessed jointly rather than independently: an order line counts as successful only when it arrives on the date promised, in the quantity requested, and in acceptable condition. The components are not averaged; their probabilities are multiplied. Even in an organization where timeliness and completeness each look strong when reported alone, the composite ratio will fall materially below both, and it will fall further the weaker the correlation between the two components happens to be. That composite figure, rather than either of the reported lines, is the number the customer sees and remembers.

The decision to ship partially is not irrational; it is, at the moment it is taken, a shortcut that genuinely lowers cost. Releasing what is ready rather than holding it frees warehouse space, recognizes a portion of the period's revenue, meets part of the customer's immediate need, and signals good faith in the commercial relationship. The difficulty lies not in the shortcut itself but in its persistence after the underlying condition has changed: once the customer has anchored its own production schedule to that delivery, or once the order line forms part of an assembly kit, a commissioning package, or a field installation, completeness stops producing linear benefit. For material that carries meaning only as a set, a high fill rate translates into zero usability the moment the single missing item happens to be the critical one.

The organizational layer is established at the moment of commitment. In most organizations, the unit that gives the delivery date and the unit obliged to hold it are not the same: sales issues the date, planning knows the capacity, procurement knows the lead time, and manufacturing controls the sequence, with none of them carrying another function's constraint inside its own decision. Absent a discipline that interrogates capacity and material availability together at the moment the promise is made, the date emerges from negotiation rather than from capability. This configuration systematically presents OTIF failure as an execution problem and directs remediation effort toward production and logistics, whereas a substantial share of the deviation originates in the meeting where the commitment was given rather than on the floor where it was pursued.

The first place the cost accumulates is not the supplier's balance sheet but the customer's. Variance in delivery performance converts, on the buyer's side, into safety stock; as reliability declines, the buyer holds more inventory, releases orders earlier, and plans against a wider tolerance window. The carrying cost of that inventory sits in the customer's working capital, but it does not remain there: it returns to the supplier in the next price negotiation as a discount demand, in the annual review as a downgraded scorecard rating, or, most quietly, as the qualification of a second source. The financial trace of OTIF failure therefore appears rarely in the late-delivery penalty line and far more often in unit price and in share of wallet — two surfaces on which it never carries an OTIF label.

The second accumulation point is the supplier's own cost structure, which escapes recognition because it is distributed rather than concentrated. Every split order line multiplies the freight, packaging, invoicing, documentation, and reconciliation work attached to it; the premium freight paid to expedite, the unplanned overtime, and the rescheduling effort scatter across separate expense categories and never gather in a single place where they can be weighed against the original decision. Short or late lines are, in addition, the most common stated ground for invoice dispute, and the dispute process extends the maturity of the receivable, lifts days sales outstanding, and ties the cash conversion cycle directly to delivery performance. For that reason, whether a working capital squeeze has supply chain origins is a question rarely framed correctly in a credit discussion.

The third surface is contractual and, ultimately, valuation-related. For large corporate buyers and retail chains, the OTIF threshold on the supplier scorecard functions less as a penalty trigger than as the condition of remaining on the list, and for a supplier that falls below it the consequence is seldom the amount deducted but the quiet cessation of new order flow. When a share or asset sale comes into view, buy-side analysts cross the OTIF series against customer concentration: in a company whose revenue is gathered in a few accounts and whose delivery performance has trended downward, those two facts together tend to move not the multiple but the structure of the transaction. The typical outcome is a portion of consideration tied to an earn-out, a higher escrow percentage, contract renewal confirmations added to the conditions precedent, and a representation and warranty package extended to cover supply obligations.

The first mechanism that neutralizes this tendency is not individual vigilance but measurement design, and four of its components require separate fixing. The unit of measurement should be the order line rather than the order, because the customer builds its plan at line level. The reference date should be the first confirmed date rather than the revised one, with revisions tracked as a distinct series. The measurement moment should be customer acceptance rather than dispatch, given that transport risk forms part of performance under most Incoterms configurations. And the timing window and quantity tolerance should be defined in writing in advance rather than negotiated at the moment of application. Once these four are fixed, the reported ratio typically falls — a movement that represents not deterioration but the first accurate visibility of the existing position.

The second mechanism is keeping the record at the moment of commitment rather than at the moment of approval. When a delivery date is issued, the capacity assumption, the material lead time, and the backlog position against which it was issued are rarely captured at that instant, and in their absence any subsequent deviation can only be debated through its outcome — a debate that resolves, predictably, into a negotiation over accountability between functions. The complement to that record is the classification of every deviation into one of six headings: capacity, supply, quality and rework, data error, commitment error, and logistics. Where no such classification exists, improvement resources migrate toward the most visible line, which is logistics, while a meaningful proportion of deviations continues to accumulate under commitment and data.

The mechanism BEIREK installs in capital-intensive projects and multi-asset operations joins these two layers: a promise record captured at the moment of commitment, a date series in which revisions are counted separately, a weekly review rhythm that classifies deviations against the six headings above, and a calibration exercise that carries the output of that rhythm into the contractual position. The contractual work consists of resetting the liquidated damages cap, the service credit thresholds, and the conditions under which partial delivery is accepted against the observed distribution of performance rather than its average, since a penalty clause drafted around mean performance protects no party in the tail events where it matters. Reporting to the board and the investment committee then carries the distribution rather than a single ratio, because the spread between average OTIF and OTIF at the three largest accounts is, in any structure with revenue concentration, more informative than the average on its own.

Delivery performance is among the least manipulable indicators through which a company reveals its operational maturity to the outside world, for the simple reason that it accumulates in the counterparty's own records independently of any contract, presentation, or management representation. The operative question is therefore not which band the ratio occupies, but whether the company knows the difference between the number it measures and the number its customer measures — and the magnitude of that difference tends to say considerably more about the institution than the delivery performance itself ever does.