A recurring scene plays out in budget review sessions: finance presents the average forecast deviation for the first three quarters, the figure looks defensible, and the meeting moves to the next agenda item. Opened up on a project-by-project basis, however, the same period tells a different story — deviation on small, repeatable work stays inside a narrow band, while on a handful of large deliveries it widens several-fold in both directions. The average compresses two distinct behaviours into one number, and what it compresses is precisely what requires management. The company has reported its hit rate accurately without reporting where risk has accumulated.

The same pattern repeats across sales forecasting, supplier lead times, currency and commodity assumptions, and working capital projections. On line items that are small, numerous and broadly similar to one another, forecasts track outcomes closely; on items that are large, few in number, and each carrying its own negotiation history, deviation opens up without much apparent constraint. The common internal reading treats this as a competence question — the team forecasting large work is presumed less disciplined. The observation itself points elsewhere: what produces the deviation is not the calibre of the team but the structural position of the item being forecast.

The name for this behaviour is heteroskedasticity — the systematic variation of error variance with the conditions of the observation. The assumption that a forecasting model's errors are distributed with constant width across the entire observed range is one of the foundational conveniences of statistical inference; where it holds, a single deviation measure represents the whole range and confidence intervals retain meaning. Where it breaks — where error is narrow at small values and wide at large ones — the computed average deviation remains arithmetically correct while describing no individual forecast at all. The corporate translation is direct: the reported accuracy figure is true, and it says almost nothing about the next large delivery.

Understanding why this compression proves so durable inside organisations requires seeing how genuinely functional the single average is. Presenting every line item with its own band at board level would multiply agenda time and drown discussion in detail, whereas one deviation figure permits period-to-period comparison, lends itself to target setting, and can be wired into an incentive scheme. The simplification really does lower the cost of deciding, and so long as conditions remain homogeneous the information loss stays bounded. The difficulty arises when portfolio composition shifts — when a few large engagements begin carrying the decisive share of revenue — and the same summary remains in use; the shortcut outlives the condition that justified it.

The first cost of that persistence surfaces in reserve placement. Contingency set against average deviation is distributed at roughly uniform rates across line items, while the actual shape of deviation shows the reserve to be excessive on small, predictable work and insufficient on large, volatile work. The result is a reserve pool that looks adequate in aggregate yet runs thin exactly where it is called upon. This rarely appears first as a cash shortfall; it appears as an approval delay, because the incremental budget request exceeds ordinary delegated authority, becomes attached to the board calendar, and shifts the project schedule for administrative rather than financial reasons.

The second cost sits on the credit side. When DSCR or net debt to EBITDA thresholds are calibrated against a base case built on average historical deviation, breach probability appears evenly distributed across periods. In practice breach risk concentrates in the quarters where large deliveries enter revenue recognition, because that is where deviation widens. The lender-side inclination to shift covenant packages from periodic toward continuous testing arises from precisely this asymmetry. For a borrower, the practical consequence is that technical breach most often emerges not at the end of a poor year but during the densest quarter of a good one.

The third and most expensive cost becomes visible at the valuation table. An acquirer or investment committee reviewing a target's forecasting history looks at the distribution rather than the mean; the question asked is not how closely budget was met but how far deviation opened, and in which category of work. A company unable to produce that breakdown cannot demonstrate that its forecasting process is repeatable independently of its founder, however respectable its headline accuracy. The consequence typically appears not in the multiple itself but in the transaction structure — an additional condition precedent, a widened representation and warranty scope, or an earn-out disaggregated into triggers by work type rather than settled against a single aggregate. The price holds; the party carrying the risk changes.

The mechanism that neutralises this tendency is not an exhortation to forecast more carefully but a change in the architecture within which forecasts are recorded, since instruction at the individual level leaves the structural source of the distribution untouched. The workable intervention separates into four components: (a) every forecast record carries, alongside the realised value, the condition tags prevailing at the moment of estimation — project size band, counterparty count, delivery duration, degree of novelty; (b) deviation is reported by band rather than as a single mean, which is to say accuracy is defined at segment level; (c) contingency allocation is tied not to average deviation but to the deviation width of the band itself; (d) covenant and liquidity thresholds are tested against the portfolio composition of the period in question rather than against a period average.

On engagements BEIREK manages, this is established not as a separate analytical report but as the format of the forecast record itself. Every budget and schedule assumption is logged with its condition tags at the moment of proposal rather than at the moment of approval; when the outcome arrives, deviation writes itself into the band it belongs to, and what emerges at period end is not a single accuracy percentage but a map of width across bands. The function of that map is not to audit the past but to determine where reserve sits on the next drawdown; contingency migrates toward the band where width has grown, not toward the mean.

The second mechanism binds that map to the governance rhythm. What gets discussed in the monthly progress session is not the magnitude of a deviation but whether it fell inside its expected band — a large deviation within band is not a management problem, while a small deviation outside band indicates that something has been mislabelled in the model and warrants examination. That distinction moves the forecasting conversation off the axis of culpability and onto the axis of calibration, and in practice it accelerates the upward flow of bad news, because the price of reporting a deviation early is now a band update rather than a personal performance note.

The diligence-day value of such a record architecture materialises long after it is built. What an acquirer searches for in a data room is less how well the company forecasts than whether the company itself knows where its forecast error concentrates; a record carrying that knowledge generates confidence despite the size of the deviation, while a record lacking it generates questions despite the height of the accuracy. What determines valuation is frequently not performance itself but the demonstrability of the conditions under which performance repeats — and the distribution of deviation is the most direct evidence available for that demonstration.

Forecasting maturity in an institution is measured, ultimately, not by how accurately it predicts but by how well it can describe the structure of its own error. A company that knows its average deviation has reported its history; a company that knows the conditions under which that deviation widens knows the confidence interval attached to its next decision. The question worth putting is not how closely this year's budget held, but how closely the places where it failed to hold resemble one another.