On the morning a monthly operating report is closed, one cell in the production series remains empty — a break in the field data feed, an instrument offline, a contractor submission not yet delivered. The analyst closing the report, wanting the formulas to hold, carries forward the prior month's figure or writes the average of the two surrounding months; the action takes seconds and is recorded nowhere. The report goes to the board the following day, travels onward to the lender as part of the monthly package, and three quarters later is uploaded to a data room. At no link in that chain does the cell announce itself as an assigned rather than an observed value; from the moment it enters the table, the fill carries exactly the same typographic weight as a measurement.

The same pattern appears more densely on the transaction side. Confronted with four gaps in a thirty-six month performance series, a buy-side modelling team will generally prefer completion to truncation: where customer contracts lack start dates, a typical term is assumed; where exit dates are absent from the attrition calculation, mid-period is adopted; where warranty cost is unreported in certain months, the average of reported months is spread across them. Each of these choices is individually defensible, and most look more reasonable than the alternative — narrowing the dataset until the analysis loses meaning. Yet although imputation is itself an assumption, it never reaches the assumption schedule; the model's assumptions page carries the discount rate, the inflation path and terminal growth, but not the count of filled cells.

The behaviour is established under the name **imputation bias** — the systematic distortion of a distribution, and of the relationships between variables, produced by inappropriate substitution for missing values — and its mechanics operate along three distinct routes. Where a series is completed with a measure of central tendency, every inserted value sits at zero distance from the mean, so variance contracts arithmetically: standard deviation falls, the weight of the tails thins, and the distribution presents as narrower than it is. Where regression-based imputation is used, the filled value is by construction the value the model predicts, with the consequence that the relationship between predictor and target validates itself as imputed observations accumulate and correlation strengthens artificially. Last-observation-carried-forward, meanwhile, manufactures a continuity the series never possessed, and in trend analysis this is the most misleading of the three.

The shortcut is not in itself an error; under certain conditions it genuinely reduces cost. Closing a report on time, keeping a model operational, meeting the calendar of a credit committee package are concrete values, and the opportunity cost of waiting for complete data is real. Imputation serves its purpose without distorting the centre of a series when the reason for the gap is independent of the quantity being measured — instrument servicing performed on a planned schedule, a reporting delay arising from an administrative cause. The difficulty lies not in the existence of the shortcut but in its persistence after the condition changes, because gaps in institutional data are seldom independently generated.

Absence itself carries information, and the information it carries typically points in an unfavourable direction. A measurement interruption most often coincides with a fault, a shutdown, a grid curtailment or a period in which the site team was occupied by an unusual event; a contractor report tends to arrive late in precisely those months where reported performance would have fallen below the contractual threshold; a disputed invoice sits blank in the income statement because it has not been finalised. When gaps cluster in adverse periods, filling them with the average of surrounding months systematically lifts those periods, and the series comes to display both a higher mean and a narrower dispersion. Two errors thus accumulate not in the same direction but toward the same result: the asset reads as both better and more predictable than it is.

The valuation consequence of that reading is direct, since capital markets price predictability explicitly. Of two cash flow series with identical means, the less volatile trades at a higher multiple and a thinner risk premium; on the debt side the effect is sharper still, because debt sizing looks not at the mean but at the lower tail. Where a minimum DSCR threshold is anchored to a P90 scenario computed from a smoothed series, the true lower tail remains fatter than the model perceives, and the result is covenant headroom tested unexpectedly in the first adverse period. Because reserve account calibration derives from the same series, the buffer itself is sized thinner than the underlying distribution requires.

On the transaction side the cost surfaces less in price than in risk allocation. Where a buyer's data quality review identifies provenance uncertainty within a series, the typical response is not a blanket reduction in valuation but a transfer of that uncertainty back across the table: escrow ratios rise, earn-out thresholds are redrawn to exclude imputed periods, a discrete data integrity head is added to the representations and warranties package, and an independent verification exercise enters the conditions precedent. Each of these defers value the seller can convert to cash; even where the headline price appears preserved, both the timing and the probability of actual receipt have changed. In operating contracts the effect emerges earlier still, since an availability guarantee computed from an imputed measurement series carries the dispute between contractor and operator into the post-closing period.

At the organisational layer the mechanism becomes self-sustaining. A value imputed once becomes the starting point for the following year's budget, so what is debated in the budget meeting is not realised performance but performance believed to have been realised, and contesting that figure requires first remembering that it was a fill — a recollection usually resident in someone who has since left the team. The moment a gap closes it ceases to be interrogable; an empty cell in a table generates a request, whereas a populated cell generates nothing. Institutional memory thereby inherits its own estimates as observations, and that inheritance passes through no approval mechanism whatsoever.

What neutralises the tendency is not individual attentiveness or analytical rigour but a record architecture that carries the provenance of each figure, and this architecture is built from four separable components. The first is a gap inventory prepared before analysis begins: how many cells in each series are missing, how the gaps cluster across the calendar, and which operational event each cluster coincides with, held in a document distinct from the analytical output. The second is a provenance flag — whether a cell is observed or assigned, carried within the dataset itself and preserved through report generation. The third is method disclosure paired with a sensitivity band: the imputation method is stated, and the same calculation is repeated under at least two alternative methods so that the range of the result is visible. The fourth is a threshold rule barring any imputed value from entering a covenant, guarantee or acceptance calculation as a single point estimate; it enters such a calculation only with its band attached.

In the projects BEIREK manages, this architecture is instituted at the point the reporting line is first constructed and operated as a delivery condition rather than an analytical habit. Templates through which site and contractor reports are transmitted carry observed and assigned values in separate columns; every cell filled at the monthly close is entered into a gap record together with its method and rationale; and that record travels not as an annex to the package going to the lender or the investment committee but as part of it. In financial closing and acquisition processes the same record is reviewed before the data room is populated, and the calculations that govern the lower tail — debt sizing, reserve calibration, availability commitments — are run once more with imputed periods excluded.

The concrete output of that rhythm is more often a negotiating position than a table. Where a counterparty's data quality review encounters a gap, the fact that the gap is already inventoried and its effect already expressed as a band materially reduces the probability that the finding converts into a risk allocation demand, since what generates leverage in negotiation is not the existence of an omission but its discovery by the other side. The same record, during the operating phase, fixes in advance the ground on which performance disputes with a contractor are conducted, and that fixing shortens the duration of the disagreement.

A cell left empty in a dataset carries no less information than a cell that has been filled; the information it carries is simply uncomfortable, and removing that discomfort is paid for not in the cost of the analysis but in the cost of its credibility. The maturity of an institution's data discipline is measured not by how many gaps it can close, but by its ability to demonstrate which gaps it chose to leave open and who recorded that choice.