In a monthly production review, the on-time delivery figure for the same period frequently reaches the table as two numbers rather than one — the rate cited by operations differing from the rate cited by the commercial side by a few percentage points — and the gap tends to close within a few minutes, not through the application of a written rule but through two people recalling, jointly and from memory, which orders were counted and which were set aside. Each side is correct within its own definition, so nothing is treated as an error and the agenda moves on. A diligence team seated in the same room, however, records not the figure but the manner in which the discrepancy was resolved, since what enters a valuation is not the month's rate but the mechanism by which that rate was produced.

Asked where the measure comes from, management typically answers with a system name and finishes with a person's name: the underlying data sits in MES or ERP, while the table that reaches the board is assembled at month-end inside a working file in which one individual pulls the raw extract and adjusts the known exceptions by hand. Most of those adjustments are defensible — work orders closed late in the system, shipments recorded twice, deferrals caused by the customer genuinely distort the raw feed, and leaving them uncorrected would produce a misleading table. The operational test applied in diligence is not whether the adjustments were warranted; the test is whether last quarter's figure can be reproduced from the raw source, within a few hours, to the same result, with the person who normally prepares that file absent from the exercise.

How such an arrangement comes into being follows a recognizable sequence across companies. Metrics rarely originate from an intention to design a measurement system; a customer complaint, a warranty claim, a lender's reporting requirement or a quality incident raises a specific question, and a measure is defined in order to answer it. The question loses its urgency over time; the measure remains. Leaving the definition unwritten is rational at the moment it is left unwritten, since a small team working in one location and sharing an intuition about what the indicator does and does not capture gains little from documenting it. The cost surfaces once the condition changes — a second shift, a second site, a newly hired engineer, a new customer segment — and what is missing at that point is not data but a shared meaning of the data.

Where the definition is unwritten, ambiguity resolves, predictably, in the direction that produces the least friction. Planned downtime is placed outside the scope, customer-caused delays are held separately, jobs extended by an engineering change request drop out of the count; each exclusion is defensible when examined on its own, and most have in fact been defended at the time. In aggregate, however, the series slowly loses the property of being comparable with its own history, and a three-year improvement curve becomes the sum of genuine improvement and a quietly narrowing definition. To the extent that an indicator also becomes an input to bonus calculation, budget setting or board assessment, the quantity of information it carries declines; where the measuring party and the measured party sit on the same line, that outcome follows independently of anyone's intent.

A second mechanism concerns the composition of the metric set. What has been instrumented gets measured; what has not been instrumented goes unmeasured even where it determines the result. For that reason the dashboard in many technical organizations reads as an archaeology of past crises rather than as a map of present value drivers — whatever once produced a serious incident still has its indicator in place, while rework rate, commissioning duration or design revision cycle, any of which may be setting today's margin, are not tracked at all. On the ownership side, an arrangement that would not survive a single audit question inside the finance function is treated as ordinary within engineering: the person who defines the indicator, the person who produces it, and the person whose performance is assessed against it are frequently the same person.

The way this reaches valuation is a matter of substitution rather than, as is often assumed, a matter of trust. A buyer's or a lender's model rests on a handful of operating assumptions — capacity utilization, maintenance cost per unit, first-pass yield, commissioning duration, mean time between failures — and those assumptions enter the model as verified series, not as assertions. Where the series cannot be reproduced, the assumption is not removed from the model; the conservative band for the comparable asset class is inserted in its place. The resulting difference is arithmetic rather than punitive, and because it typically forms in the operating assumptions, the less visible layer of the model, well before it reaches the multiple, it is rarely named at the negotiating table.

The counterpart on the transaction side is more tangible. An earn-out constructed on an indicator whose definition has never been fixed will, in all likelihood, become a scope dispute after closing; good faith on both sides does not alter that outcome, since what is disputed is not intent but the boundary of the count. A diligence team that identifies this exposure will typically move in four directions: shifting part of the consideration into escrow, broadening the representation and warranty coverage relating to operating data, shortening the measurement window in order to contain the uncertainty, or requiring that the indicator definitions be agreed in writing as a condition precedent. Each of these instruments carries a price, and by the nature of transaction architecture that price is borne by the seller.

The third channel is continuity. An arrangement in which management information is produced inside one individual's working file is coded in diligence notes not as a measurement weakness but as key-person dependency, and that coding feeds directly into the integration plan, the retention agreements and, at times, the deferral of a portion of the consideration. In scaling, the same weakness presents a different face: a measurement practice that runs on intuition at a single site does not produce comparable numbers once there are three, and management consolidating three sites into one table is in substance consolidating three different definitions. A comparable mechanism operates in insurance and warranty provisioning, where failure and claim history that is not maintained in traceable form leaves premiums and reserves set against the class average rather than against the company's own record.

The mechanism that neutralizes this tendency is system design rather than individual attentiveness, and it separates into four components. The first is a metric dictionary: for each indicator, the formula, the source system and field, the scope boundary, the exception rules, the measurement frequency and the approving authority recorded in a single document kept current, approved and accessible to a review team. The second is source lineage — every reported figure traceable back to the raw record, with intermediate adjustments applied under a predefined rule rather than at the preparer's discretion. The third is a definition change log; when the scope of an indicator changes, the historical series is restated and the change recorded with its effective date, so that trend and definition remain distinguishable. The fourth is decision linkage: an indicator for which no threshold and no consequent decision has been written down remains a reporting line item.

BEIREK's intervention in this area is built on constructing the production chain rather than on increasing the number of indicators. A source-system mapping is typically prepared first for every figure in the existing management pack, with the dictionary written in the company's own terminology; ownership is then separated, so that the line producing an indicator and the line accountable for the performance it measures do not converge on one person. A reproduction exercise follows: the figures for a selected quarter are rebuilt from the raw source with the person who normally prepares them held out of the process, and the variance is analyzed item by item — an exercise that, in a single session, usually surfaces the entries the dictionary is still missing. The monthly review rhythm is then reset from a format in which a number is merely reported to one in which it is tied to a threshold and to a decision.

The value of a measurement system in an investment review lies less in the accuracy of the number it produces than in that number remaining the same irrespective of who produces it. Performance itself may depend on individuals, and in most companies does so to some degree; but for as long as the manner of measuring performance also depends on individuals, demonstrated results are priced as a temporary outcome rather than as institutional capacity. The question that determines valuation is therefore, more often than not, not how well the company runs, but whether a party other than the incumbent preparer could show how well it runs and arrive at the same answer.