---
title: "Performance Metrics: The Accuracy of the Number, or the Way It Is Produced?"
description: "In diligence, performance metrics are tested for reproducibility: whether the same figure can be rebuilt from the raw source, within a limited window, to the same result, with the person who normally prepares it held out of the exercise. Where it cannot, the assumption is not deleted from the model; a conservative class band replaces it, and the difference reappears as escrow, earn-out conditionality or broadened representations."
url: https://www.beirek.com/en/blog/performance-metrics-due-diligence
canonical: https://www.beirek.com/en/blog/performance-metrics-due-diligence
published: 2026-07-09
modified: 2026-07-09
category: "Technology & Engineering"
category_url: https://www.beirek.com/en/blog/category/technology-engineering
language: en-US
reading_time_minutes: 8
publisher: BEIREK LLC
publisher_url: https://www.beirek.com
license: "© BEIREK LLC — citation with attribution and link permitted"
keywords: ["performance metrics","KPI definition governance","operational due diligence","data lineage","key-person dependency","valuation discount","earn-out structuring"]
topics: ["Investment readiness and valuation review","Technology and engineering operations","Management reporting and KPI architecture","Transaction structuring and diligence findings"]
alternate_language_url: https://www.beirek.com/tr/blog/performance-metrics-due-diligence
---

# Performance Metrics: The Accuracy of the Number, or the Way It Is Produced?

> **In short:** In diligence, performance metrics are tested for reproducibility: whether the same figure can be rebuilt from the raw source, within a limited window, to the same result, with the person who normally prepares it held out of the exercise. Where it cannot, the assumption is not deleted from the model; a conservative class band replaces it, and the difference reappears as escrow, earn-out conditionality or broadened representations.

*In an investment review, technical performance indicators are assessed less on the accuracy of the numbers they produce than on the manner in which those numbers are produced. A metric set whose definitions are unwritten, whose lineage to the source system is broken, and which is assembled inside one individual's working file converts, in the model, into a conservative substitution and, in the transaction, into escrow and broadened representation coverage.*

---

In a monthly production review, the on-time delivery figure for the same period frequently reaches the table as two numbers rather than one — the rate cited by operations differing from the rate cited by the commercial side by a few percentage points — and the gap tends to close within a few minutes, not through the application of a written rule but through two people recalling, jointly and from memory, which orders were counted and which were set aside. Each side is correct within its own definition, so nothing is treated as an error and the agenda moves on. A diligence team seated in the same room, however, records not the figure but the manner in which the discrepancy was resolved, since what enters a valuation is not the month's rate but the mechanism by which that rate was produced.

Asked where the measure comes from, management typically answers with a system name and finishes with a person's name: the underlying data sits in MES or ERP, while the table that reaches the board is assembled at month-end inside a working file in which one individual pulls the raw extract and adjusts the known exceptions by hand. Most of those adjustments are defensible — work orders closed late in the system, shipments recorded twice, deferrals caused by the customer genuinely distort the raw feed, and leaving them uncorrected would produce a misleading table. The operational test applied in diligence is not whether the adjustments were warranted; the test is whether last quarter's figure can be reproduced from the raw source, within a few hours, to the same result, with the person who normally prepares that file absent from the exercise.

How such an arrangement comes into being follows a recognizable sequence across companies. Metrics rarely originate from an intention to design a measurement system; a customer complaint, a warranty claim, a lender's reporting requirement or a quality incident raises a specific question, and a measure is defined in order to answer it. The question loses its urgency over time; the measure remains. Leaving the definition unwritten is rational at the moment it is left unwritten, since a small team working in one location and sharing an intuition about what the indicator does and does not capture gains little from documenting it. The cost surfaces once the condition changes — a second shift, a second site, a newly hired engineer, a new customer segment — and what is missing at that point is not data but a shared meaning of the data.

Where the definition is unwritten, ambiguity resolves, predictably, in the direction that produces the least friction. Planned downtime is placed outside the scope, customer-caused delays are held separately, jobs extended by an engineering change request drop out of the count; each exclusion is defensible when examined on its own, and most have in fact been defended at the time. In aggregate, however, the series slowly loses the property of being comparable with its own history, and a three-year improvement curve becomes the sum of genuine improvement and a quietly narrowing definition. To the extent that an indicator also becomes an input to bonus calculation, budget setting or board assessment, the quantity of information it carries declines; where the measuring party and the measured party sit on the same line, that outcome follows independently of anyone's intent.

A second mechanism concerns the composition of the metric set. What has been instrumented gets measured; what has not been instrumented goes unmeasured even where it determines the result. For that reason the dashboard in many technical organizations reads as an archaeology of past crises rather than as a map of present value drivers — whatever once produced a serious incident still has its indicator in place, while rework rate, commissioning duration or design revision cycle, any of which may be setting today's margin, are not tracked at all. On the ownership side, an arrangement that would not survive a single audit question inside the finance function is treated as ordinary within engineering: the person who defines the indicator, the person who produces it, and the person whose performance is assessed against it are frequently the same person.

The way this reaches valuation is a matter of substitution rather than, as is often assumed, a matter of trust. A buyer's or a lender's model rests on a handful of operating assumptions — capacity utilization, maintenance cost per unit, first-pass yield, commissioning duration, mean time between failures — and those assumptions enter the model as verified series, not as assertions. Where the series cannot be reproduced, the assumption is not removed from the model; the conservative band for the comparable asset class is inserted in its place. The resulting difference is arithmetic rather than punitive, and because it typically forms in the operating assumptions, the less visible layer of the model, well before it reaches the multiple, it is rarely named at the negotiating table.

The counterpart on the transaction side is more tangible. An earn-out constructed on an indicator whose definition has never been fixed will, in all likelihood, become a scope dispute after closing; good faith on both sides does not alter that outcome, since what is disputed is not intent but the boundary of the count. A diligence team that identifies this exposure will typically move in four directions: shifting part of the consideration into escrow, broadening the representation and warranty coverage relating to operating data, shortening the measurement window in order to contain the uncertainty, or requiring that the indicator definitions be agreed in writing as a condition precedent. Each of these instruments carries a price, and by the nature of transaction architecture that price is borne by the seller.

The third channel is continuity. An arrangement in which management information is produced inside one individual's working file is coded in diligence notes not as a measurement weakness but as key-person dependency, and that coding feeds directly into the integration plan, the retention agreements and, at times, the deferral of a portion of the consideration. In scaling, the same weakness presents a different face: a measurement practice that runs on intuition at a single site does not produce comparable numbers once there are three, and management consolidating three sites into one table is in substance consolidating three different definitions. A comparable mechanism operates in insurance and warranty provisioning, where failure and claim history that is not maintained in traceable form leaves premiums and reserves set against the class average rather than against the company's own record.

The mechanism that neutralizes this tendency is system design rather than individual attentiveness, and it separates into four components. The first is a metric dictionary: for each indicator, the formula, the source system and field, the scope boundary, the exception rules, the measurement frequency and the approving authority recorded in a single document kept current, approved and accessible to a review team. The second is source lineage — every reported figure traceable back to the raw record, with intermediate adjustments applied under a predefined rule rather than at the preparer's discretion. The third is a definition change log; when the scope of an indicator changes, the historical series is restated and the change recorded with its effective date, so that trend and definition remain distinguishable. The fourth is decision linkage: an indicator for which no threshold and no consequent decision has been written down remains a reporting line item.

BEIREK's intervention in this area is built on constructing the production chain rather than on increasing the number of indicators. A source-system mapping is typically prepared first for every figure in the existing management pack, with the dictionary written in the company's own terminology; ownership is then separated, so that the line producing an indicator and the line accountable for the performance it measures do not converge on one person. A reproduction exercise follows: the figures for a selected quarter are rebuilt from the raw source with the person who normally prepares them held out of the process, and the variance is analyzed item by item — an exercise that, in a single session, usually surfaces the entries the dictionary is still missing. The monthly review rhythm is then reset from a format in which a number is merely reported to one in which it is tied to a threshold and to a decision.

The value of a measurement system in an investment review lies less in the accuracy of the number it produces than in that number remaining the same irrespective of who produces it. Performance itself may depend on individuals, and in most companies does so to some degree; but for as long as the manner of measuring performance also depends on individuals, demonstrated results are priced as a temporary outcome rather than as institutional capacity. The question that determines valuation is therefore, more often than not, not how well the company runs, but whether a party other than the incumbent preparer could show how well it runs and arrive at the same answer.

## Key Points

- The value of an indicator lies not in the sophistication of its definition but in that definition being written, approved and traceable down to the raw record.
- Absent a written definition, scope ambiguity resolves in the direction of least friction, and the series gradually loses the property of being comparable with its own history.
- An operating series that cannot be reproduced is not removed from the model; a conservative class average is substituted for it, and the resulting difference is borne by the seller.
- Where the person who defines an indicator, produces it and is assessed against it is the same person, diligence notes record key-person dependency rather than clear ownership.
- An indicator that has changed no decision across four consecutive quarters is not management information but a reporting cost.

## Questions

### How exactly are performance indicators tested in an investment review?

The test runs across three layers. The first examines whether a written and approved definition exists: formula, source system, scope boundary and exception rules. The second asks whether the indicator is actually used in daily operations — which decision it changes, and at which threshold. The third attempts reproduction: the figures for a selected period are rebuilt from the raw record with the person who normally reports that period held out, and the variance is analyzed.

### The company tracks KPIs but the definitions are not written down; why does that create a problem?

Absent a written definition, scope ambiguity resolves over time in the direction of least friction; planned downtime, customer-caused delays or jobs extended by revision are excluded from the count on individually defensible grounds. Each decision may be reasonable, yet in aggregate the series loses the property of being comparable with its own history. In diligence this is recorded not as an error but as an unverifiable historical series, which leads to conservative substitution in the model.

### Through which channel does weak performance measurement reach valuation?

Three channels operate. In the model, the conservative band of the class average replaces the operating series that cannot be verified, and the difference forms within the operating assumptions. In the transaction structure, indicators whose definitions are unfixed increase the escrow proportion, widen representation and warranty coverage and tighten earn-out conditionality. On the continuity side, numbers produced by a single individual are coded as key-person dependency and provide grounds for deferring part of the consideration.

### How is a performance measurement system shown to be independent of the founder?

Independence is evidenced by repetition rather than by assertion. The metric dictionary should be current and approved, every reported figure traceable to the raw record, intermediate adjustments governed by written rule rather than discretion, and definition changes logged together with a restatement of the historical series. On top of that, a reproduction attempt conducted with the customary preparer held out of the process demonstrates independence at the operational level rather than merely at the documentary level.

---

Source: https://www.beirek.com/en/blog/performance-metrics-due-diligence
Publisher: BEIREK LLC — https://www.beirek.com
