Examine the promotion and pay decisions a company makes in the three months following the close of its annual performance cycle, and a recurring pattern emerges: a meaningful share of those decisions does not map cleanly onto the individuals who received the highest ratings on the appraisal form. The merit list itself is internally coherent, defensible, and in most cases substantively correct — but its correctness derives from the founder's or general manager's accumulated judgment about who is actually carrying weight, not from the output of the review instrument. The form was completed, archived, and filed in the HR system, and it was not on the table at the moment of decision. A diligence team placing those two lists side by side reads the divergence not as a question of candor but as a question of architecture.
A second manifestation of the same divergence appears in the distribution of scores itself. When ratings issued by departmental managers are consolidated into a single table and one department's mean sits materially above another's without any operational correlate — no supporting difference in throughput, error rate, customer escalations, or delivery performance — what the instrument is measuring is not performance but the relationship each manager maintains with a team. The manager who rates strictly cannot secure raises for the people reporting to that function, the manager who rates generously captures a disproportionate share of whatever merit budget the system releases, and within two or three cycles this arithmetic becomes common knowledge. Once it does, scores compress upward, discriminating power disappears, and the system exhausts its own capacity to measure anything.
The mechanism underneath this behavior is not a weakness but, under the prevailing conditions, an entirely rational choice. For a line manager, the cost of assigning a low rating is immediate and borne personally: a difficult conversation, a damaged working relationship, unease propagating through a small team, and quite possibly a resignation. The cost of a generous rating is deferred and diffused — absorbed by the budget, by the company, by next year's planning cycle. Where the cost of a decision falls on the decision-maker while its benefit disperses across the institution, the direction in which the decision-maker will predictably lean is not difficult to anticipate. The same asymmetry sharpens further in any structure that compresses the development conversation and the compensation conversation into a single sitting, since expecting candid feedback from the person who determines a subordinate's raise loads two incompatible functions onto one table.
The founder's habit of deciding outside the system rests on comparable logic. Below a certain organizational scale, the founder's judgment genuinely constitutes the better information source — direct observation of who does what, no intermediary required, and an instrument that reports back, more slowly, something already known. The difficulty lies not in the shortcut itself but in its persistence after the underlying condition has changed: once headcount exceeds the founder's capacity for direct observation and the same method continues, the company no longer possesses a performance management system, it possesses one person's memory. On the balance sheet these two look identical; at the diligence table, the second does not appear at all.
What the reviewing party is looking for here is not the existence of the system but whether the system produces decisions. The date and approval signature on the policy document closes the first question; the second concerns how many individuals received a lower rating than in the prior cycle over the past two rounds, and what documented events those reductions rest on. The third question is more uncomfortable: can any statistical relationship be established between review output and the merit increase schedule, or were the two lists generated independently of one another? The fourth sits at the intersection of performance data and attrition data — do the ratings of departed employees differ meaningfully from those of retained employees, or are high-rated individuals concentrated among the leavers? That last intersection typically reveals whether the system is ceremonial faster than the other three combined.
Ownership is the dimension most frequently answered incorrectly in diligence. Asked who is responsible for the system, management almost invariably points to the HR function — yet HR's responsibility is to run the calendar, distribute the instrument, and maintain the records. The genuine owner of performance appraisal is the line manager who answers for the ratings issued to a team and for how those ratings correspond to operational outcomes. Where no such accountability mechanism exists — that is, where the inconsistency between one manager's ratings and that team's actual output is never discussed anywhere — the system is unowned, and the records of an unowned system amount, a year later, to archival volume and nothing more.
The continuity dimension reduces to a single question: had the founder participated in no performance decision for an entire year, would the same decisions have been produced? This question is rarely asked internally, because the answer is usually known and usually uncomfortable. Rather than posing it directly, the reviewing party tests it obliquely, probing in second-tier management interviews what proportion of each manager's compensation and promotion proposals for their own teams was accepted, on what grounds the rejected proposals were rejected, and whether those grounds exist in writing anywhere. The real boundary of delegated authority is visible not in the titles conferred but in the rejection rate of proposals originating one level down.
The valuation consequence of this deficiency typically surfaces in deal structure before it surfaces in the multiple. In files where founder dependency concentrates in the human capital line, the buyer's characteristic reflex is to condition a portion of the purchase price on the founder's post-closing tenure, to require a separate retention plan for key personnel, and to attach key-employee attrition to the representation and warranty package or to earn-out triggers. Each of these converts seller consideration from cash into a contingent receivable; the headline price appears preserved while the risk profile of that price shifts quietly. Part of the discount never becomes visible at all: the buyer carries the managerial turnover expected in the first eighteen months of an uninstitutionalized system as a transition cost line in its own model and buries that line inside the offer.
Structural remediation is built through decision architecture rather than individual awareness, and it comprises four separable components. The first is separation: the development conversation and the compensation conversation are not scheduled at the same point in the calendar, with at least one quarter placed between them, so that candor in feedback escapes the shadow of a pay negotiation. The second is calibration: departmental managers see one another's distributions in a joint session before finalizing individual ratings, and any manager whose mean departs from the others justifies that departure in that session. The third is the evidence chain: every rating must attach to at least one concrete event recorded within the period — work delivered, a target missed, an incident resolved, a customer response received — with the record captured at the moment of the event rather than at the moment of approval. The fourth is linkage: the relationship between review output and compensation or promotion decisions is set down in a written rule set, where departure from the rule is not prohibited but its rationale is documented.
BEIREK's intervention on this line is not the design of an appraisal instrument but the construction of a record that makes visible where the decision was actually taken. In practice, the review output of the last two cycles is first reconciled against the compensation, promotion, and attrition data of the same periods in a single table; that reconciliation demonstrates whether the system produces decisions without requiring anyone to argue the point. The calibration session is then bound to a calendar discipline and its output — which rating changed, and on what stated basis — is recorded, since that record is the single hardest piece of evidence to fabricate when diligence tests whether the system genuinely operates. In the third step, consistency between team performance and each line manager's own rating behavior is placed under a regular review rhythm, so that ownership migrates from HR to line management.
Building this architecture is not a software selection question, and it rarely produces results before the second cycle; the first cycle is consumed by learning calibration itself, while by the second the distribution of scores begins to carry genuine differentiation. Yet this is precisely what holds value at the diligence table: a single year's record is a statement of intent, whereas two consecutive years of mutually consistent records constitute evidence of capability. The distinction an investor is searching for lies between a company that has good people and a company that possesses a mechanism capable of recognizing and retaining good people; the first is the outcome of a period, the second is an institutional asset, and only the second is transferable.
Ultimately the performance appraisal system functions less as an instrument of talent policy than as an indicator of how a company records what it knows about itself. The distance between correct information residing in the founder's head and that information being institutionally held by the company is the same distance that determines, at the point of sale, how much of the consideration arrives as cash and how much arrives as contingency. Closing that distance costs two appraisal cycles; leaving it open costs, in most files, an amount already embedded in the structure of the transaction.
