# The Selected Regulator: Operator Endogeneity and the Closure of the Measurement Critique

Author: Nourizadeh, Moreno
ORCID: 0009-0006-0174-2585
Version: v1_2

---

## Abstract

A mature critical literature has established that metric-based management instruments cannot represent the substrates they act upon. Ashby's law of requisite variety bounds what any regulator can absorb; the Conant-Ashby theorem requires the regulator to model its system; Goodhart's law, Campbell's law, and the surrogation literature document the collapse of measures under targeting; the audit-society and reactivity literatures document the institutional consequences. Yet this literature, across its cybernetic, economic, and sociological branches, preserves a single repair it never closes: the substrate-calibrated human operator who supplements the blind instrument with judgement. The critique quarantines the operator by assumption, treating the population that reads the dashboard as an exogenous draw of competent persons handed an inadequate tool. This paper closes the repair. Drawing together the finance literature on managerial overconfidence and selection (Goel and Thakor, 2008; Gervais, Heaton and Odean, 2011; Campbell et al., 2011; Otto, 2014), the personnel-economics evidence on promotion (Lazear and Rosen, 1981; Benson, Li and Shue, 2019), and the social psychology of confidence expression (Anderson et al., 2012; Radzevick and Moore, 2011; Kennedy, Anderson and Moore, 2013), the paper argues that the operator channel is itself filtered by the same measurement apparatus whose limits the critique has proven, and that the filtering selects for calibration to the instrument rather than calibration to the substrate. The result is what I call operator endogeneity: the population operating the instrument is an output of the instrument, and the repair by human judgement fails as a class for the same reason the repair by better instruments fails as a class. A formal statement is developed as a perverted completion of the Conant-Ashby theorem: under measurement-driven promotion, every retained operator tends toward a model of the regulator rather than a model of the system. Three consequences follow. First, the common-knowledge structure attributed to institutionalised metric regimes stratifies: acknowledged inadequacy at the working layer coexists with sincere belief at the selected layer, because selection manufactures sincerity. Second, the individualising genres that dominate public discussion of competence, the Dunning-Kruger genre and the executive success narrative, operate as the regime's alibi, converting a selection effect into a personal defect or a personal virtue and deleting the selecting environment in both directions. Third, the same variance-harvesting tournament that staffs the firm reappears at the scale of capital allocation, where the selection of extreme belief is priced rather than promoted. The paper concludes that no reform conducted through the operator channel can succeed while the selection channel remains measurement-driven, and states what any reform would instead have to alter.

**Keywords:** requisite variety, Conant-Ashby theorem, Goodhart's law, managerial overconfidence, CEO selection, tournament theory, Peter Principle, promotion, Dunning-Kruger effect, organisational silence, management fashion, venture capital

---

## 1. Introduction: The Exemption Clause

The critical literature on measurement in organisations has, over five decades, assembled a result of unusual solidity. The result holds across at least four traditions that rarely cite one another. In cybernetics, a regulator can absorb only as much disturbance variety as it commands (Ashby, 1956), and can regulate well only insofar as it contains a model of the system it regulates (Conant and Ashby, 1970). In economics, any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes (Goodhart, 1975), a proposition anticipated for social indicators by Campbell (1976) and compressed by Strathern (1997) into the aphorism that a measure which becomes a target ceases to be a good measure. In management accounting, the mechanism has a name, surrogation, and an experimental demonstration: the metric substitutes for the construct it proxies, the proxy becomes the target, and the target is met by spending the construct (Choi, Hecht and Tayler, 2012). In sociology, the audit explosion (Power, 1997), the reactivity of rankings (Espeland and Sauder, 2007), and the performativity of models (MacKenzie, 2006) document what happens when whole institutional fields reorganise themselves around instruments of this kind, and Muller (2018) surveys the cross-sectoral record. The traditions converge on a single proposition: the instrument cannot carry what the organisation runs on, and acting through the instrument consumes what the instrument cannot see.

Read the canonical statements of this critique closely, however, and a recurring clause appears at the point where the argument might otherwise implicate persons. The clause exempts the operator. The limit is located in the instrument; the men and women who read it are held to be competent, well-intentioned, and analytically separable from the apparatus they operate. Vaughan's (1996) reconstruction of the Challenger launch decision insists that the engineers and managers were conforming to structure, and that the disaster required no amoral calculators. The surrogation experiments are run on ordinary subjects to show that the substitution requires no defect of character. The audit-society literature treats auditors as carriers of a ritual whose logic exceeds them. The clause does honourable analytical work: it keeps the critique architectural, blocks the conspiratorial misreading, and matches the phenomenology of institutional life, in which each participant does the work the role specifies while the aggregate does something no participant chose. This paper does not dispute the clause's local truth. It disputes what the clause silently smuggles in.

What the exemption smuggles in is an assumption about the operator population: that it is exogenous to the instrument. The critique models the situation as a random draw of capable humans handed a blind tool. Under that model, one repair survives everything the critique proves. Grant that the dashboard cannot represent tacit coordination, relational structure, or slow substrates; a substrate-calibrated human can still supplement it. The executive can walk the floor. The board can weigh the numbers against judgement. The senior figure who knows the work can read the play behind the score. This repair is not a straw position; it is the standing recommendation of the practitioner literature that has absorbed the critique, from management-by-walking-around to the injunction that leaders must attend to what the metrics miss, and it is licensed by the critique's own exemption clause, since a population of competent operators exogenous to the instrument could in principle correct for it. The critique has closed the instrument channel; it has left the operator channel open.

This paper closes the operator channel. The closing thesis, stated at once, runs as follows. The operator population is not exogenous to the instrument. Ascent through an organisation is itself a measurement-driven allocation conducted through the same per-role, short-cycle apparatus whose representational limits the critique has established, and what that allocation selects for is legibility to the apparatus: performance on the metrics the apparatus can read, and, at the interpersonal layer where metrics run out, the expressed confidence that evaluators read as competence. The selection is not an error the firm commits against its own interest. A body of formal and empirical work in financial economics shows that value-maximising governance deliberately advances the overconfident (Goel and Thakor, 2008), optimally hires them (Gervais, Heaton and Odean, 2011), retains the moderately optimistic while dismissing the calibrated and the extreme alike (Campbell et al., 2011), and prices their beliefs into cheaper contracts (Otto, 2014). A parallel body of personnel economics shows that firms knowingly promote on current-role output at the documented expense of managerial capability (Benson, Li and Shue, 2019). A parallel body of social psychology shows that overconfidence purchases status because it produces the behavioural signature evaluators use as their cue for competence, while the actually competent frequently do not display the cue (Anderson et al., 2012). The three bodies of work describe one filter operating at three registers. The filter's output is an operator stratum calibrated to the instrument.

The consequence for the measurement critique is a completion it has lacked. The Conant-Ashby theorem states that every good regulator of a system must be a model of that system. The selection dynamics described here enact a perverted twin of the theorem, which I state informally now and formally in section 7: under measurement-driven promotion, every retained operator tends toward a model of the regulator. The population that would have to supply the supplementing judgement, the population calibrated to the substrate rather than to the instrument, is the population the filter attenuates, because substrate calibration presents inside the tournament as hesitancy, as negativity, and as numbers that refuse to move. The repair by human judgement therefore fails as a class, and it fails for the same structural reason the repair by better instruments fails as a class: both repairs are conducted through a channel the apparatus itself governs. I call the general condition operator endogeneity, and the completed result the selected regulator. A reformulation developed in section 7 states the same result with the vocabulary of bias removed: since calibration is a two-place relation, the trait the literature measures as overconfidence against the substrate is, against the instrument, a fit, and the filter is best described as allocating the operator stratum's modelling capacity to the regulator's horizon rather than as advancing an error.

Three further results follow from the closure, and the paper develops each. The first concerns the epistemic structure of institutionalised metric regimes. The literature on such regimes, from Soviet production statistics (Berliner, 1957; Kornai, 1992) through the normalisation of deviance at NASA (Vaughan, 1996) to the pre-2008 career of Value-at-Risk (MacKenzie and Spears, 2014), models them as common-knowledge configurations: everyone knows the framework is inadequate, everyone knows that everyone knows, and the framework persists anyway. Operator endogeneity forces a correction. Recognition of the framework's inadequacy is not uniform across the hierarchy; it thins with height, because ascent is filtered against it, and the sorting mechanisms of belief homogenisation (Van den Steen, 2005; 2010) plus the upward attenuation of disconfirming signal (Rosen and Tesser, 1970; Morrison and Milliken, 2000) complete what selection begins. At the working layer the inadequacy is acknowledged; at the selected layer it is sincerely invisible. The regime's stability then requires no cynicism anywhere, which resolves a residual difficulty in the common-knowledge accounts. The second result concerns the genres through which public discussion processes questions of competence. The Dunning-Kruger genre locates miscalibration in the individual skull; the executive success narrative locates achievement in the individual will. The paper argues that the two genres perform the same deletion in opposite directions, removing the selecting environment from the explanation, and that the deletion functions as the regime's alibi: it converts the filter's output into evidence of personal defect below and personal virtue above, so that the filter itself never appears in the story. The recent statistical dismantling of the Dunning-Kruger effect (Krueger and Mueller, 2002; Nuhfer et al., 2016; 2017; Gignac and Zajenkowski, 2020) is read here not as a curiosity of psychometrics but as the removal of the alibi's empirical warrant. The third result concerns scale. The variance tournament that staffs the firm reappears in capital allocation, where the selection of extreme belief is conducted by pricing rather than promotion; the paper reads the venture market's power-law payoff structure (Kerr, Nanda and Rhodes-Kropf, 2014), its documented overvaluation mechanics (Gornall and Strebulaev, 2020), and a contemporary case in point through the same filter logic, and connects the result to the pay-for-luck evidence that the reward system decouples from contribution wherever monitoring weakens (Bertrand and Mullainathan, 2001; Daniel, Li and Naveen, 2020; Andreani, Dou, Xu and You, 2025).

A remark on what the argument does not claim, before the exposition begins. The argument does not claim that overconfident managers destroy value in every instance; the finance literature is explicit that moderate overconfidence can offset risk aversion and raise investment toward first-best levels (Goel and Thakor, 2008; Campbell et al., 2011; Hirshleifer, Low and Teoh, 2012), and the argument concedes this throughout. The claim concerns composition, not the welfare effect of any single trait in any single decision. Whatever the per-decision value of selected traits, the selection principle determines who occupies the stratum from which supplementing judgement would have to come, and it is the composition of that stratum, not the sign of any decision, that decides whether the operator repair can work. The argument likewise does not claim that evaluators are foolish to read confidence as competence; under the information conditions evaluators face, the cue is the best signal available, which is exactly why the filter operates without requiring anyone's defect. The argument, throughout, is about what a selection principle does to a population over iterations, and about what the resulting population can and cannot supply.

The paper proceeds as follows. Section 2 assembles the standing results of the measurement critique in the form the later argument requires. Section 3 isolates the surviving repair and documents the exemption clause that licenses it. Sections 4 through 6 assemble the selection evidence at three registers: the interpersonal reading of confidence, the promotion tournament, and the deliberate design of selection in governance. Section 7 states the formal closure. Section 8 develops the stratification of recognition. Section 9 analyses the individualising genres as alibi. Section 10 carries the filter to the scale of capital allocation. Section 11 answers objections, section 12 draws the implications for reform, and section 13 concludes.

## 2. The Measurement Critique and Its Standing Results

The argument requires four results from the existing critique, each stated here in the form that will bear load later. None of the four is original to this paper; the paper's contribution begins where they stop.

### 2.1 The variety bound and the model requirement

Ashby's law of requisite variety states that a regulator can hold a system's essential variables within bounds only if the regulator's own variety, the count of distinct states it can assume and deploy, at least matches the variety of the disturbances it must absorb; in Ashby's compression, only variety can destroy variety (Ashby, 1956, p. 207). The law is a counting identity rather than an empirical conjecture: where disturbances present more distinguishable states than the regulator can distinguish and answer differentially, some distinctions map to a single response and control over the difference lapses by construction. Conant and Ashby (1970) strengthened the law into a condition on the regulator's internal organisation: any regulator that is maximally both successful and simple must be isomorphic with the system being regulated, so that good regulation entails that the regulator embody a model of its system. The two results compose into the criterion this paper will pervert in section 7. A regulator with variety but no model has states and no map; a regulator with a model but no variety has a map and no states; regulation requires both, and the model requirement fixes what the regulator's internal organisation must be a model *of*, namely the system, the substrate, the thing regulated.

A temporal corollary, implicit in Ashby and developed across the later organisational-cybernetics literature, matters for the argument's promotion application. A regulator that acts on its account of the system every reporting cycle, and whose account is updated and rewarded only at that cycle, carries an operative model whose resolution cannot exceed the cycle; disturbances and accumulations that mature over horizons long against the cycle appear within any single cycle as approximately constant inputs the account cannot distinguish, a bound with its exact form in the rate-distortion theory of lossy source coding (Shannon, 1959; Cover and Thomas, 2006, ch. 10). The corollary will reappear as the temporal face of the promotion filter: an evaluation conducted quarterly cannot register substrate contributions that accrue over years, whatever the goodwill of the evaluator, and so a promotion system keyed to such evaluation cannot select on them.

### 2.2 Collapse under targeting

Goodhart's (1975) observation that any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes, Campbell's (1976) parallel law for social indicators, and Strathern's (1997) canonical compression together establish that the representational limits of section 2.1 are not static. An instrument adequate as description degrades as target, because the regulated population reallocates behaviour toward what the instrument reads. The surrogation experiments give the mechanism its psychology: under incentive compensation tied to a measure, decision-makers come to treat the measure as the construct, and do so even when the substitution is pointed out to them (Choi, Hecht and Tayler, 2012). The collapse result matters here for one reason above all: it applies to the evaluation of persons exactly as it applies to the evaluation of processes. A promotion criterion is a measure placed under the most intense targeting pressure an organisation contains, since careers depend on it; by Goodhart's dynamic, whatever the criterion once tracked, the population ascending under it converges on the criterion itself. Section 5 will document that this is what the promotion evidence shows.

### 2.3 Reactivity and performativity

Espeland and Sauder (2007) demonstrated, in the case of law-school rankings, that public measures do not report a field so much as reorganise it: actors internalise the measure's categories, redistribute effort toward its inputs, and the field comes to resemble the measure's picture of it. MacKenzie (2006) established the stronger form for financial models, an engine rather than a camera: the model participates in producing the price patterns it purports to describe. The reactivity results extend the critique from representation to production, and they matter here because a promotion criterion is reactive in exactly this sense. A criterion that reads confidence as competence does not simply misread a fixed population; it teaches the population to produce the reading, redistributing self-presentation toward the cue (a dynamic section 4 documents experimentally in Radzevick and Moore, 2011). The instrument manufactures the trait it then selects.

### 2.4 The tacit remainder

Polanyi (1958; 1966) established that the knowledge on which skilled performance runs exceeds what its bearer can articulate, that the articulable rides on a tacit base transmitted by apprenticeship and proximity rather than by prescription, and that forcing the tacit into explicit form does not transfer it. The organisational consequence, developed across the literature from Nelson and Winter (1982) through the knowledge-management corrective wave, is that the substrate an organisation runs on, the mutual models, the calibration of colleagues to one another's blind spots, the sense of how the product actually behaves, does not survive transcription into any per-role record. The consequence for this paper is a sharpening of what "substrate calibration" means when predicated of a person: it names a stock accumulated through years of proximity to the work, carried tacitly, invisible to per-role instruments by the nature of the stock rather than by any deficiency of instrument design. The promotion filter of section 5 cannot select on this stock even in principle, because the stock is not in the record the filter reads; the filter can select only on what the record carries, and on the interpersonal cues that fill the record's gaps.

These four results define the situation the exemption clause then interprets. The instrument cannot carry the substrate (2.1, 2.4); it degrades further under targeting (2.2); it remakes what it reads (2.3). The critique concludes that the instrument cannot be repaired from inside the instrument class: a richer dashboard, a longer-horizon metric, an added leading indicator is another instrument of the same class and inherits the same bounds. The conclusion is correct and this paper relies on it. The question is what remains once the instrument class is closed, and the standing answer is: the operator.

## 3. The Surviving Repair: The Supplementing Operator

### 3.1 The repair stated at full strength

The repair deserves its strongest formulation, because the paper's claim is that even the strongest form fails, and a weak formulation would make the closure trivial. The strong form runs as follows. Concede every result of section 2. The instrument is blind to relational structure, tacit stocks, and slow substrates; it collapses under targeting; it remakes what it reads. None of this entails organisational blindness, because organisations are not run by instruments; they are run by people who use instruments. A person of judgement holds the number and the work together. She reads the velocity figure and also reads the codebase; she sees the cross-sell metric and also sees the branch; she receives the quarterly report and also walks the floor, in the practice Peters and Waterman (1982) canonised and a long practitioner tradition sustains. Her judgement is exactly the substrate model the Conant-Ashby theorem demands and the instrument lacks; supplied by her, the coupled system of instrument-plus-operator can satisfy the model requirement even though the instrument alone cannot. On this account the critique's results are real and bounded: they show what dashboards cannot do, and thereby show what leaders must do. The practical literature that has metabolised the measurement critique almost always lands here. Muller (2018) closes with a checklist for the judicious use of metrics by judicious people; the balanced-scorecard tradition (Kaplan and Norton, 1996) responds to the narrowness of financial measures with a wider instrument read by wiser readers; the recurring executive injunction to attend to what the numbers miss presupposes an executive who can.

The repair has a second, institutional form, which also deserves statement. Where the first form supplements the instrument with an individual's judgement, the second supplements it with a body's: the board, the review committee, the promotion panel. The panel hears the metrics and also hears testimony; it can, in principle, weigh the record against what the record cannot carry. Governance design in the post-crisis period leans on this form heavily, multiplying committees whose stated function is to supply the judgement instruments lack.

### 3.2 The exemption clause that licenses it

Both forms of the repair rest on one assumption, and the assumption is visible in the critique's own texts at the moment they exempt the operator. The assumption is exogeneity: the judgement-bearing population exists independently of the instrument regime and is available to be deployed against it. Vaughan (1996, pp. 394-399) closes her reconstruction by relocating causation from individual failing to organisational structure, and in doing so leaves the individual's competence intact and unexplained, an input to the analysis rather than an output of the system analysed. Power (1997, p. 123) describes auditors as operating within rituals of verification whose logic exceeds any auditor, again holding the person analytically apart from the apparatus. The surrogation literature demonstrates that ordinary, unimpaired subjects surrogate, which establishes that the pathology needs no defective person, and in establishing this, once more treats the person's formation as outside the story. In each case the exemption performs indispensable work against the moralising misreading. In each case it also fixes the operator population as a given.

Once fixed as a given, the population can carry the repair. If competent, substrate-calibrated persons exist in the ordinary distribution and rise through organisations in the ordinary way, then at every altitude of the hierarchy there are readers who can supplement the instrument, and the critique's results, however strong, describe a tool problem with a staffing solution. The entire weight of the repair rests on the phrase "rise through organisations in the ordinary way". The next three sections examine what the ordinary way is.

## 4. Confidence as Signal: The Interpersonal Reading of Competence

The promotion filter operates on two inputs: the record the instrument carries, and the impression the candidate makes where the record runs out. This section concerns the second input. The finding it assembles from three decades of judgement research can be stated in one sentence: interpersonal evaluation reads expressed confidence as competence, rewards it with status and influence regardless of its calibration, punishes its absence, and thereby installs, at the heart of every evaluative encounter the record does not settle, a cue that selects on self-assessment rather than on ability.

### 4.1 Conceptual hygiene: three overconfidences

The term overconfidence covers three dissociable phenomena, and the argument requires them kept apart (Moore and Healy, 2008). Overestimation names the belief that one's absolute performance exceeds what it does. Overplacement names the belief that one ranks above others when one does not. Overprecision names excessive certainty in the accuracy of one's beliefs, the assignment of too-narrow confidence intervals. The selection results below run principally on overplacement and overprecision: the tournament of section 5 selects on beliefs about relative standing, and the interpersonal cue of the present section transmits certainty. Overprecision matters additionally because it is the variant with the strongest field evidence at the executive level: Ben-David, Graham and Harvey (2013) found that the 80 per cent confidence intervals of chief financial officers, elicited across more than 13,000 forecasts of market returns, contained the realised outcome 36 per cent of the time, a miscalibration that persisted across a decade of quarterly feedback. The persistence under feedback deserves emphasis, because it forecloses the response that experience would calibrate the selected: the selected population's miscalibration survives exactly the kind of repeated, consequential feedback that calibration theories predict should extinguish it.

### 4.2 The status returns to overconfidence

Anderson, Brion, Moore and Kennedy (2012) tested, across six studies, what they named the status-enhancement account of overconfidence: the hypothesis that overconfidence pervades self-judgement because it purchases social status. Three findings from that programme carry the present argument. First, overconfident individuals attained higher status, measured as respect and influence accorded by peers, in both short-lived and longer-term task groups, and did so regardless of whether their confidence was justified by measured ability. Second, and decisive for the filter argument, a Brunswikian lens analysis identified the pathway: overconfidence produces a behavioural signature, more speaking time, a confident vocal tone, earlier and more frequent answer offering, an expansive posture, that observers use as their cue for competence, so that the overconfident appear competent through the display. Third, and still more damaging to the repair thesis, the cue-display correlations for actual competence were null: the actually competent did not display the behavioural cues that signal competence to others. The instrumented reading of persons, in other words, is not a noisy reading of ability; it is a clean reading of self-assessment, and self-assessment and ability come apart by construction wherever miscalibration exists. Anderson and Kilduff (2009) had established the antecedent result that dominance-related behaviour predicts influence attainment independent of competence; Kennedy, Anderson and Moore (2013) then closed the escape route the 2012 findings left open, showing that the status returns to overconfidence survive even when the group later learns the overconfident individual's true performance: the discovered overconfident were not penalised down to the level their ability warranted, because the confident display had already been encoded as a competence attribution the disconfirmation did not fully reverse.

### 4.3 The market for confidence

Price and Stone (2004) named the underlying heuristic: evaluators treat an adviser's confidence as a proxy for the adviser's knowledge, preferring the more confident of two advisers even when the confidence is uninformative. Sah, Moore and MacCoun (2013) established the interaction with accuracy: confidence increases an adviser's credibility and persuasiveness, calibration information moderates the effect only weakly, and advisers face incentives to inflate. Radzevick and Moore (2011) supplied the market dynamics, and their result converts the heuristic from a static bias into a selection engine. In experimental markets where advisers competed to sell judgement, advisers whose services were chosen more often were the more confident, not the more accurate; competition between advisers escalated expressed confidence over rounds; and the escalation persisted when advisers varied in underlying quality, so that the market's selection pressure operated on the expression rather than on the substance. Tenney, Meikle, Hunsaker, Moore and Anderson (2019) added the boundary condition: the social penalty for exposed overconfidence attaches to explicit verbal claims more than to the nonverbal display, which means the display channel, the channel the lens analysis identified as carrying the competence attribution, is close to penalty-free. Tetlock (2005), across two decades of real-world expert prediction, documented the field version: the experts most in demand from media and institutions, the confident single-framework hedgehogs, were systematically the worst calibrated, while the tentative, many-model foxes predicted better and were consulted less.

### 4.4 What the interpersonal layer contributes to the filter

Assembled, the judgement literature specifies the second input to the promotion filter with unpleasant exactness. Where the record does not settle an evaluation, the evaluation is settled by a cue; the cue is expressed confidence; the cue is read as competence; the cue is displayed by the miscalibrated and, by the null cue-display result, frequently not displayed by the able; competition amplifies the display; exposure does not fully claw back its returns; and the display channel that carries the attribution is the channel least exposed to sanction. An evaluator inside this structure is not behaving foolishly. Under the information conditions of most evaluations, no better cue is available at the point of decision, which is why the structure needs no defect anywhere to operate. But a selection system that iterates this cue over careers is running a filter on self-assessment error, and the direction of the filter is fixed: it advances overplacement and overprecision, and it attenuates the accurate self-assessor, whose calibrated hedging the cue reads as incapacity. The substrate-calibrated person of section 3, the one whose judgement the repair requires, is calibrated in exactly the register the cue punishes.

## 5. The Promotion Filter

The first input to the filter is the record, and the record is read through a tournament. This section assembles the personnel-economics and organisational-sociology evidence that ascent through the record channel selects on legibility to the instrument, with the same direction of bias the interpersonal channel supplies.

### 5.1 Tournaments and the reward for variance

Lazear and Rosen (1981) established the canonical result that rank-order tournaments can implement efficient effort: where individual output is costly to measure exactly, paying for rank rather than level economises on measurement and sustains incentives. The design carries a known side effect that the subsequent literature developed at length: because a tournament pays for finishing first, and because a contestant behind in expectation raises her win probability by raising the variance of her outcome, tournaments reward variance-seeking, a dynamic documented in settings from mutual-fund interim rankings (Brown, Harlow and Starks, 1996; Chevalier and Ellison, 1997) onward. The variance property is the bridge between tournament design and belief selection, because a contestant's willingness to take variance is a function of what she believes about herself. The overplacer enters more contests and takes larger gambles because she expects to win them (Camerer and Lovallo, 1999, whose experimental entrants produced systematic excess entry exactly when payoffs depended on relative skill); the overprecise mistakes her estimate's reliability and bets accordingly. Across repeated tournament rounds, the winners' circle therefore fills disproportionately with the miscalibrated, not because miscalibration causes ability, but because miscalibration causes the risk posture that rank-order selection rewards, and because the tournament observes outcomes, never the calibration that produced the position taken.

### 5.2 The deliberate promotion of the overconfident

Goel and Thakor (2008) converted this dynamic from an incidental bias into a theorem about value-maximising governance, and their result is the load-bearing formal object of the present section. In their model, managers of a priori unknown ability compete in an internal promotion tournament for the chief executive position; overconfident managers, meaning managers who underestimate the riskiness of their projects, choose riskier projects than their rational peers; and because the tournament promotes the best observed outcome, and risk-taking spreads the outcome distribution, an overconfident manager has a strictly higher probability than an equally able rational manager of being promoted, under governance that is maximising firm value and knows everything the modeller knows. The qualifier deserves the emphasis Goel and Thakor give it: the promotion of the overconfident is not a governance lapse the model diagnoses; it is the equilibrium behaviour of correctly designed internal governance, because the same underestimation of risk that distorts the overconfident manager's project selection also offsets managerial risk aversion, so that a moderately overconfident chief executive raises firm value relative to a rational one, the effect turning negative only beyond an interior optimum. The model yields the composition result the present paper needs in exact form: the population reaching the top of the record channel is enriched in overconfidence relative to the population entering it, by the operation of the selection mechanism itself, under the mechanism's own objective function. Subsequent empirical work confirms the direction: overconfident managers, identified through option-exercise behaviour, are more likely to be promoted to chief executive positions than rational ones, and Malmendier and Tate's (2005; 2008) foundational measurement work, classifying as overconfident those executives who persistently hold deep-in-the-money options on their own firm, established that the classification predicts investment distortion, cash-flow sensitivity, and value-destroying acquisition in the field. A preregistered experimental result closes, from below, the causal gap the observational promotion studies leave open. Weidmann, Vecci, Said, Deming and Bhalotra (2024), randomising managers repeatedly across teams to isolate managerial contribution from team composition, find that participants who nominate themselves for the manager position produce worse team performance than managers appointed by lottery, in part because the self-nominated are overconfident, above all about their social skills, while actual managerial contribution is predicted by measured decision-making skill and fluid intelligence, neither of which the self-selection channel tracks. The entry margin into authority is therefore adversely selected before any organisational record exists: the population that steps forward is already enriched in the trait the subsequent filter amplifies, and a lottery would outperform the volunteers.

### 5.3 The Peter Principle with microdata

Benson, Li and Shue (2019) supplied the record channel's other jaw. Using microdata on 53,035 sales workers across 131 firms, they provided the first large-scale test of the Peter Principle and found it: promotion probability rises steeply with current-role sales performance, current-role sales performance negatively predicts subsequent managerial value-added, and the observable characteristic that does predict managerial performance, collaboration experience, is discounted in promotion decisions relative to the sales figure. The firms in their data forgo up to 30 per cent of the managerial value-added a potential-based promotion rule would capture, and the authors' interpretation matters as much as the estimate: the pattern is consistent with firms knowingly accepting managerial mismatch as the price of preserving the tournament's incentive effect on the current role. Lazear (2004) had shown that part of the post-promotion decline is mechanical, the regression to the mean of a noisy signal selected on its maximum; the Benson, Li and Shue evidence shows that beyond the mechanical component, the promotion rule itself weights the instrument-legible number over the instrument-illegible capability, deliberately, as policy. The record channel, in other words, does more than fail to observe substrate calibration (section 2.4 established that it cannot); where a proxy for the substrate-relevant capability is observable, the channel discounts it in favour of the metric.

### 5.4 Homosocial reproduction and credential gating

Two sociological mechanisms complete the record channel's specification, both operating where the record and the interpersonal cue leave residual discretion. Kanter (1977) documented, under the name homosocial reproduction, the tendency of managers facing uncertainty about candidates to resolve it by selecting persons like themselves, on the reasoning that similarity is the only available warrant of trustworthiness where performance is hard to verify; the mechanism converts each cohort of the selected into the template for the next, compounding whatever composition the tournament has already produced. Collins (1979) and Abbott (1988) documented credential gating: evaluative systems weight certified standing over demonstrated content, so that claims arriving without the recognised credential are discounted regardless of their substance. Credential gating matters for the closure argument because it governs the reception of substrate knowledge specifically. The substrate-calibrated speak from experience of the work, a warrant the credential system does not register; their claims arrive, in the terms of the gate, from unaccredited sources, and are attenuated accordingly, while equivalent claims arriving under fashionable credentials pass. The gate thereby completes, at the level of discourse, the same selection the tournament conducts at the level of position: what ascends is what the apparatus can read.

### 5.5 The filter assembled

The two inputs now compose. The record channel promotes on instrument-legible output, discounts the observable proxies of managerial capability, rewards the variance posture of the miscalibrated, and does so as deliberate policy under value-maximising governance. The interpersonal channel, operating wherever the record runs out, promotes on a confidence display that tracks self-assessment error and that the able frequently fail to produce. Homosocial reproduction compounds the composition across cohorts; credential gating silences the substrate register in which the filtered-out would testify. Every stage selects on calibration to the apparatus, none on calibration to the substrate, and the stages are serial, so their selection effects multiply. The population emerging at altitude is therefore not a random draw of the able; it is the filter's image, and the assumption of operator exogeneity on which section 3's repair rests is empirically false in the only direction that matters: the operators available to supplement the instrument are exactly those the instrument has chosen.

## 6. Selection by Design: Governance Prices the Beliefs It Promotes

A natural objection to section 5 holds that the filter, once named, would be corrected: boards and governance systems, observing that promotion selects miscalibration, would counteract it. The finance literature answers the objection with a stronger and stranger finding: governance does observe the selection, and ratifies it, because within the frame governance optimises over, the selected beliefs are an asset. This section assembles the ratification evidence, which converts the filter from an anomaly awaiting repair into an equilibrium.

### 6.1 Optimal hiring and cheap conviction

Gervais, Heaton and Odean (2011) modelled the contracting problem directly. A risk-averse manager's overconfidence makes him less conservative; it is therefore cheaper for the firm to motivate him toward valuable risky projects, and moderate overconfidence leads firms to offer flatter compensation that leaves both parties better off within the model's terms. Overconfident managers are, in the authors' phrase, more attractive to firms than their rational counterparts, in part because overconfidence commits them to effort. Otto (2014) tested the pricing implication in the field across 1,766 firms: optimistic chief executives, identified through option-exercise and forecast behaviour, receive smaller incentive grants and lower total compensation than their peers, consistent with firms exploiting the optimists' overvaluation of performance-contingent pay. The firm, in other words, does not simply tolerate the selected belief structure; it prices it, purchasing conviction below the market rate for calibration. Campbell, Gallmeyer, Johnson, Rutherford and Stanley (2011) established the retention face of the same policy: forced chief-executive turnover is elevated at both tails of the optimism distribution and minimised at moderate optimism, exactly as a board maximising shareholder value should behave given that moderate optimism offsets risk aversion toward first-best investment. The board's revealed preference is for the interior of the miscalibration distribution, and against calibration as such. Hirshleifer, Low and Teoh (2012) added the innovation register: overconfident chief executives invest more in research and development, obtain more patents and citations per research dollar in innovative industries, and achieve greater innovative output for given expenditure, evidence that the selected trait pays along at least one dimension the market prices.

### 6.2 Sincerity as a product of selection

The measurement work underlying this literature carries an implication for section 8 that must be registered now. Malmendier and Tate's Longholder classification identifies overconfidence through executives' persistent retention of exercisable in-the-money options on their own firms, that is, through voluntary, personally costly exposure of the executive's private wealth to the beliefs in question. The classification therefore certifies sincerity: the selected do not stop at displaying conviction for evaluators; they bet their own portfolios on it, holding undiversified positions a calibrated actor would unwind. Whatever else the filter produces, it produces believers. The point forecloses, in advance, the reading of executive belief as strategic performance, and it supplies the microfoundation for the stratification result of section 8: a selection system that advances sincere miscalibration manufactures a stratum for whom the instrument's picture of the world is not a convenient fiction maintained cynically, but the world.

### 6.3 Decoupled reward and the weakness of the outcome check

The last component of ratification concerns the feedback that would, on the repair account, eventually discipline the selected: outcomes. The pay-for-luck literature documents how weakly outcomes bind. Bertrand and Mullainathan (2001) established that chief-executive pay responds to observable luck, industry and macroeconomic movements outside any executive's control, as strongly as to general performance, and that the sensitivity concentrates in weakly governed firms. Daniel, Li and Naveen (2020) documented the asymmetry: executives are rewarded for good luck and penalised for bad luck to a smaller degree. Andreani, Dou, Xu and You (2025), exploiting the one-off windfalls of the 2017 United States tax legislation as clean luck, confirmed that weakly monitored chief executives capture pay from outcomes their decisions did not produce. The literature's import for the present argument is not the injustice of the pattern but its informational consequence: where reward decouples from contribution, the outcome record loses its power to identify the selected miscalibration, because the selected can accumulate favourable records out of variance and luck that the reward system books as performance. The outcome check that the repair account relies upon to eventually surface calibration operates, in the documented governance environment, at a lag and a noise level that let a variance-selected cohort clear it for the length of a career. Section 10 shows the same decoupling operating at market scale.

### 6.4 The equilibrium stated

Sections 4 through 6 together specify an equilibrium, and it is worth stating as one object before the formal treatment. Evaluators read confidence as competence because no better cue is available at the point of decision (section 4). Tournaments promote the outcome distribution's right tail, which variance-takers and hence the miscalibrated disproportionately occupy (section 5). Governance ratifies the resulting composition because, within the quarterly, instrument-legible frame governance optimises over, moderate miscalibration is a priced asset: cheaper to incentivise, offsetting of risk aversion, productive of innovation output, retained by boards acting on shareholder value (section 6). Every agent in the chain behaves defensibly under its local information and objective. The aggregate is a staffing system that concentrates, at exactly the altitudes from which supplementing judgement would have to come, a population selected for calibration to the instrument and against calibration to the substrate, sincere in its beliefs, and insulated by reward-decoupling from the outcome feedback that would otherwise expose the selection. No conspiracy appears anywhere in the chain, which is why the exemption clause of section 3.2, true of each agent, is false of the system: the operators are not exogenous to the instrument; they are its most consequential product.

## 7. The Formal Closure: A Perverted Conant-Ashby

This section states the argument's core result in a form exact enough to be attacked. The register is that of a theorem-sketch composed from established components; the components are cited, the composition is mine, and the composition's analytical responsibility is mine alone.

### 7.1 Definitions

Let an organisation comprise a substrate $S$, the productive complex of tacit stocks, relational coordination, and slow-cycle capabilities established in section 2.4, and a regulator $R$, the formal measurement-and-control apparatus of section 2.1, operating at action-cycle $T_R$ against substrate cycles $T_S \gg T_R$. Let the operator stratum $O$ denote the population of persons who read $R$'s outputs and exercise the discretion $R$ does not mechanise: managers, executives, boards. Let a person's calibration be the correspondence between her internal model and a target; call her substrate-calibrated to the degree her internal model tracks $S$ (its cycles, its tacit stocks, its relational structure), and instrument-calibrated to the degree her internal model tracks $R$ (its metrics, its cycles, its vocabulary, the movements of its numbers). The two calibrations are distinct targets of one modelling capacity, and nothing prevents an individual holding both; the argument concerns which of them selection can see.

Define a promotion system as measurement-driven when the allocation of persons across $O$'s positions is a function principally of ($i$) records legible to $R$ and ($ii$) evaluators' interpersonal impressions. Sections 4 and 5 established the empirical content of ($i$) and ($ii$): the record channel reads instrument-legible output and discounts substrate proxies (Benson, Li and Shue, 2019); the impression channel reads the confidence display, which tracks self-assessment error and not ability (Anderson et al., 2012).

### 7.2 Premises

**P1 (Instrument under-representation of the substrate).** Records legible to $R$ under-carry substrate calibration, and under-weight what they carry. The stock itself is tacit, accrues over $T_S$, and cannot enter per-role records resolving at $T_R$ (sections 2.1, 2.4); where observable proxies of substrate-relevant capability do exist, the promotion rule discounts them relative to instrument output, as the collaboration-experience result shows directly (Benson, Li and Shue, 2019). The premise required is directional, and deliberately weaker than absolute opacity: wherever the two calibrations compete for weight in an allocation, the record channel weights instrument calibration above substrate calibration.

**P2 (Cue inversion at the interpersonal layer).** Where records run out, evaluation runs on the confidence display; the display is produced by miscalibrated self-assessment; and the display is not reliably produced by ability (the null cue-display correlation of Anderson et al., 2012). Hence the interpersonal channel selects on a variable that diverges from ability exactly where miscalibration exists, and diverges in miscalibration's favour.

**P3 (Variance selection in rank order).** Rank-order promotion advances the right tail of observed outcomes; contestants' outcome variance is increasing in their risk posture; risk posture is increasing in overplacement and overprecision (sections 5.1, 5.2, with Goel and Thakor, 2008, supplying the equilibrium form). Hence iterated rank-order selection enriches the advanced population in miscalibration relative to the entering population.

**P4 (Ratification).** The composition produced by P2 and P3 is stable under governance oversight, because governance, optimising over $R$-legible objectives at $T_R$, prices moderate miscalibration as an asset (Gervais, Heaton and Odean, 2011; Otto, 2014; Campbell et al., 2011) and receives outcome feedback too decoupled from contribution to identify the selection within career horizons (Bertrand and Mullainathan, 2001; Daniel, Li and Naveen, 2020; Andreani et al., 2025).

### 7.3 The proposition and its corollaries

**Proposition (The selected regulator).** Under a measurement-driven promotion system satisfying P1 through P4, the operator stratum $O$ converges, across selection iterations, toward instrument calibration and away from substrate calibration: every retained operator tends toward a model of the regulator.

*Argument.* By P1, substrate calibration contributes systematically less to any candidate's standing through the record channel than instrument calibration contributes; a trait the allocation under-weights is under-selected, and drifts or declines in the advanced population relative to the traits the allocation weights. By P2, the impression channel contributes standing in proportion to a display that miscalibration produces and ability does not reliably produce; the channel therefore selects for self-assessment error. By P3, the record channel itself, beyond its blindness, actively enriches miscalibration through variance selection. Substrate calibration, moreover, correlates negatively with the selected traits along an identifiable pathway: the substrate-calibrated person's model contains the substrate's long cycles and accumulated fragilities, which surface in her evaluative behaviour as hedged estimates, resistance to metric-favourable actions that spend the substrate, and articulated concern in registers the credential gate discounts (section 5.4); each of these presents, inside the filter, as the negative of the confidence cue and the variance posture. The filter therefore does not treat substrate calibration as noise; it reads substrate calibration's behavioural expression as evidence against advancement. By P4, no oversight layer reverses the composition, because the oversight layer prices it. Across iterations, with homosocial reproduction compounding cohort upon cohort (Kanter, 1977), the stratum's composition converges toward the filter's fixed point: operators whose internal models are models of what the filter reads, which is the regulator. $\blacksquare$

The convergence claim admits an exact population-dynamic statement, which removes any residual dependence on absolute opacity and replaces it with a condition on gradients. For operator $i$, let $r_i$ denote calibration to the regulator, $s_i$ calibration to the substrate, $z_i = r_i - s_i$ the calibration displacement, and $w_i$ the advancement-and-retention weight the selection system assigns. The Price identity decomposes the change in mean displacement across one selection cycle as $\Delta \bar{z} = \operatorname{Cov}(w_i, z_i)/\bar{w} + \mathbb{E}(w_i \, \Delta z_i)/\bar{w}$ (Price, 1970; Frank, 2012). The first term carries selection; the second carries within-person change between selections, the socialisation channel through which survivors learn which registers advance and which registers cost. The Proposition then requires only the gradient condition $\operatorname{Cov}(w_i, r_i) > \operatorname{Cov}(w_i, s_i)$: advancement covaries more strongly with instrument calibration than with substrate calibration. P1 supplies the record channel's contribution to the gradient, P2 the impression channel's, P3 the tournament's; and where the behavioural expression of $s_i$ is read as evidence against advancement, as the argument above holds, the second covariance turns negative and the drift accelerates. Substrate calibration need not be invisible for the theorem; it need only be less selectable, and the composition drifts in every cycle by the identity's first term while the second term compounds it. The decomposition also names the three mechanisms of operator endogeneity in one line: selection (the covariance term, sections 4 to 6), socialisation (the expectation term, section 8.2 read from the transmitting side, the survivor learning silence), and attrition, the exit margin through which those whose $s_i$ persistently collides with the apparatus remove themselves, which the sorting results of section 8.1 supply (Van den Steen, 2005). With the selection criterion now specified, the triad recovers, and redirects, the attraction-selection-attrition cycle of the organisational-homogeneity literature; section 7.6 states the relation to that tradition exactly.

**Corollary 1 (The operator repair collapses as a class).** The repair of section 3 requires that $O$ contain, at decision-relevant altitudes, operators whose judgement supplies the substrate model $R$ lacks. By the Proposition, measurement-driven promotion attenuates exactly this population as a function of altitude. The repair therefore fails not contingently, for want of the right individuals, but structurally, because the channel through which individuals reach the deciding positions is governed by the apparatus the repair is meant to correct. The operator repair and the instrument repair fail for the same reason at one remove: the instrument repair proposes a better instrument, which remains an instrument (section 2); the operator repair proposes a supplementing judgement, which must arrive through a selection conducted by the instrument.

**Corollary 2 (The perverted Conant-Ashby).** Conant and Ashby (1970) require that every good regulator of a system be a model of that system. In the coupled arrangement instrument-plus-operator, the model requirement could in principle be met by the operator's judgement. Under the Proposition, the operator's internal model is selected toward isomorphism with the instrument, so the coupled arrangement contains, with increasing fidelity across iterations, two models of the regulator and none of the system. The composition can be written exactly. Let $M_R(S)$ denote the instrument's lossy model of the substrate. The repair of section 3 requires the operator to hold an independently formed $M_O(S)$, a second channel to the substrate that could confront the first; the selected operator instead holds, to the degree that her model has been trained, corrected, and validated on the instrument's outputs and on standing allocated by them, a model of the model, $M_O(M_R(S))$. The consequence is bounded by the data-processing inequality: no processing of a signal can recover information the signal does not carry (Cover and Thomas, 2006, ch. 2), so a model formed downstream of the instrument's compression cannot restore what the compression discarded, however able its holder, and the coupled arrangement's information about $S$ remains bounded by what $M_R$ carries. The bound is scoped, and the scope is the mechanism's own: it binds exactly insofar as the operator's non-instrument channel to the substrate has been thinned, and sections 8.1 and 8.2 document the thinning, sorting narrowing whom she hears, silence narrowing what they say. The human layer, added to supplement the compression, comes to duplicate it. The arrangement thereby satisfies the theorem's letter, there are models present, while inverting its point: the modelling capacity that regulation requires has been captured, through the staffing channel, by the wrong target. I call an operator stratum in this condition a selected regulator: a regulator whose human layer has been made isomorphic to its instrumental layer by the instrumental layer's own selection dynamics.

**Corollary 3 (Reflexive stability).** The configuration is self-maintaining. Operators calibrated to the instrument respond to evidence of the instrument's inadequacy in the instrument's own register, by refining metrics, adding indicators, intensifying measurement, since that register is what their models contain; the responses extend the apparatus's coverage of the promotion channel and sharpen the filter. The critique of measurement, when it penetrates such a stratum at all, is metabolised as a demand for better measurement. This corollary predicts the pattern the reform literature documents as the recurring conversion of measurement critique into measurement expansion (Power, 1997; Strathern, 2000; Muller, 2018), and explains it without imputing bad faith at any node.

### 7.4 What the formalisation does and does not claim

The proposition is a convergence claim about composition under iterated selection, not a universal claim about individuals. It is compatible with the existence, at any altitude, of substrate-calibrated operators: filters leak, tenure predates regime changes, founder-operators bypass the tournament, and some individuals hold both calibrations and survive on the instrument-facing one. The claim is that the filter's direction is fixed and its operation compounds, so that the density of the population the repair requires falls with altitude and with the regime's age, and falls fastest exactly where the repair needs it most, in the mature, intensively measured organisations where the instrument's blindness does the most damage. The proposition is likewise not a claim that firms err by their own lights in operating the filter; P4 states the opposite, and the depth of the trap lies there: the filter is locally optimal at every stage under the objectives the apparatus can express, and its aggregate product is a stratum constitutively unable to see what the apparatus cannot show it.

### 7.5 The temporal restatement: miscalibration as horizon calibration

One reformulation of the Proposition deserves separate statement, because it dissolves an apparent paradox in the selection literature and sharpens what the filter selects. Throughout sections 4 to 6, the selected trait has been named as the literature names it, overconfidence, a miscalibration. But miscalibration is a two-place relation: an internal model is miscalibrated with respect to a target. Specify the target as the substrate, its base rates, its long cycles, its tail risks, and the selected executives are miscalibrated, as the forecast-interval evidence shows against realised outcomes (Ben-David, Graham and Harvey, 2013). Specify the target as the regulator, the quarterly figures, the metric responses to action, the movements the dashboard will register within the evaluation cycle, and the same executives are frequently well calibrated: they model, with considerable accuracy, what the instrument will show, which is why they succeed within it repeatedly enough to ascend. What the literature registers as a bias is, under this redescription, an allocation of one modelling capacity to the target the selection environment pays for. The dissolved paradox is the persistence result: a decade of quarterly feedback does not correct the executives' intervals (section 4.1) because the feedback that reaches them, filtered through the reward system's decoupling (section 6.3) and the upward signal-stripping (section 8.2), is feedback about the regulator's states, on which their models are already performing. The calibration that never improves is calibration to a target the environment never scores. The Proposition can accordingly be restated without the vocabulary of bias at all: measurement-driven promotion allocates the operator stratum's modelling capacity to the regulator's horizon and away from the substrate's, and the trait the psychology literature measures as overconfidence is the substrate-facing shadow of an instrument-facing fit. The restatement matters for reform (section 12), because it implies that the selected are not defective reasoners awaiting debiasing, the framing the training industry sells, but accurate reasoners about the wrong object, and no debiasing intervention changes the object; only the selection principle does.

### 7.6 Neighbouring traditions and the exact difference

Three established traditions stand near the Proposition, and stating the distance to each locates the contribution. Merton (1940) described the bureaucratic personality: sustained discipline under rules transfers attachment from purposes to procedures, and the resulting trained incapacity turns the official's very reliability into blindness when circumstances outrun the rules. Merton's mechanism operates after appointment, through socialisation; the Proposition's operates before and during, through selection, and the two compound: the hierarchy is enriched in persons whose relation to the world already fits the blindness ascent requires, and socialisation then intensifies what selection began. What Merton called trained incapacity arrives here as selected incapacity, with training as its second term. Schneider (1987) established the attraction-selection-attrition cycle: organisations homogenise because they attract, select, and retain those who fit, and shed those who do not, so that the people make the place. The Proposition specifies the fit criterion that tradition leaves open, the measurement apparatus, and thereby reverses the slogan's direction of explanation: the place, through its instrument, selects the people who will remake the place in the instrument's image. Hambrick and Mason (1984) founded upper-echelons theory on the claim that the organisation reflects the cognitive bases of its top managers. The Proposition supplies the missing antecedent: the top managers reflect the apparatus that selected them, so that upper-echelons effects are, one step back, apparatus effects, and the loop closes, apparatus selecting operators, operators governing through the apparatus's categories, governance extending the apparatus, the extended apparatus selecting the next cohort.

## 8. Sorting and Sincerity: The Stratification of Recognition

The literature on entrenched metric regimes carries an epistemic model that the closure now forces open. The model is common knowledge of inadequacy: in the Soviet reporting system, factory managers knew the statistics were inflated, ministries knew the factories knew, planners knew the ministries knew (Berliner, 1957; Kornai, 1992; Nove, 1977); at NASA before Challenger, the O-ring concerns were documented, circulated, and known to be known (Vaughan, 1996); in the pre-2008 risk apparatus, traders, risk officers, and regulators alike knew Value-at-Risk did not carry tail risk, and the knowledge was itself public (MacKenzie and Spears, 2014; Taleb, 2007). The common-knowledge model explains persistence through the framework's monopoly on legitimate discourse: everyone knows, and the knowing changes nothing, because only framework-language has standing. The model carries a residual difficulty its users have registered without resolving: it requires, at the top of each hierarchy, either a stratum of cynics maintaining what they know to be false, for which the biographical evidence is thin, or a mechanism by which the top does not know what everyone below knows. Operator endogeneity supplies the mechanism, and in doing so replaces the uniform common-knowledge structure with a gradient.

### 8.1 Selection manufactures sincerity

Section 6.2 established the microfoundation: the overconfidence the filter advances is measured through personally costly exposure, executives holding undiversified positions in their own beliefs, and is therefore sincere by construction of the measure. Van den Steen (2005) supplies the sorting theory that generalises the point beyond the overconfidence construct. Where managers hold visions, strong beliefs about the right course, employees who share the belief sort into the firm and those who do not sort out, since shared belief raises the subjective value of working under the vision; the firm's belief distribution homogenises endogenously, without persuasion, through the joining and leaving decisions themselves, and Van den Steen (2010) extends the mechanism to the origin of shared beliefs in organisations generally. Rotemberg and Saloner (2000) show a complementary result: a visionary chief executive, meaning one with biased beliefs, can raise firm value by making credible the commitments a calibrated executive could not credibly sustain; Bolton, Brunnermeier and Veldkamp (2013) formalise the adjacent virtue of resoluteness, the leader's imperviousness to disconfirming signal, as a coordination asset. Across these models, the belief structure at altitude is not a residue that better information would wash out. It is functional, selected, sorted for, and sincerely held.

### 8.2 Ascent strips the disconfirming signal

Two communication mechanisms complete what selection and sorting begin, ensuring that whatever calibration survives the filter receives progressively less material to calibrate on. The MUM effect (Rosen and Tesser, 1970) names the robust reluctance of subordinates to transmit bad news upward; organisational silence (Morrison and Milliken, 2000) names the institutionalised form, in which whole categories of concern are withheld because employees infer, from prior processing of such concerns, that voice is futile or dangerous, and Detert and Edmondson (2011) document the implicit theories that sustain the withholding even absent any explicit sanction. The upward channel therefore delivers, at each altitude, a signal from which the substrate's distress has been progressively edited. Power's own perceptual consequences compound the editing: Galinsky, Magee, Inesi and Gruenfeld (2006) established experimentally that induced power reduces perspective-taking, the tendency to represent others' informational states, so that altitude diminishes the very capacity that would be required to reconstruct the edited signal from its residue.

### 8.3 The gradient stated

Composing selection (8.1), sorting (8.1), signal-stripping (8.2), and the perceptual effect of power (8.2), the epistemic structure of a mature metric regime stratifies as follows. At the working layer, where persons are coupled to the substrate daily, the framework's inadequacy is directly experienced, commonly known, and discussed in registers the framework does not carry; the common-knowledge model is accurate here, and the ethnographic record that generated it was largely collected here. Ascending, the density of substrate calibration falls by the Proposition; the belief distribution homogenises by sorting; the disconfirming signal thins by silence and the MUM effect; and the perspective-taking that might compensate declines with power. At the selected layer the framework is not a maintained fiction; it is the sincere content of the operators' models, funded with their own portfolios, reinforced by a filtered signal that largely confirms it. Recognition of the gap between framework and substrate, in short, is a decreasing function of height, and the regime's stability requires cynicism nowhere: below, those who know lack standing; above, those with standing do not know, and were selected for not knowing. The gradient explains the recurring historical observation that such regimes end through external shock rather than internal recognition (Berliner's Soviet case ended through glasnost; Vaughan's through disaster and subpoena; the risk-model case through the crisis itself): the altitude at which recognition would have authority is the altitude at which selection has made recognition scarcest.

## 9. The Alibi: Individualising Genres and the Deletion of the Room

If the argument to this point holds, a question of reception follows it. A selection structure of this consequence, operating in plain sight across the institutions of an entire economy, should be a standing object of public analysis, and it is not. What occupies its place in public discussion are two genres, apparently opposed, which this section argues are the same explanatory operation run in opposite directions, and whose joint effect is to keep the selection structure out of the explanation of its own products.

### 9.1 The Dunning-Kruger genre

Kruger and Dunning (1999) reported that participants scoring in the bottom quartile of tests of humour, grammar, and logic grossly overestimated their relative standing, and attributed the pattern to a metacognitive deficit: the skills required to perform well are the skills required to recognise poor performance, so the unskilled are doubly cursed, incompetent and unaware. The finding escaped the journals into general culture with a speed and totality few psychological results have matched, becoming the standing public explanation for confident incompetence: the confidently wrong colleague, the assertive novice, the manager certain of what he does not understand, each is processed, in the genre, as an instance of a cognitive defect residing in the individual skull.

The empirical career of the finding matters to the argument, and it has been unkind. Krueger and Mueller (2002) showed that the canonical quartile plot is generated by the combination of a better-than-average effect and regression to the mean: with any imperfect correlation between self-assessment and measured performance, the bottom quartile's self-estimates regress upward and the top quartile's downward, producing the celebrated crossing pattern from statistical structure alone. Nuhfer, Cogan, Fleisher, Gaze and Wirth (2016) and Nuhfer, Fleisher, Cogan, Wirth and Gaze (2017) demonstrated that random simulated data reproduce the canonical graphs, locating the effect substantially in a graphical convention and its numeracy assumptions. Gignac and Zajenkowski (2020), applying the Glejser test of heteroscedasticity and nonlinear regression to self-assessed and measured intelligence, found the pattern mostly a statistical artefact, with subsequent exchanges (Hiller, 2023; Gignac and Zajenkowski, 2023) leaving, at most, a weak residual effect (Dunkel, Nedelec and van der Linden, 2023). The present argument does not require the strong artefact conclusion, and takes no side beyond the published record; what it requires is already conceded on all sides of that exchange: the population-level phenomenon the genre invokes, a distinctive incapacity of the unskilled to know their standing, is far weaker than its cultural career, and much of what the genre explains through the individual skull is generated by the statistical structure of comparing any imperfect self-signal with any noisy measure.

### 9.2 The vanished answer key

Before the genre's statistical career is weighed, its migration must be described, because the migration contains a logical event the statistics do not touch. The original studies possess an answer key: participants estimate their performance on a bounded test, and the estimate is compared with a defined score. Whatever the comparison's artefacts, the two terms of the relation exist independently. When the genre travels into organisational and public life, the second term quietly changes hands. The colleague accused of overestimating his contribution, the employee told she overestimates her strategic understanding, the practitioner informed he overestimates his competence: overestimation relative to what independent score? In institutional use, the standard occupying the answer key's position is an evaluation the institution itself has produced, a manager's rating, a promotion decision, a compensation band, a title, a market outcome, which is to say, an output of exactly the selection apparatus sections 4 through 6 described. The inference then closes on itself: the system rates the person low; the person disputes the rating; the dispute is taken as confirmation of impaired self-assessment, since accurate self-assessors would accept the score. Disagreement with the instrument has become evidence of the disorder the instrument diagnoses. The circularity is independent of the artefact debate below: even granting the original effect in full, within bounded tasks against defined keys, its institutional form substitutes the apparatus's own verdict for the key, and thereby converts a psychological finding into a device by which any selection system can rule its own contested judgements self-certifying. The device has a formal shape. Performance observable to an institution is a joint product, $P = f(C, E)$, of formed capacity $C$ and the affording environment $E$, the room that solicits, receives, and renders the capacity actionable, which is the situated-cognition result assembled below. The institution observes $P$ and writes the observation back into $C$ alone, $P \Rightarrow C$, deleting $E$ from the inference; and since the institution itself supplies most of $E$, the deletion removes the institution's own contribution from every judgement it renders about persons.

### 9.3 What the genre deletes

Set the genre beside the selection evidence of sections 4 through 6 and its function becomes visible. The genre explains confident incompetence as a property of persons. The selection evidence explains the observed distribution of confident incompetence, its concentration in positions of influence, as a property of filters: environments that read confidence as competence, reward variance, price conviction, and advance accordingly. The genre, in other words, takes the filter's output and attributes it to the filtered, converting a selection effect into a dispositional one. The operation is the fundamental attribution error (Ross, 1977) institutionalised as an explanatory style: the situation's contribution, here the selecting environment's, is deleted, and the residue is booked to character. The deletion is the more consequential because the situated-cognition tradition has established, across independent literatures, that competence itself is not a skull-resident quantity of the kind the genre presumes. Rationality is bounded jointly by the mind and the structure of the environment, the two blades of Simon's (1990) scissors, and heuristics are rational relative to the environments they are adapted to (Gigerenzer and colleagues' ecological-rationality programme); cognition in real work settings is distributed across persons, artefacts, and representational media, so that performance belongs to the system rather than the individual (Hutchins, 1995); skill is constituted within communities of practice and does not travel as an individual possession (Lave and Wenger, 1991); and dispositions formed in one field misfire in another not through decay of the person but through hysteresis of the habitus, the lag of embodied history behind a changed environment (Bourdieu, 1990). Against this background, the same person is competent in one room and incompetent in another because competence names a coupling between a formed capacity and an affording environment; a judgement of incompetence rendered without specifying the room is not an incomplete judgement but a category mistake. The genre renders exactly this judgement as a matter of form, and its cultural success has made the category mistake the default public analysis of the phenomenon whose structural analysis sections 4 through 6 assembled.

### 9.4 The success narrative: the same deletion, opposite sign

The executive success narrative, the genre of the founder's vision, the leader's conviction, the decisive bet vindicated, appears to be the Dunning-Kruger genre's opposite: it explains achievement where the other explains error. Structurally it is the same operation. The narrative takes the tournament's right tail and attributes the position to the person: the conviction, the work, the refusal to listen to doubters. The selection evidence shows what the attribution deletes. A tournament that advances variance advances, in its right tail, those whose gambles resolved favourably; for each, the tournament simultaneously produced a left tail of statistically indistinguishable contestants whose identical posture resolved badly, and who appear in no narrative because narration samples on survival. The pay-for-luck results (section 6.3) add that the reward system itself cannot distinguish the vindicated bet from the lucky one, and pays both as contribution. The success narrative therefore performs the identical deletion of the selecting environment, with the sign reversed: below, the filter's discards are booked to personal defect; above, the filter's survivors are booked to personal virtue; and in both directions the filter vanishes from the account. The two genres jointly constitute what I call the alibi: the explanatory regime under which a selection structure's outputs are systematically re-described as facts about individuals, so that the structure never appears as an object of analysis in the culture it organises. The alibi requires no author, any more than the filter requires a conspirator; it is the fundamental attribution error operating at the scale of public discourse, energised below by the comfort of diagnosing others and above by the interest of the narrated in their own narration.

### 9.5 The fashion cycle and the capture of the correction

A further mechanism completes the alibi's machinery, and it deserves more weight than its position here suggests, because it operates on the correction itself: it governs what happens when substrate knowledge does, eventually, become sayable. The management-fashion literature (Abrahamson, 1991; 1996; Kieser, 1997; Benders and van Veen, 2001) documents that managerial discourse moves in short waves: techniques and vocabularies surge on rhetorics supplied by a fashion-setting market of consultancies, gurus, and business media, and recede on cycle times measured in quarters, with the swings bearing little relation to the techniques' demonstrated performance. The cycle interacts with the credential gate of section 5.4 to produce a phenomenon the diffusion literature has described in its own terms (Strang and Meyer, 1993; Zbaracki, 1998): when a slow-cycle body of knowledge crosses into fashion, its label is taken up by actors positioned at the fashion-setting nodes, while its long-standing carriers, whose warrant is experience rather than position, find the label they carried now spoken back to them by the newly credentialled, and their own priority reclassified as late arrival. Zbaracki's (1998) study of total quality management records the general form: the rhetoric detaches from the practice and circulates on its own, appropriated by those skilled in rhetoric's market rather than in the practice's substance. The mechanism completes the alibi temporally. The substrate register is discounted while unfashionable (the gate), and expropriated when fashionable (the cycle); at no phase does the apparatus credit the substrate's carriers, because at every phase the apparatus reads position and label, never the underlying stock, which is one more instance, at the discourse layer, of the single incapacity this paper has tracked through every layer: the apparatus cannot distinguish possession of the word from possession of the capacity. The mechanism warrants statement as a result in its own right, the semantic form of Corollary 3. Corollary 3 showed the apparatus metabolising criticism of measurement as demand for more measurement; the present mechanism shows it metabolising previously excluded knowledge as newly assignable vocabulary. When a substrate register becomes institutionally valuable, the apparatus need not admit the register's carriers; it can expropriate the register, granting its labels to operators whose standing the apparatus produced elsewhere and assigning them interpretation rights over the very knowledge it previously discounted. The apparatus discounts the knowledge while it is only knowledge, and rewards it once it has become vocabulary, and at both phases it reads position, never the stock. The consequence for the selected regulator is direct: fashion does not diversify the stratum's composition, however thoroughly it transforms the stratum's speech; the correction arrives, and is routed through the population that required correcting.

## 10. The Market Register: Variance Tournaments in Capital Allocation

The filter of sections 4 through 7 operates wherever allocation is conducted through instrument-legible records and rank-order selection. Nothing in its construction confines it to the interior of firms, and this section follows it to the allocation system that stands above firms: the market for capital, where selection is conducted by pricing rather than promotion, and where the filter's operation is, if anything, more visible, because the tournament's mathematics are published.

### 10.1 Excess entry and the power law

Camerer and Lovallo (1999) demonstrated experimentally that overplacement drives excess entry into competitive markets: when payoffs depended on relative skill, participants entered at rates producing systematic aggregate losses, each entrant confident of occupying the profitable ranks, and the effect strengthened when participants had self-selected into the experiment on the belief that they were skilled, a design the authors named reference-group neglect. The venture-capital market institutionalises the corresponding payoff structure. Returns to venture portfolios follow a power law: the bulk of aggregate returns derives from a small number of extreme outcomes, most investments return little or nothing, and the industry's economics consist in purchasing claims on the right tail (Kerr, Nanda and Rhodes-Kropf, 2014, who document the centrality of experimentation under extreme skew). A power-law payoff converts the allocation problem into a pure variance tournament: the rational strategy prices the tail, and the tail is supplied by founders whose beliefs about their own ventures sit far enough from the base rates to sustain the attempt. The market therefore selects, through the entry margin and the funding margin jointly, for exactly the belief structure the internal tournament selects for, and does so as a matter of portfolio arithmetic rather than evaluative error.

### 10.2 Priced conviction and its measurement

Gornall and Strebulaev (2020) established the pricing counterpart of the record channel's legibility bias. Reported unicorn valuations, computed by multiplying the price of the latest, most protected share class across all shares outstanding, overstate fair value by 48 per cent on average in their sample of 135 unicorns, with almost half losing unicorn status under contractually exact valuation; the headline number that circulates, and against which subsequent tournaments of founders, funds, and business coverage are run, is an instrument-legible figure systematically decoupled from the construct it is read as measuring. The parallel with section 5.3 holds term by term: as the sales figure stands to managerial capability, the post-money headline stands to enterprise value, a legible number selected on and celebrated while the illegible construct diverges beneath it.

### 10.3 A contemporary illustration

The mechanics admit a current illustration, offered here as illustration and nothing more, with no claim about the venture's eventual fortunes. Lovable, a Stockholm software company founded in November 2023, reached a reported 500 million dollars in annualised recurring revenue by June 2026 (TechCrunch, 2026); it raised 330 million dollars at a 6.6 billion dollar valuation in December 2025, and by July 2026 was reported in talks to raise a further 300 million dollars at 13.2 billion dollars, twice the December figure within seven months and roughly twenty-six times annualised revenue (Temkin, 2026). The point of the illustration is not the company, whose product may justify every dollar; the point is the tournament's visible arithmetic. A doubling of headline valuation in seven months prices a right tail; the price is set in a market whose payoff structure makes tail-pricing rational (10.1) and whose headline instrument overstates by construction (10.2); the round's announcement then enters the fashion cycle of section 9.5 as evidence, recruiting the next cohort of entrants under reference-group neglect (10.1), and the success narrative of section 9.4 stands ready to convert whichever ventures survive into stories of founder conviction, deleting the cohort that carried the same conviction into the left tail. Every component of the intra-firm filter reappears: variance selection, legible-number targeting, sincerity of the selected (founders are, of the entire chain, the participants most personally exposed to their beliefs), reward decoupled from contribution wherever monitoring thins, and the alibi waiting at the end. The filter is one structure at two scales; the scales differ in that the firm promotes its selected beliefs into authority over a substrate, while the market prices its selected beliefs into claims on one, and the downstream absorption, the employees, suppliers, and customers of the ventures the tail-price eventually reprices, is conducted by parties who never entered the tournament whose variance they absorb.

### 10.4 The nested loop

The relation between the two scales must be stated more strongly than analogy, because the levels are causally coupled, and the coupling changes what reform can mean. The capital market selects firms and founders; the selected firm's objective function is set against the valuation it must service; the objective selects executives and managers through the internal filter of sections 4 to 7; the selected operators determine what the internal regulator measures and demands; the regulator's demands determine which substrate expenditures occur and which outputs become legible as performance; and the legible outputs, revenue growth, user counts, margin movement, return to the capital market as the signals on which the next round of capital selection runs. Each level constitutes part of the fitness environment of the level below it, and the loop closes: the allocation layer selects the firms whose internal selection systems generate the signals the allocation layer selects on. Within the loop, a valuation is not a passive estimate awaiting confirmation; it functions as a selection signal that partly manufactures its own evidence, since the headline number confers the capital, recruitment power, and standing that enable the visible growth subsequently read as the number's vindication, the engine-not-camera dynamic of section 2.3 operating at the allocation layer. And the loop requires the selected operator as its transmission: someone must convert the wager into internal demands, translate the thirteen-billion figure into hiring targets, product commitments, acceleration, and intensified extraction from the substrate, and the operator selected for instrument calibration is the point at which the financial promise acquires operational hands. The coupling's consequence for reform is developed in section 12, and it is severe: an internal selection function that protected substrate registers would present to the allocation layer as slowness and illegibility, and the allocation layer would select against the firm that adopted it, so that the filter this paper describes is not one arrangement a firm might discard, but the locally stable configuration of a nested selection stack.

## 11. Objections

Five objections carry weight, and each receives its answer here at the strongest form I can give it.

### 11.1 Overconfidence creates value; the filter is a feature

*Objection.* The finance results the paper relies on cut against its conclusion. Goel and Thakor (2008) show moderate overconfidence raising firm value; Gervais, Heaton and Odean (2011) show it cheapening incentive provision; Hirshleifer, Low and Teoh (2012) show it purchasing innovation; Campbell et al. (2011) show boards optimally retaining it. If the selected trait pays, the filter is not a pathology; it is competent institutional design, and the paper has documented a virtue while narrating a vice.

*Answer.* The objection is correct about every result it cites and wrong about what the paper claims, and the distinction between the two is the paper's core. The cited results establish the per-trait, per-decision value of moderate miscalibration within the objective function the instrument can express, at the horizon the instrument can see. The paper's claim concerns the composition of the operator stratum with respect to a different variable: the capacity to model the substrate the instrument cannot express, at horizons the instrument cannot see. The two claims are compatible, and their compatibility is the trap's exact shape. A filter can be locally optimal in every ratifiable respect, this is premise P4, and still, by the same operation, evacuate from the deciding stratum the calibration on which the system's long-horizon viability depends, because that calibration is priced by no stage of the filter. The objection, pressed fully, amounts to the observation that the filter maximises what the instrument measures, which the paper affirms; the paper's addition is that what the instrument measures does not contain the substrate, by the measurement critique's own standing results, so that maximising it is compatible with, and under sustained pressure conduces to, the depletion the critique documents. Value on the dashboard and the argued conclusion do not compete; the first is the mechanism of the second.

### 11.2 Markets discipline the selection

*Objection.* Whatever internal filters select, product and capital markets eventually price outcomes. Firms run by miscalibrated strata underperform and are competed away, corrected by activist investors, or repriced; the selection the paper describes cannot persist against market selection at the level above it.

*Answer.* Three findings bound the discipline's reach. First, the horizon mismatch: the substrate depletion at issue matures over $T_S$, years to decades, while the outcome records on which market discipline operates clear at $T_R$; a cohort selected on variance can traverse full careers inside the gap, and section 6.3's pay-for-luck results show the reward system booking the gap's noise as contribution throughout. Second, the market's own selection runs the same filter: section 10 showed capital allocation pricing the identical belief structure through power-law arithmetic, so the level above does not correct the level below; it recapitulates it. Third, survivorship: the discipline that does operate removes its evidence, since the disciplined disappear from the panels on which persistence would be measured, and the surviving population, which is the population studied, observed, and narrated, is the doubly selected one. Market discipline does operate; it arrives late, and it arrives in the form the paper's section 8.3 predicts, as external shock, after the internal recognition that would have pre-empted it has been selected out.

### 11.3 Boards and professionalised governance already correct for this

*Objection.* Modern governance is aware of overconfidence; the behavioural-corporate-finance literature the paper cites is taught in the business schools whose graduates staff the boards. Selection-aware governance can discount the confidence cue, weight calibration, and repair the filter from the oversight layer.

*Answer.* The objection describes a possibility the evidence to date runs against. Campbell et al. (2011) show boards, behaving exactly as shareholder-value maximisation prescribes, retaining moderate optimism and dismissing calibration toward the tails, which is the filter operating through the oversight layer rather than despite it; Otto (2014) shows the belief structure priced into contracts, which is ratification, not correction. The deeper difficulty is Corollary 3: a board is itself an operator stratum, staffed through the same selection channels (director appointments run on records, reputations, and interpersonal impression, with homosocial reproduction documented at least since Kanter, 1977), and its response to filter-awareness is therefore predicted to take the instrument-calibrated form, an assessment framework, a competency matrix, a structured interview protocol, that is, more instrument, which extends the filter's coverage rather than escaping it. The objection would hold in a governance layer selected outside the measurement channel; identifying such a layer, and protecting its selection from reabsorption, is exactly the reform problem section 12 states, and its statement there is the objection's concession.

### 11.4 The constructs are too heterogeneous to compose

*Objection.* The paper composes overplacement in laboratory groups (Anderson et al., 2012), overprecision in CFO forecasts (Ben-David, Graham and Harvey, 2013), option-holding classifications (Malmendier and Tate, 2005), model-theoretic risk-underestimation (Goel and Thakor, 2008), and sales-record promotion (Benson, Li and Shue, 2019) into a single filter. These are different constructs, differently measured, in different populations; the composition may be an artefact of the survey.

*Answer.* The composition is legitimate because the argument does not require the constructs to be one trait; it requires them to share one selection-relevant property, and they do. Each construct names a divergence between an agent's internal model and a calibration target, expressed in behaviour the respective channel reads favourably: display (the impression channel), variance posture (the tournament), forecast conviction (the contracting layer), metric output (the record). The proposition of section 7 quantifies over the property, divergence-read-as-merit, not over any single trait, and its conclusion, the compositional drift of the stratum toward instrument calibration, follows from the channels' shared direction, which each literature establishes independently with its own construct and method. The independence of the literatures is, on this reading, the argument's strength: four traditions that do not cite one another, using non-overlapping measures on non-overlapping populations, each find their channel selecting against calibration, and the probability that four independent methodological artefacts share a direction is the probability the objection must defend.

### 11.5 The Dunning-Kruger critique undermines the paper's own materials

*Objection.* Section 9 leans on the statistical dismantling of the Dunning-Kruger effect, but the paper's sections 4 through 6 rely throughout on overconfidence research from the same broad literature. If self-assessment research is methodologically fragile, the fragility infects the paper's evidentiary base.

*Answer.* The dismantled object and the load-bearing objects are distinct, and the distinction is exactly the one Moore and Healy (2008) drew. The Dunning-Kruger claim concerns the relation between skill level and self-assessment error, that the unskilled err distinctively; the artefact critiques show this relational claim to be substantially statistical structure. The results the paper builds on concern the existence, selection, and consequences of self-assessment error wherever it occurs: that miscalibration exists (Ben-David, Graham and Harvey, 2013, measured against realised outcomes, immune to the artefact mechanics), that its display purchases status (Anderson et al., 2012, an effect of expressed confidence on evaluators, not a claim about who errs), that markets amplify its expression (Radzevick and Moore, 2011, a treatment effect), and that selection systems advance it (Goel and Thakor, 2008; Benson, Li and Shue, 2019, results about mechanisms, not about the skill-error relation). None of these inherits the artefact. The paper's use of the critique in section 9 is, moreover, load-bearing in the opposite direction: the argument there is that the culturally dominant individual-deficit account has lost its empirical warrant, which strengthens the case for the structural account exactly because the structural account never rested on the dismantled result.

## 12. Implications: What Cannot Be Repaired From Inside

The practical upshot of the closure is austere, and stating it plainly is preferable to decorating it. The measurement critique, completed by operator endogeneity, entails that two entire repair classes are unavailable. Instrument repairs, better metrics, longer horizons, richer dashboards, remain instruments and inherit the representational bounds of section 2; this was the critique's own conclusion. Operator repairs, better judgement, wiser leaders, walking the floor, more discerning boards, must arrive through a selection channel the instrument governs, and the channel's direction is fixed against them; this is the present paper's conclusion. What remains is the channel itself. The single lever the analysis leaves standing is the selection principle: not who is promoted, which the filter determines, but what promotion is a function of, which the filter presupposes. Any reform with purchase on the argued structure must alter the function, decoupling some fraction of ascent from instrument-legible records and confidence display, and coupling it instead to assessments rendered in the substrate's registers, by assessors positioned in the substrate's horizons, under protections that prevent the assessment's reabsorption into the metric apparatus, the reabsorption Corollary 3 predicts as the system's first response.

Section 10.4 enlarges the problem a level before any specification begins. The internal selection function is itself an object of selection: the allocation layer above the firm rewards the internal regimes that generate legible signals, and would read a firm that reformed its promotion function toward substrate registers as slower, less decisive, and less investable, selecting against the reform from above. The reform question is therefore misstated as the design of a better promotion function; stated fully, it asks how a counter-selection mechanism can survive inside an environment whose higher levels reward the selection regime it interrupts. The general requirement can be given exactly: a corrective institution must preserve a population whose standing does not depend on the signals produced by the regulator it is authorised to contest. This is more than better selection; it is institutionalised counter-selection, standing derived from a different coupling, direct substrate exposure, long temporal continuity, answerability for delayed consequences, protection from short-cycle retaliation. And even this cannot be a terminus, because every counter-selection body is a candidate selected regulator: grant one wise committee permanent authority and the filter resumes on the committee. The design aim that survives its own logic is heterogeneity of standing, several partially non-substitutable routes by which operators and correctors acquire authority, maintained so that no single selection logic can colonise the whole operator population; the constitutional analogue states it in one line, that no regulator should control every channel through which its own operators and correctors gain their standing.

The paper does not pretend this specification is easy to satisfy, and its difficulty is itself a finding. Every concrete proposal that comes to hand, peer assessment by long-tenured practitioners, promotion juries drawn from the working layer, mandatory tours of duty in the substrate before eligibility for altitude, apprenticeship models of succession, weighted terms for demonstrated calibration (Tetlock's forecasting-tournament instruments show calibration is measurable where outcomes are recorded and scored), faces the same pair of pressures: the fashion cycle stands ready to convert it into a labelled technique administered by the fashion-positioned (section 9.5), and the apparatus stands ready to operationalise it into metrics, at which point Goodhart's dynamic resumes (section 2.2). A reform that survives both pressures would have to be, in effect, institutionally boring and procedurally entrenched: unglamorous enough to escape the fashion market, and constitutionally protected enough, in the organisation's own governance, to escape metricisation. The analysis can specify these conditions; it cannot manufacture the will to meet them, and the stratification result of section 8 explains why the will is scarce where the authority resides.

The completed argument is falsifiable, and its predictions travel together in a way description after the fact does not: instrument calibration should rise with altitude under sustained measurement-driven selection; recognition of substrate loss should fall where decision authority rises; confidence display should purchase the most standing where direct performance information is weakest; substrate-calibrated warning should be recoded as behavioural or cultural deficit; newly fashionable knowledge should transfer authority to the already selected rather than to its long carriers; power-law payoff environments should intensify the selection of extreme belief; internal measurement reform should return, within a cycle or two, as measurement expansion; and substrate loss should become organisationally visible principally through external shock. Several of these already carry direct evidence in the sections above; the remainder are open, and the conjunction is the test, since any one alone admits other explanations.

A second implication is diagnostic and addresses the reader directly. The analysis implies a discipline for the interpretation of confident authority and diagnosed incompetence alike: both presentations are filter outputs before they are character facts, and the first analytical question in either case is not what the person lacks or possesses, but what the surrounding selection system reads, rewards, and cannot see. The discipline runs against the grain of the attribution machinery section 9 described, one's own included; the fundamental attribution error is not suspended for those who can name it. The paper's argument, consistently applied, therefore ends by qualifying its own author and readers: whatever standing this analysis attains will be mediated by the same credential gates, fashion cycles, and legibility filters it describes, and its reception will be, in miniature, one more datum on the structure it argues.

## 13. Conclusion

The measurement critique proved that the instrument cannot carry the organisation, and stopped at the operator's edge, exempting the human reader whose judgement might supplement what the apparatus cannot show. The exemption was honourable and locally true, and it left the critique incomplete in a way that mattered: as long as the operator population stood outside the analysis, the critique's own results licensed the hope that staffing could compensate for instrumentation. This paper has argued that the operator population stands inside the analysis. Ascent is an allocation conducted by the apparatus; the allocation reads records the substrate cannot enter and cues that ability does not reliably produce; its tournament form enriches miscalibration by variance; its governance layer ratifies and prices the result; and the composition it converges on is a stratum of sincere, personally invested, instrument-calibrated operators, a selected regulator, every retained member tending toward a model of the regulator rather than the system, in perverted satisfaction of the theorem that defines good regulation. Recognition of the apparatus's limits consequently thins with altitude; the genres of public explanation delete the filter in both directions, booking its discards to personal defect and its survivors to personal virtue; and the same tournament reappears at the scale of capital allocation, priced rather than promoted, with its variance absorbed downstream by parties who never entered it. The critique of measurement is hereby closed at its second flank: neither the instrument nor its operator can repair the blindness, because the operator is the instrument's product, and the only object the analysis leaves on the table is the selection principle itself, the one component of the apparatus that decides all the others and that no dashboard will ever propose to change.

---

## References

Abbott, A. (1988) *The System of Professions: An Essay on the Division of Expert Labor*. Chicago: University of Chicago Press.

Abrahamson, E. (1991) 'Managerial fads and fashions: the diffusion and rejection of innovations', *Academy of Management Review*, 16(3), pp. 586-612.

Abrahamson, E. (1996) 'Management fashion', *Academy of Management Review*, 21(1), pp. 254-285.

Anderson, C. and Kilduff, G. J. (2009) 'Why do dominant personalities attain influence in face-to-face groups? The competence-signaling effects of trait dominance', *Journal of Personality and Social Psychology*, 96(2), pp. 491-503.

Anderson, C., Brion, S., Moore, D. A. and Kennedy, J. A. (2012) 'A status-enhancement account of overconfidence', *Journal of Personality and Social Psychology*, 103(4), pp. 718-735.

Andreani, M., Dou, Y., Xu, G. and You, H. (2025) 'Are CEOs rewarded for luck? Evidence from corporate tax windfalls', *Journal of Finance*, 80(3).

Ashby, W. R. (1956) *An Introduction to Cybernetics*. London: Chapman and Hall.

Ben-David, I., Graham, J. R. and Harvey, C. R. (2013) 'Managerial miscalibration', *Quarterly Journal of Economics*, 128(4), pp. 1547-1584.

Benders, J. and van Veen, K. (2001) 'What's in a fashion? Interpretative viability and management fashions', *Organization*, 8(1), pp. 33-53.

Benson, A., Li, D. and Shue, K. (2019) 'Promotions and the Peter Principle', *Quarterly Journal of Economics*, 134(4), pp. 2085-2134.

Berliner, J. S. (1957) *Factory and Manager in the USSR*. Cambridge, MA: Harvard University Press.

Bertrand, M. and Mullainathan, S. (2001) 'Are CEOs rewarded for luck? The ones without principals are', *Quarterly Journal of Economics*, 116(3), pp. 901-932.

Bolton, P., Brunnermeier, M. K. and Veldkamp, L. (2013) 'Leadership, coordination, and corporate culture', *Review of Economic Studies*, 80(2), pp. 512-537.

Bourdieu, P. (1990) *The Logic of Practice*. Translated by R. Nice. Cambridge: Polity Press.

Brown, K. C., Harlow, W. V. and Starks, L. T. (1996) 'Of tournaments and temptations: an analysis of managerial incentives in the mutual fund industry', *Journal of Finance*, 51(1), pp. 85-110.

Camerer, C. and Lovallo, D. (1999) 'Overconfidence and excess entry: an experimental approach', *American Economic Review*, 89(1), pp. 306-318.

Campbell, D. T. (1976) 'Assessing the impact of planned social change', Occasional Paper Series, 8. Hanover, NH: Dartmouth College, Public Affairs Center.

Campbell, T. C., Gallmeyer, M., Johnson, S. A., Rutherford, J. and Stanley, B. W. (2011) 'CEO optimism and forced turnover', *Journal of Financial Economics*, 101(3), pp. 695-712.

Chevalier, J. and Ellison, G. (1997) 'Risk taking by mutual funds as a response to incentives', *Journal of Political Economy*, 105(6), pp. 1167-1200.

Choi, J., Hecht, G. W. and Tayler, W. B. (2012) 'Lost in translation: the effects of incentive compensation on strategy surrogation', *The Accounting Review*, 87(4), pp. 1135-1163.

Collins, R. (1979) *The Credential Society: An Historical Sociology of Education and Stratification*. New York: Academic Press.

Conant, R. C. and Ashby, W. R. (1970) 'Every good regulator of a system must be a model of that system', *International Journal of Systems Science*, 1(2), pp. 89-97.

Cover, T. M. and Thomas, J. A. (2006) *Elements of Information Theory*. 2nd edn. Hoboken, NJ: Wiley.

Daniel, N. D., Li, Y. and Naveen, L. (2020) 'Symmetry in pay for luck', *Review of Financial Studies*, 33(7), pp. 3174-3204.

Detert, J. R. and Edmondson, A. C. (2011) 'Implicit voice theories: taken-for-granted rules of self-censorship at work', *Academy of Management Journal*, 54(3), pp. 461-488.

Dunkel, C. S., Nedelec, J. and van der Linden, D. (2023) 'Reexamining the Dunning-Kruger effect', *Personality and Individual Differences*, 205, 112113.

Espeland, W. N. and Sauder, M. (2007) 'Rankings and reactivity: how public measures recreate social worlds', *American Journal of Sociology*, 113(1), pp. 1-40.

Galinsky, A. D., Magee, J. C., Inesi, M. E. and Gruenfeld, D. H. (2006) 'Power and perspectives not taken', *Psychological Science*, 17(12), pp. 1068-1074.

Frank, S. A. (2012) 'Natural selection. IV. The Price equation', *Journal of Evolutionary Biology*, 25(6), pp. 1002-1019.

Gervais, S., Heaton, J. B. and Odean, T. (2011) 'Overconfidence, compensation contracts, and capital budgeting', *Journal of Finance*, 66(5), pp. 1735-1777.

Gignac, G. E. and Zajenkowski, M. (2020) 'The Dunning-Kruger effect is (mostly) a statistical artefact: valid approaches to testing the hypothesis with individual differences data', *Intelligence*, 80, 101449.

Gignac, G. E. and Zajenkowski, M. (2023) 'Maybe the Dunning-Kruger effect is a statistical artefact after all: a reply to Hiller (2023)', *Intelligence*, 97, 101733.

Gigerenzer, G., Todd, P. M. and the ABC Research Group (1999) *Simple Heuristics That Make Us Smart*. New York: Oxford University Press.

Goel, A. M. and Thakor, A. V. (2008) 'Overconfidence, CEO selection, and corporate governance', *Journal of Finance*, 63(6), pp. 2737-2784.

Goodhart, C. A. E. (1975) 'Problems of monetary management: the UK experience', *Papers in Monetary Economics*, 1. Sydney: Reserve Bank of Australia.

Gornall, W. and Strebulaev, I. A. (2020) 'Squaring venture capital valuations with reality', *Journal of Financial Economics*, 135(1), pp. 120-143.

Hambrick, D. C. and Mason, P. A. (1984) 'Upper echelons: the organization as a reflection of its top managers', *Academy of Management Review*, 9(2), pp. 193-206.

Hiller, A. (2023) 'Comment on Gignac and Zajenkowski, "The Dunning-Kruger effect is (mostly) a statistical artefact"', *Intelligence*, 97, 101732.

Hirshleifer, D., Low, A. and Teoh, S. H. (2012) 'Are overconfident CEOs better innovators?', *Journal of Finance*, 67(4), pp. 1457-1498.

Hutchins, E. (1995) *Cognition in the Wild*. Cambridge, MA: MIT Press.

Kanter, R. M. (1977) *Men and Women of the Corporation*. New York: Basic Books.

Kaplan, R. S. and Norton, D. P. (1996) *The Balanced Scorecard: Translating Strategy into Action*. Boston: Harvard Business School Press.

Kennedy, J. A., Anderson, C. and Moore, D. A. (2013) 'When overconfidence is revealed to others: testing the status-enhancement theory of overconfidence', *Organizational Behavior and Human Decision Processes*, 122(2), pp. 266-279.

Kerr, W. R., Nanda, R. and Rhodes-Kropf, M. (2014) 'Entrepreneurship as experimentation', *Journal of Economic Perspectives*, 28(3), pp. 25-48.

Kieser, A. (1997) 'Rhetoric and myth in management fashion', *Organization*, 4(1), pp. 49-74.

Kornai, J. (1992) *The Socialist System: The Political Economy of Communism*. Princeton: Princeton University Press.

Krueger, J. and Mueller, R. A. (2002) 'Unskilled, unaware, or both? The better-than-average heuristic and statistical regression predict errors in estimates of own performance', *Journal of Personality and Social Psychology*, 82(2), pp. 180-188.

Kruger, J. and Dunning, D. (1999) 'Unskilled and unaware of it: how difficulties in recognizing one's own incompetence lead to inflated self-assessments', *Journal of Personality and Social Psychology*, 77(6), pp. 1121-1134.

Lave, J. and Wenger, E. (1991) *Situated Learning: Legitimate Peripheral Participation*. Cambridge: Cambridge University Press.

Lazear, E. P. (2004) 'The Peter Principle: a theory of decline', *Journal of Political Economy*, 112(S1), pp. S141-S163.

Lazear, E. P. and Rosen, S. (1981) 'Rank-order tournaments as optimum labor contracts', *Journal of Political Economy*, 89(5), pp. 841-864.

MacKenzie, D. (2006) *An Engine, Not a Camera: How Financial Models Shape Markets*. Cambridge, MA: MIT Press.

MacKenzie, D. and Spears, T. (2014) '"The formula that killed Wall Street": the Gaussian copula and modelling practices in investment banking', *Social Studies of Science*, 44(3), pp. 393-417.

Malmendier, U. and Tate, G. (2005) 'CEO overconfidence and corporate investment', *Journal of Finance*, 60(6), pp. 2661-2700.

Malmendier, U. and Tate, G. (2008) 'Who makes acquisitions? CEO overconfidence and the market's reaction', *Journal of Financial Economics*, 89(1), pp. 20-43.

Merton, R. K. (1940) 'Bureaucratic structure and personality', *Social Forces*, 18(4), pp. 560-568.

Moore, D. A. and Healy, P. J. (2008) 'The trouble with overconfidence', *Psychological Review*, 115(2), pp. 502-517.

Morrison, E. W. and Milliken, F. J. (2000) 'Organizational silence: a barrier to change and development in a pluralistic world', *Academy of Management Review*, 25(4), pp. 706-725.

Muller, J. Z. (2018) *The Tyranny of Metrics*. Princeton: Princeton University Press.

Nelson, R. R. and Winter, S. G. (1982) *An Evolutionary Theory of Economic Change*. Cambridge, MA: Belknap Press.

Nove, A. (1977) *The Soviet Economic System*. London: Allen and Unwin.

Nuhfer, E., Cogan, C., Fleisher, S., Gaze, E. and Wirth, K. (2016) 'Random number simulations reveal how random noise affects the measurements and graphical portrayals of self-assessed competency', *Numeracy*, 9(1), Article 4.

Nuhfer, E., Fleisher, S., Cogan, C., Wirth, K. and Gaze, E. (2017) 'How random noise and a graphical convention subverted behavioral scientists' explanations of self-assessment data: numeracy underlies better alternatives', *Numeracy*, 10(1), Article 4.

Otto, C. A. (2014) 'CEO optimism and incentive compensation', *Journal of Financial Economics*, 114(2), pp. 366-404.

Peters, T. J. and Waterman, R. H. (1982) *In Search of Excellence: Lessons from America's Best-Run Companies*. New York: Harper and Row.

Polanyi, M. (1958) *Personal Knowledge: Towards a Post-Critical Philosophy*. Chicago: University of Chicago Press.

Polanyi, M. (1966) *The Tacit Dimension*. Garden City, NY: Doubleday.

Power, M. (1997) *The Audit Society: Rituals of Verification*. Oxford: Oxford University Press.

Price, P. C. and Stone, E. R. (2004) 'Intuitive evaluation of likelihood judgment producers: evidence for a confidence heuristic', *Journal of Behavioral Decision Making*, 17(1), pp. 39-57.

Price, G. R. (1970) 'Selection and covariance', *Nature*, 227(5257), pp. 520-521.

Radzevick, J. R. and Moore, D. A. (2011) 'Competing to be certain (but wrong): market dynamics and excessive confidence in judgment', *Management Science*, 57(1), pp. 93-106.

Rosen, S. and Tesser, A. (1970) 'On reluctance to communicate undesirable information: the MUM effect', *Sociometry*, 33(3), pp. 253-263.

Ross, L. (1977) 'The intuitive psychologist and his shortcomings: distortions in the attribution process', in Berkowitz, L. (ed.) *Advances in Experimental Social Psychology*, 10. New York: Academic Press, pp. 173-220.

Rotemberg, J. J. and Saloner, G. (2000) 'Visionaries, managers, and strategic direction', *RAND Journal of Economics*, 31(4), pp. 693-716.

Sah, S., Moore, D. A. and MacCoun, R. (2013) 'Cheap talk and credibility: the consequences of confidence and accuracy on advisor credibility and persuasiveness', *Organizational Behavior and Human Decision Processes*, 121(2), pp. 246-255.

Schneider, B. (1987) 'The people make the place', *Personnel Psychology*, 40(3), pp. 437-453.

Shannon, C. E. (1959) 'Coding theorems for a discrete source with a fidelity criterion', *IRE National Convention Record*, 7(4), pp. 142-163.

Simon, H. A. (1990) 'Invariants of human behavior', *Annual Review of Psychology*, 41, pp. 1-19.

Strang, D. and Meyer, J. W. (1993) 'Institutional conditions for diffusion', *Theory and Society*, 22(4), pp. 487-511.

Strathern, M. (1997) '"Improving ratings": audit in the British university system', *European Review*, 5(3), pp. 305-321.

Strathern, M. (ed.) (2000) *Audit Cultures: Anthropological Studies in Accountability, Ethics and the Academy*. London: Routledge.

Taleb, N. N. (2007) *The Black Swan: The Impact of the Highly Improbable*. New York: Random House.

Temkin, M. (2026) 'Lovable reportedly in talks to double its valuation to USD 13.2B', *TechCrunch*, 8 July.

Tenney, E. R., Meikle, N. L., Hunsaker, D., Moore, D. A. and Anderson, C. (2019) 'Is overconfidence a social liability? The effect of verbal versus nonverbal expressions of confidence', *Journal of Personality and Social Psychology*, 116(3), pp. 396-415.

Tetlock, P. E. (2005) *Expert Political Judgment: How Good Is It? How Can We Know?* Princeton: Princeton University Press.

Van den Steen, E. (2005) 'Organizational beliefs and managerial vision', *Journal of Law, Economics, and Organization*, 21(1), pp. 256-283.

Van den Steen, E. (2010) 'On the origin of shared beliefs (and corporate culture)', *RAND Journal of Economics*, 41(4), pp. 617-648.

Vaughan, D. (1996) *The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA*. Chicago: University of Chicago Press.

Weidmann, B., Vecci, J., Said, F., Deming, D. J. and Bhalotra, S. R. (2024) 'How do you find a good manager?', NBER Working Paper 32699. Cambridge, MA: National Bureau of Economic Research.

Zbaracki, M. J. (1998) 'The rhetoric and reality of total quality management', *Administrative Science Quarterly*, 43(3), pp. 602-636.
