The Difference Between Being Tested and Being Watched

In the autumn of 1998, a hedge fund called Long-Term Capital Management lost roughly $4.6 billion in about four months and had to be unwound by a consortium of banks assembled at the Federal Reserve’s request, for fear its collapse would take the rest of Wall Street down with it. The fund’s principals included two Nobel laureates in economics and a team that had, by any ordinary measure, done the math correctly. Their models of bond spreads were not naive. They had been checked, refined, and checked again. And they failed anyway — not because the arithmetic was wrong, but because the thing being modeled had started paying attention to the model.

That sounds like a category error. Bond markets don’t pay attention to anything; they’re just prices. But by 1998 a great many trading desks were running strategies close enough to LTCM’s that when the fund needed to unwind its positions, it was selling into a market full of people who had learned, directly or by imitation, to think the way LTCM thought. The spreads LTCM’s models treated as independent, mean-reverting facts about the world were partly an artifact of how many people were betting on them reverting. The model had not simply been tested against the market. At some point, it had started to become part of what the market was.

The imitation channel is the one everyone tells, and it’s real: LTCM’s returns were public, its principals were famous, and prop desks across Wall Street built similar relative-value books, modeled on the same reasoning. But there was a second channel running underneath the first, and it worked by a different logic entirely — different enough that folding it into the first, the way the paragraph above almost does, is a mistake worth catching before it goes any further. LTCM’s positions were large, leveraged, and financed through a small set of the same prime brokers and counterparties financing everyone else’s similar positions. When the fund needed to unwind, it wasn’t just selling into a market of believers — it was selling into a market of balance sheets wired together by shared collateral and shared lenders, forcing correlated selling whether or not the sellers had ever heard of LTCM’s models, agreed with its reasoning, or knew the fund existed. A margin call doesn’t check whether the position it’s closing out was a good idea. It fires the same way whether the underlying thesis was right or wrong, and it would have fired just as readily if LTCM had never published a word of its own thinking. This deserves a name of its own, and gets one below.

This is worth taking seriously as a general problem, not a war story about one famous blowup, because it points to a distinction most of us use every day without noticing: the difference between a claim being tested and a claim being merely watched — and, it turns out, a third way a check can fail to be a real test that isn’t quite either.

What a real test requires

When an engineer runs a stress test on a beam, the beam does not know it is being tested. It does not adjust its behavior because it has heard about the test, has an opinion about the engineer, or stands to gain from a particular outcome. The test and the thing being tested are causally separate — the checking event doesn’t reach back and change what’s being checked. This is what makes the result trustworthy: if the beam holds, that fact was not manufactured by the act of testing it.

Historical sciences that can never run controlled experiments at all — geology, cosmology, evolutionary biology at the scale of continents and epochs — still meet this condition. You cannot rerun the Cretaceous, but the rock strata, isotope ratios, and fossil distributions do not rearrange themselves in response to being studied. The test there is indirect (a convergence of independent traces rather than a repeated experiment), but it still isn’t watching you back.

There are at least two different ways a check can fail to meet this condition, and it matters which one you’re dealing with, because they call for different fixes. The first is what people usually mean by “reflexive”: the target reads the claim, or is affected by people who have, and reacts in a way that’s directional and content-dependent — it moves toward whichever outcome the belief itself predicts, and moves harder the more people hold the belief. Financial markets, policy targets, and populations that have read the study about themselves work this way. Robert Lucas made a version of this argument about macroeconomic policy models in the 1970s: a model of consumer behavior estimated from historical data will fail once policymakers start acting on it, because the estimated relationships were never independent of the policy regime that produced them — change the regime, and you change the thing you thought you’d measured. Notably, this doesn’t require that anyone read Lucas, or the model itself; it requires only that people’s decisions depend on the policy regime, and that the regime is what changed. Goodhart’s Law needs a sharper kind of agent: once a measure becomes an explicit target, someone optimizing against it will bend their behavior to satisfy the letter of the measure rather than the thing it was meant to track. Ian Hacking’s looping needs sharper still — an interpreter, someone who reads a classification, understands what it says about them, and changes because of that understanding, which changes what the classification is now true of. Regime-sensitivity, optimization, interpretation: three different minimum requirements, too easily collapsed into one slogan about “people reacting to being modeled.”

The second way doesn’t need any of that — no belief, no optimizer, no interpreter. It needs scale. A margin call doesn’t check whether the position it’s closing out was a good idea; it fires the same way whether the underlying thesis was right or wrong, and it fires whether or not anyone involved has ever heard of the model behind the position. Call this congestion: not the target reacting to the claim, but the act of enacting the claim at scale consuming a resource — liquidity, in LTCM’s case — that the model’s own math had quietly assumed would stay available no matter how hard the model itself leaned on it. Popularity still matters here, but only upstream: more imitators running similar leveraged books through the same handful of counterparties is what puts the coupling in place to begin with. Once it’s in place, the mechanism that fires doesn’t check anyone’s beliefs. That’s the tell that separates it from the first category, and it’s why congestion earns a name of its own instead of folding quietly into “reflexivity” — which is the easy, and wrong, place to file it.

The objection worth taking seriously

An honest reader will point out that no measurement is ever perfectly independent of the thing measured. A thermometer left in a small volume of liquid changes the liquid’s temperature. Heisenberg’s uncertainty principle says something stronger still: certain pairs of properties cannot, even in principle, be measured without disturbing each other. If physics itself doesn’t get untouched access to its targets, what’s really being claimed about reflexive systems that physics doesn’t already concede?

The objection deserves a real answer. A thermometer doesn’t warm the liquid more because you guessed a higher number before dipping it in — the disturbance is indifferent to which hypothesis you’re rooting for. A crowded trade disturbs a market in the direction of the theory’s own prediction, and more forcefully the more people believed the theory in advance. That’s the cleaner way to state the difference, and it doesn’t require arguing that a large disturbance becomes different in kind from a small one — a move some readers will resist on principle regardless of how it’s argued. What differs is not magnitude but coupling to intention: physical disturbance is symmetric with respect to the hypothesis space, indifferent to content; reflexive disturbance is directional and hypothesis-contingent, pushing toward whichever outcome the belief itself predicts, and pushing harder the more the belief spreads. A test whose result improves in proportion to how much people trusted it beforehand is not a test in the sense engineering means the word.

This is also the precise reason congestion isn’t the same failure as reflexivity, even though both showed up in the same collapse. A margin call is symmetric in exactly this sense — it doesn’t fire harder because the underlying thesis was more widely believed, only because a price moved and a collateral threshold was crossed. It’s a thermometer wired to enough leverage to move markets when it registers a reading, but it’s still a thermometer: indifferent to content, not pushing toward confirmation of anyone’s theory. LTCM’s imitators were reflexive disturbance. LTCM’s counterparties were congestion. Telling that story as one thing, the way the opening of this essay almost did, erases a distinction worth keeping.

There’s a second objection worth crediting rather than dodging: game theory and mechanism design model strategic, self-aware agents explicitly, and they work — spectrum auctions designed on these principles have raised many billions of dollars for governments in outcomes close to their theoretical predictions. Doesn’t that show reflexive domains can be handled with engineering-grade rigor after all?

It does, and it’s worth being precise about what it represents. There are three ways to manage the channel between a claim and its target — this is one of them — and a fourth kind of fix, aimed at the claimant rather than the channel, comes later. Physics insulates: in the domains where the channel barely exists, there’s nothing to manage. Medicine blocks: it severs a channel that’s real but narrow, through blinding. Mechanism design does neither — it harnesses the channel, assuming from the outset that agents are watching and will game whatever rules they’re given, then writing the rules so that the gaming and the intended outcome coincide. That isn’t a counterexample to the argument here; it’s the argument’s positive form, applied to the one condition under which it’s tractable: mechanism design only works when someone actually gets to write the rules in advance. A spectrum auction has a designer with the authority to set the format before anyone bids. Culture, organic markets, and most of politics don’t offer that position to anyone — you’re playing inside a rule set you didn’t write and can’t rewrite from outside it.

It’s also worth being honest about what a successful auction actually proves, because it’s a more modest and more useful claim than it sounds. An auction landing near its predicted outcome doesn’t confirm the designer’s model of the bidders — it shows the design was robust even if that model was wrong. The best mechanisms are prized precisely because they don’t require correctly predicting what any bidder wants; they work whatever the bidders turn out to want. That’s a different achievement than confirmation, and it’s the honest one to reach for in exactly the domains where confirmation is unavailable: not “we predicted the market” but “we didn’t need to predict it correctly to get a good outcome.” So mechanism design is genuine proof that reflexivity can be engineered for. It is not evidence that most reflexive domains have access to an engineer, and even where it applies, what it delivers is robustness rather than the confirmation the rest of this essay has been chasing.

What the fix has always looked like, in the one place we already trust it

Medicine is the domain most people point to for the remedy to this problem, and it’s a good place to start — but it’s worth being honest about how much of the problem it actually solves. A patient’s cholesterol doesn’t scheme. But a patient’s behavior does respond to knowing they’re in a study, and a physician’s assessment of a patient’s improvement responds to knowing which arm of the trial that patient is in. This is not a hypothetical risk; it’s the reason the double-blind randomized trial exists as a specific piece of institutional machinery rather than a generic good idea. Blinding manufactures independence exactly where this one leak would otherwise open — it makes sure that neither the measurer nor the measured knows enough to let the assertion shape the confirmation.

But that’s all it does. Blinding says nothing about a trial that gets run, comes back null, and never gets published, while three positive trials of the same drug do. It says nothing about a company funding the studies of its own drug most likely to succeed and quietly declining to fund the ones less likely to. It says nothing about a trial’s outcome measures shifting after the data comes in, so that whatever moved becomes the headline result and whatever didn’t gets called secondary. Medicine isn’t the domain where this problem got solved; it’s the domain where one specific leak — assignment-knowledge reaching behavior and assessment — got solved cleanly enough that everyone points to it, while a slower replication crisis over everything blinding doesn’t touch has been running in parallel for a generation.

The medical practice that actually generalizes isn’t blinding. It’s preregistration: committing to a hypothesis and an analysis plan before the data exists to shape either. Where blinding works by severing a channel — cutting the wire between what someone knows and what they do — preregistration works by binding a commitment in time, so the confirmation, whatever it turns out to be, can’t be retrofitted afterward to look like it was expected all along. That makes preregistration kin to a different fix than blinding, one this essay is about to need.

Where blinding can’t reach

For the domains where blinding isn’t merely expensive but incoherent — you can’t blind a currency market to a central bank’s forward guidance, because the guidance isn’t a fact about the intervention that could be concealed while the intervention proceeds; the guidance is the intervention — there’s nothing left to sever, at any price. That’s a structural ceiling, not a budget line, so the fix has to look more like preregistration than like blinding: a commitment fixed before the outcome is known, rather than a channel that gets cut. The strongest version of this in open, reflexive systems is a genuine bet: tie the fate of the person making the claim to the same uncertain outcome the audience is being asked to act on, so the confirmation can’t be retrofitted after the fact any more than a preregistered hypothesis can. It doesn’t try to restore the causal independence blinding manufactures; it can’t, because the pathway from claim to target is exactly what makes the domain worth predicting in the first place. What it does instead is make sure the claimant loses if the claim turns out false, regardless of whether the claim was believed. This doesn’t stop the market from moving in response to the forecast. It stops the forecaster from profiting merely from having been believed, independent of having been right.

This fix has its own failure conditions, worth stating because they show exactly when it degrades back into the thing it was meant to fix. It fails if the claimant can exit before the reckoning arrives — sell the position, leave the advisory role, retire before the long-run data comes in. It fails if the population the claim is actually about isn’t the population bearing the cost of believing it — a forecast about people not yet born, or about a market the claimant has no stake in. Where exit is available or the timescales don’t line up, “skin in the game” collapses into the same unfalsifiable territory as everything else, dressed up as accountability.

The same resourcing problem shows up here too, worth naming because it cuts against the instinct to treat funding as protective. Rigorous preregistration and credible skin-in-the-game both require someone with the standing and capital to commit to them, which correlates with a claim’s popularity — but in reflexive systems, popularity is also the thing that causes the contamination these fixes are meant to catch. The same success that buys a claim the resources to be checked rigorously is often the success that has already made the check moot, because the crowding or compliance the popularity produced is now baked into the outcome. More money doesn’t uniformly protect against this. Sometimes it just buys a better-funded version of the same corrupted loop.

When the gap is opened on purpose

Everything so far has treated the failure to confront a claim honestly as an emergent hazard — a byproduct of success nobody necessarily intended, the way LTCM’s models were undone by their own credibility rather than by any dishonesty in building them. But the same gap that opens by accident can be opened on purpose, and the two need to be kept apart, because the remedies aren’t the same. Deliberately withholding or distorting information to gain advantage is an old and well-understood move; less obviously, so is a whole genre of tactics that don’t touch the information at all and instead attack the time available to check it. “Act now.” “This offer is available to the next thousand callers.” A countdown clock on a purchase page. These aren’t persuasion in the ordinary sense — they’re a deliberate compression of the window between hearing a claim and being able to confront it, engineered so that commitment is extracted before verification, deliberation, or a second opinion can occur. The claim doesn’t have to survive scrutiny if scrutiny is made to arrive too late to matter.

This is the intentional cousin of the contamination the rest of this essay describes, and the distinction matters. LTCM’s models weren’t dishonest; they were sincerely wrong in a way their own success helped bring about, and nobody was trying to stop anyone from checking anything. A false-urgency tactic has no such innocence — denying the audience a confrontation window isn’t a side effect of the claim’s popularity, it is the entire strategy. It’s also, not coincidentally, the exact inversion of the fix above: skin-in-the-game works by tying the claimant’s exposure to the same outcome the audience faces; manufactured urgency works by making sure the claimant collects before any exposure has the chance to bind at all. The question this essay has been building toward — could the test have embarrassed the person running it, without their having had any way to arrange in advance that it wouldn’t — has a sharper twin for this case: what would have had to go right, outside the claimant’s control, for this claim to fail before you acted on it? Where the honest answer is “nothing had the chance to,” you aren’t looking at a claim that was tested and passed. You’re looking at a claim that was deliberately kept from ever being tested at all.

None of this says reflexive knowledge is impossible, or that markets, policies, and populations are unknowable in the way flat-earthers say the moon landing is unknowable. It says something narrower and more useful: for any confident claim about a system that can read the claim, or a claim delivered under manufactured urgency, the question to ask is not “was this tested?” but whether anything outside the claimant’s control ever had the chance to prove it wrong before it was acted on. Most of what passes for confirmation in reflexive domains — cited studies, backtested strategies, celebrated forecasts, limited-time offers — never has to answer that question, because almost nobody asks it. LTCM’s models were tested constantly, right up until the day the market, which had been listening the whole time, tested them back. The harder cases are the ones built so that nothing was ever listening at all.

That question is itself a matter of degree, not a switch — there is almost always some chance, however slight, that a claim could have failed before anyone acted on it, so the honest version isn’t “could this have failed” but “was there a mechanism, independent of the claimant’s cooperation, that gave failure a real chance to show up.” It’s fair to ask that of this essay, too. Its central distinction faced exactly one real test while being written: a reader pointed out that the balance-sheet coupling in the LTCM story didn’t behave the way the essay’s own criterion said reflexive disturbance should — that it was symmetric, not directional, more thermometer than crowd. The essay was wrong about it, in a way outside the author’s ability to arrange otherwise, and the extra section above is what actually confronting that looks like: not a hedge wrapped around the original claim, but a category added because the original one didn’t hold. A distinction that never turns its own test on itself is asking for a trust it hasn’t extended to anything else.

Leave a comment