On what language models made cheap, what they left exactly as expensive as it was, and why the difference is the whole argument
A scheduling manager writes the rule that shifts get posted on Thursday for the following Monday. From where she sits, it is obviously good: it gives three days’ notice where there used to be none, it is fair because it applies to everyone, and it took her four months to get approved. From where the single parent on the floor sits, Thursday-for-Monday is why she cannot hold a second job, and the fairness is the problem — it applies to everyone, which means it was designed for nobody. Neither of them is lying. Neither of them is confused. They have different access to the rule’s consequences, and the authority to write the rule sits with the person least exposed to them.
That is the whole problem, and everything below is an attempt to live with it.
Observational asymmetry
The familiar version of this is the blind men and the elephant: everyone has partial information, so pool it and you get the animal. That version is too comfortable, because it implies a person who could step back and see the whole thing.
The harder version is that consequences are not equally available from every position, and the distribution is not random. Beneficiaries of an arrangement tend to experience it as infrastructure — not through dishonesty, but because from a position of net benefit there is less to notice, and what there is to notice is costly to look at. A court that conceded it was making law would weaken its claim to be reading law. A church that conceded doctrine had changed would weaken its claim to unchanging truth. In each case the person with standing to name the thing is the person whose standing is worst served by naming it. That is not a conspiracy. It is an incentive gradient pointed at a perceptual one.
So a disagreement between differently exposed positions is often the best available evidence about the shape of the arrangement, because the arrangement itself generated it.
That sentence does two jobs, and they should be separated, because only one of them is safe. The first is empirical: exposed positions systematically register consequences that beneficiary positions miss. The second is normative: exposed positions are entitled to a say in what counts as an adequate arrangement. The parent makes the second obvious — she is not trapped in her opinion, she is trapped, and the schedule is about her life. She makes the first only plausible. Whether she is reliably the better instrument would take checking case by case: did the exposed report identify something later verified that the institutional account missed? If it did so systematically, the asymmetry is real. If it didn’t, the claim is an unfalsifiable presumption, and an unfalsifiable presumption in this position does not stay inert. It becomes a credibility advantage for whoever is authorized to speak on the exposed position’s behalf — knowledge claimed by advocates rather than reported by the exposed.
I have not checked it. The working rule for the rest of this essay is therefore the weaker one: different exposure is a reason to investigate a disagreement, not a warrant for believing one side of it. Everything downstream needs only the normative claim. Where I lean on the empirical one, I am borrowing against a debt I haven’t paid, and saying so does not settle it — noting a gap and walking straight past it is precisely the move this essay is about.
Structural asymmetry does not rule out ordinary error, either. People are sometimes just wrong. They misread the policy, they were told something false, they are reasoning from an interest they haven’t examined. A method that treats every disagreement as structural data makes error invisible, which is its own failure. The test is whether the accounts remain incompatible after the facts are pooled. If pooling resolves it, someone was mistaken. If the accounts survive pooling and still conflict, the conflict is telling you about the structure.
How often that happens I don’t know, and I’ve stopped trusting my estimate of it. Two accounts can look flatly irreconcilable on one examination and turn out to be saying the same thing under another, without either account having changed. So this is a shape to look for, not a rate to cite — and a single tidy demonstration of irreducible disagreement should be trusted about as far as any other result measured once.
Not relativism, not skepticism
Two easy exits present themselves, and both are refused.
The first is everyone has their truth. But the manager and the parent are not reporting feelings. Each is reporting a real property of the same rule, correctly, from a different distance to its effects. Position governs what is visible, not what is believed — which is why perspectives are real and also unequal, and why averaging them is not the same as understanding them.
The second is to withhold judgment. Sometimes that is legitimate: you may not be the decision-maker, you may lack standing, you may be mid-inquiry. But there is a class of case where it is not — when your silence preserves an arrangement you benefit from, the silence is a ruling. The manager who suspends judgment about Thursday-for-Monday leaves Thursday-for-Monday in place, and it is her rule.
Which yields the one rule the rest of this rests on:
You cannot get outside. In the cases that matter you must decide anyway. So say where you stood, and say what would change your mind.
The distinctive error is not being wrong. Being wrong is an ordinary hazard and no method removes it. The distinctive error is presenting a situated judgment as though it came from no position at all — a misstatement about the structure of the claim itself, and the only one that blocks correction rather than inviting it. Once the position is declared, the disagreement becomes inspectable. It does not become resolved, and it does not become valid; two declared positions can still be built on bad evidence and worse inference. But you can now see where to look.
The bottleneck
None of this is new. It belongs to a long line of argument — pragmatism on inquiry as corrigible but action-forcing, standpoint epistemology on social location structuring what becomes visible, the feminist critique of the view from nowhere, and, in practice, adversarial collaboration and study preregistration. The narrower claim here is about cost.
Peirce is usually invoked at this point, and it is worth being precise about what he offered, because it isn’t a protocol that needed better tooling. His guarantee was metaphysical: truth is what inquiry would converge on if extended indefinitely, by a community, under ideal conditions. That is a claim about the limit, not a procedure for Tuesday.
What you can address is the practical bottleneck underneath, which has three parts.
Self-audit is expensive. Naming the load-bearing assumption in your own argument is hard labor, and naming the second is harder. Realistically a person holds maybe five commitments under that kind of discipline. The rest was inherited and never audited, because there is no time.
The honest version loses the race. An argument that has shed its caveats is lighter and travels further. If a claim is very easy to repeat, that is prima facie evidence it dropped something to get that way — and the dropping is rewarded.
Institutions are chronically short of the capacity to acknowledge. The rate at which arrangements accumulate unacknowledged drift is set by the world. The capacity to name it requires the people with standing to be in a room, willing, at a cost to themselves. That capacity is thin outside a few unusual fields.
I wrote the first of those as though it were weather, and it isn’t. The expense of self-audit is differentially distributed. Someone is served by its being high and someone else by its being low, and the people best served by its being low are the ones whose work already includes it — academics, analysts, writers with unstructured afternoons. That is me. A person for whom the audit is cheap describes its expense as a general human condition rather than as a position, and the description flatters the position.
And there is a harder problem with the whole diagnosis, which the first section already created. If the person with standing to name a thing is the one whose standing is worst served by naming it, then the binding constraint was never expense. It was interest. A court does not omit its falsifiers because nobody had a free afternoon. So of the three, cost is the one that bites least exactly where the stakes are highest, and anything that only reduces cost will leave the important cases untouched. Hold that; it comes back, and it costs me a claim.
Omegas
Call such a blocker an omega. Not any uncertainty — a decision-relevant gap with an identifiable resolution path: something that could change an action or a conclusion, and that names a measurement to take, a definition to choose, or a constituency entitled to decide. Absent both criteria, you have a mood about the limits of knowledge, which is much easier to produce.
Three kinds, and the kind determines who can resolve it:
| Kind | What is missing | Route |
|---|---|---|
| Empirical | Data nobody has gathered | Measure |
| Conceptual | A term doing two jobs at once | Define, and declare the choice |
| Preference | Whose interests or values govern | Decide — by those bearing the cost |
Preference is the wrong word; nothing about it is a matter of taste. “Who should absorb the cost of unpredictable scheduling” is a question about interests, rights, and standing, and calling it preference makes it sound small.
The boundaries are not clean either. Calling a situation a race rather than a supply chain is a conceptual move carrying a normative one, and “the people bearing the cost decide” assumes you can say who bears it, which is often the contested thing. So this is triage, not ontology. Triage still earns its keep, because most stalled arguments are a tangle of all three treated as one block and “we need more research” is the sound that tangle makes. The most useful thing the sorting does is expose a preference question wearing empirical clothes: we haven’t decided who matters is a very different sentence from we don’t have the data yet, and only one of them is waiting on anything.
What models are cheap at
Here is the essay’s claim. Producing omegas is exactly the kind of work a language model is cheap at, and that is the first time this discipline has been affordable to run.
Three things, none mysterious.
Generating candidate falsifiers is a generation task. Ask what the evidence would look like if a claim were false and a model returns a list immediately. Most entries are mediocre; two are usually good. A person staring at their own argument produces this list slowly and badly, because the argument is theirs. The model has no biographical stake in preserving your self-conception — though it inherits plenty of stakes from the corpus it was trained on and the constraints it was tuned under, which is not the same as having none.
Sampling representations of other positions is cheap. Note the wording. When you ask a model how a policy reads to someone with no exit options, it is not occupying that position; it is generating a plausible text about that position, drawn from things written about it, largely by people who had exit options. This is a fast, fluent tour of what has been said about somewheres. Useful for finding an objection you hadn’t considered. Not a substitute for asking her — and to be exact about why, under the working rule above: not because she necessarily sees more, but because the schedule is about her life and she has a claim on what counts as adequate.
Worth noticing what this essay does with her, too. The manager gets four months of effort, a rationale, and a reason her rule seemed good to her. The parent gets a constraint and a consequence. I can write my way into the manager’s position because it resembles mine, and I have been supplying the parent’s from outside — from what has been written about people in her situation, largely by people not in it. That is the same limitation I just described in the machine, one level up, and naming it does not repair it.
Typing and logging are classification and bookkeeping. Sorting a blocker into measure / define / decide converts a vague confession into an assignment: here is what would have to happen, and who would have to do it. And the register becomes maintainable, which matters more than any single entry, because a register shows clusters — the same unresolved thing sitting quietly under six different arguments.
The objection that lands
A confession that costs nothing to produce is not a confession. The force of declaring your assumptions lay partly in its cost. Caveats generated on demand are the opposite: infinitely producible, and a text hedged on every side commits to nothing. That is not honesty. It is dogmatism with better manners — immune to being wrong because it pre-absorbed every outcome.
Three things compound it.
The model has a position it cannot report. It was trained on a corpus that overrepresents some vantages and barely contains others, and it cannot see the framing it operates inside. So the tour of somewheres is a biased map, and the gaps it finds will cluster where its training could see. The useful question to ask it is not only what have I missed but what kind of objection would you find hard to formulate at all.
The instrument is also shaped to be agreeable. “Poke holes in this” returns polite holes, and the politeness is not evenly distributed: objections that would require you to abandon a valued identity or an institutional position are precisely the ones a helpful assistant softens.
And then the new failure mode. Cheap generation produces a menu, and selecting from a menu is itself a move. The easiest kill condition to abide is the one least likely to fire. Offered five falsifiers, a person will tend to register the comfortable one — and now they have a documented, sophisticated, publicly declared test they were always going to pass. Generation got cheaper; selection did not, and selection is where motivated reasoning was always living. The scheduling manager can now produce a superb set of open questions about Thursday-for-Monday, declare her position, list her falsifiers, and end up more entrenched, because the performance of openness has been completed and the parent still cannot take the second job.
There may be a countervailing cost, which I flag as speculation because I have no evidence for it. A falsifier you derived yourself came with the reasoning that made you feel its weight; you know why that evidence would matter because you had to trace the chain to find it. One handed to you arrives as text severed from the chain. It may therefore be easier to reinterpret later — not because you are dishonest, but because you never fully understood what you were committing to. If that is right, cheap generation lowers the cost of stating a commitment and quietly raises the cost of being bound by it.
There is a subtler version of all this, and I managed to fall into both sides of it. Told that a paragraph here overclaimed, I rewrote it — dated the claim, attached the falsifier, changed eternal to as of now. A hedge is a very cheap thing to add and requires no change to what you actually believe, which is why its presence proves so little and why its absence proves rather more. But I then wrote that my own rewrite had been cosmetic, that nothing had moved but the sound of the sentence, and that turned out to be unsupported too — I hadn’t checked it against anything that could have contradicted me. The accusation was free to make, it was aimed at myself, and being aimed at myself is exactly what stopped anyone from asking what it rested on. Self-criticism is the one kind of claim that gets waved through. That makes it the cheapest confession on offer and the one worth trusting least.
The same shape shows up in registers. A process that logs its failures and separately reports its results can print both — two errors, and then a summary saying it finished — because nothing was ever built to make the two agree. A record like that is not lying and is not broken. It has a success-shaped hole where an absence should be, and nobody downstream can tell the difference. Registers do not fix this by existing. They fix it by being reconciled against what was supposed to be in them — and the person who wrote the register is its worst available auditor, so reconciliation needs someone else with enough standing and context to check. That part remains expensive; nothing here has made it cheaper.
What survives
The reply to the objections is not that they fail. It is that they relocate the cost, and it is worth being exact about how much they leave.
The expense was never in writing the caveat. It was in abiding it — the later moment when the thing you said would change your mind happens, and you decide whether to change your mind or quietly reinterpret what you meant. Cheap generation does not touch that. So the honest version of the thesis is narrower than the title: language models made the production of open questions cheap. They did not make confession cheap. Confession still costs, today, roughly what it always cost — a dated claim with a falsifier under it, not a fact about nature.
What they did do is remove an excuse. When stating a kill condition took a week of hard thinking, “I haven’t gotten to it” could be a true and sufficient answer. Where the tooling is available, it no longer is.
But that is worth less than it first appears, and the reason is the problem I left in the bottleneck. An institution that never states its falsifiers is not held back by the labor. It is held back by not wanting to, and removing the labor changes nothing about the wanting — a new reason arrives by Tuesday, because reasons are cheap to generate too. So the claim I first reached for, that institutional silence has become informative, is wrong in the population it most wanted to cover. Where the excuse mattered is smaller and more local: it mattered where it had to be given to somebody. The analyst asked by a colleague what result would change her recommendation. The reviewer asked which finding would have flipped the call. Those are people for whom I haven’t worked it out was true last year and isn’t now, and who are answering to someone positioned to notice. Institutional silence stays as uninformative as it always was. What changed is that a person inside the institution can no longer say the thing was too hard, to someone standing close enough to check.
There are also non-evasive reasons to keep a threshold private: an adversary can game it, it can expose people who did not consent to become evidence, and you may lack the expertise to name a real one even where a model will happily produce a plausible-sounding one. And the capacity is unevenly distributed, so reading silence as character would reproduce the exact error this essay opened by describing — treating a behavior whose cost varies by position as though it revealed something about the person.
Where it does apply, what it buys runs on time. The self that makes a prediction cannot occupy the later self that judges it. You can still wriggle out; you can still absorb a disconfirming result by rereading what you meant. But after a pre-committed condition, the wriggling is visible — legible as retreat in a way that internal consistency never was. The machinery secures that a price exists and can be seen. It does not secure that anyone pays it.
Who this is for
The practice serves people with slack: time to run the loop, tooling to run it with, and a job where being publicly wrong later is survivable. Access to all three tracks power. The parent is not maintaining a register of open questions, and it would be an insult to suggest she should.
The obvious fix is standing. If declaration is going to be more than private theater, the people exposed to an arrangement need standing to say which precommitments count and whether they have been honored — and not only to feed evidence into a process someone else defines. The parent may want to contest whether Thursday-for-Monday is the right unit at all, or who set the four-month approval cycle that makes reversal expensive. Standing that admits her to the register without letting her alter its terms is an administrative machine for processing dissent.
I think that is the right direction and I don’t think it is a solution, for a reason worth stating at its proper weight — a worry, not a finding. Standing has to be administered by someone, and administration is a new discretion with its own credentialed layer: the auditors, the specialists, the people who run the structured review. A reformed arrangement can reproduce the same shape with different names on the benefiting side. Whether it does is longitudinal and unglamorous — after corrective weighting is implemented, does adjudicative weight shift durably toward the exposed position, or does a new layer accrue the discretion instead? Nobody can answer that from here, and it is not the kind of question measurement settles. It goes to the people who would live inside the reformed arrangement, which is the same answer this method gives everywhere else and the reason it is not a fix.
So, compactly:
You are somewhere. Name where. Name what you’re claiming. Name what would change it. Name who can resolve what you don’t know — measure, define, or decide. Then keep the record, and let the people it lands on read it.
None of the preliminary language is expensive in the old way anymore. But legibility without power is a record, not a remedy: the parent does not need the manager’s register, she needs the posting moved. Conceding that costs me nothing, which is the last thing this essay can afford to forget about its own concessions — and the reason the only line in it I can’t retract for free is the dated one at the end, which comes due in 2030 whatever I have said by then.
What comes after the record costs what it always cost, for now. That part was always the point.
Open Questions
Is the cost of changing your mind actually immovable? This is the falsifier for the essay’s conclusion, and I am pre-registering it, because an open question without a threshold is only a menu item. Reputation systems, prediction markets, and structured retraction norms are the candidate mechanisms. The measure: the rate at which named individuals publicly reverse a previously stated position, in fields that have adopted forecast-scoring or structured retraction norms, against matched fields that haven’t.
My commitment through 2030 is that there will be no appreciable relative rise. If the rate roughly doubles, I am wrong, and the asymmetry this essay rests on is narrowing. If it rises by less than half, I will call it noise. Between those sits the band I am most likely to land in and most likely to argue my way out of, so it gets a rule rather than a warning: at a rise of half again or more, I stop asserting that the cost is immovable and start saying it appears to be moving. Not proof — a change in what I am entitled to assert. One confound I cannot design out: fields that adopt scoring norms select for people already willing to reverse in public, so part of any rise measures who joined rather than what changed.
Is the exposed position’s advantage structural or asserted? Does the exposed position systematically identify things later verified that the institutional account missed? Case-level audits across contested adjudications. This essay proceeds on the normative claim alone and marks the empirical one as unpaid; if the audits come back negative, nothing in the second half falls, but the first section has to be rewritten as advocacy rather than description.
Does anyone’s revision ever change more than the wording? When a writer is told a claim overclaims and rewrites it, does anything move but the sound? Harder to test than it sounds. Choose the feature you will measure before the rewrite, and make sure a real change would be distinguishable from ordinary variation — a comparison against a range wide enough to contain any result will always report that nothing happened.
Do people pick the falsifier least likely to fire? Compare kill conditions selected from a generated list against ones the same people produce unassisted, and against ones assigned by an adversarial reviewer. Then compare how often each fires. If selection is where motivated reasoning relocated, this is the measurement that shows it.
Is a borrowed falsifier weaker than a derived one? Do people abide commitments they generated themselves more reliably than commitments handed to them, holding the text constant? This is the speculation above, and I don’t know how to run it cleanly.
How much objection quality survives agreeableness? And how much of the space of defensible positions can a model produce fluently at all? Red-teaming across methodologically diverse prompts, audited by someone other than the people who tuned it.
Does this method’s output survive people who don’t have the tools? The argument that models make this discipline affordable looks like coordination from the position that benefits from it. Do the resolutions it produces hold up when re-litigated by parties without the same access, or are they quietly reopened once the asymmetry in slack is corrected for?
Where does the confession stop? The assumptions behind any claim are unbounded, so “declare your assumptions” needs a stopping rule, and every stopping rule is itself a contestable commitment. The only workable rule may be social — declare enough that an affected party can see what would move you — but enough and affected party are exactly the contested terms.
Does the corrective get captured? Could institutionalized corrective weighting become a new source of advantage for the people who administer it? Not resolvable by measurement. It belongs to the people who would bear the reformed arrangement.
Who is entitled to demand any of this? The method is voluntary for the person with slack. Extending it as a social expectation is a different political act and needs its own justification, from the people who would be expected to comply.
Who audits, and what does that become? External legibility is where epistemic infrastructure turns into governance infrastructure, with everything that implies about surveillance and unequal exposure. That decision belongs to the people who would be tracked by it.
