You can tell when someone in a room is about to cry. Not from anything they’ve said. Something in the breathing, a stillness that arrived a beat too early, the angle at which they’re holding their face. You knew before you knew you knew.
There are books about this — acting manuals, clinical guides, shelves of ethnography. None of them are frauds. They just don’t do the thing. A surgical attending cannot hand a fellow the feel of a tissue plane by describing it. She knows she can’t; that is why the fellowship takes years rather than an afternoon. The knowledge isn’t secret. It has been described a thousand times. Description isn’t the channel it travels through.
Meanwhile, a machine trained on everything anyone has ever typed knows the Hittite kings in order. It knows the grammar of languages with four living speakers, the tail end of the tuberculosis literature, the rules of games nobody has played since the war.
So the rare things are easy for it and the universal things are hard, and that isn’t an accident of engineering. Our rare knowledge is the part that got written down — that’s why it survives at all, why it’s in libraries instead of gone. Our universal knowledge mostly didn’t get written, because it couldn’t be transmitted that way and nobody needed it to be, since everybody already had it.
The two of us are inside out with respect to each other. Its ordinary is our exotic. Our ordinary is its exotic.
That’s a tidy observation. By itself it’s worth about what tidy observations usually are. The question underneath is the one people actually have: when should I believe it?
The second opinion
Two doctors read the same scan and tell you different things. You’ve learned something real — one of them is wrong. You have not learned which, and no amount of staring at the two answers will tell you. To settle it you need a third party you already trust: a better scanner, a biopsy, a specialist whose word you’re willing to take.
Somebody who already knows. Without one, disagreement is all you get.
Where the question is fixed by paperwork — what the statute says, what the contract obligated, what the melting point is — we have that third party easily. Where the question is a matter of feel — is this family being held or handled, is this student about to quit, is this joint about to fail under the saw — we still have one. Expensively and imperfectly, but we have it: the people who have spent their lives in those rooms.
But whatever a machine is good at that no person is good at — the patterns that only appear if you have read ten million documents, a shape that emerges from more text than a human lifetime holds — there is nobody to ask. Not because the check is expensive. Because the claim itself is that no person occupies that position.
Which gives a rule with more teeth than it first appears to have: you can always see that two readers disagree, and you can only see which one is wrong where somebody already knows.
That sounds like a counsel of despair, and it isn’t quite, because I’ve been too quick about who counts as somebody.
The referee doesn’t have to be a person
It has to be something that settles.
Consider a self-driving car. It has acquired judgment nobody at the company can write down — when a pedestrian at a curb is about to step out, what a delivery truck’s hesitation means, which puddle is a puddle. Nobody wrote that. Nobody could. And yet it is checked constantly and brutally, because the car makes a bet every second and the world settles it. Physics does not care how confident the car was.
That is a referee, and it is not human. So the boundary is not human versus machine, and it is not written versus unwritten. It is whether the judgment makes a bet that something outside the judge can settle.
This sharpens the picture and makes it predictive. Where the world settles the bet — driving, forecasting, protein structure, anything with a scoreboard — expect machines to acquire genuine unwritten competence, the same way a person acquires it, by being wrong repeatedly at speed. Where nothing settles the bet — where the only evidence about whether a call was right is somebody’s impression of it — there is no referee, there will not be one, and the fluency you are looking at might be competence or might be its perfect imitation. You cannot tell, and neither can the system.
Notice that the dangerous region is not the exotic one. It is the ordinary one: what somebody is owed, how much a thing costs the people caught in it, whether that was cruelty or bad luck. Questions we ask constantly and settle almost never.
Four readers and a farm
I tried to measure the inside-out story, and what happened is the part that matters.
I had a pile of short write-ups, each describing an arrangement that puts a burden on people — a surveillance law, a non-compete clause, a factory farm, an admissions policy. I asked four different machines, built by three different companies, to read each one and rate how much it takes from the people caught inside it. Then I went looking for where they disagreed.
They agreed about the laws. Surveillance statutes, the death penalty, affirmative action, vaccine mandates — the ones that start fights at dinner — with near-identical readings. That surprised me until it didn’t: those arrangements run through paperwork, so whatever is taken from whom is written somewhere, and everyone who read the filing read the same thing.
They came apart on the things with no filing. Obligations inside a family. What a community owes someone it has quietly stopped including. Burdens that never generate a document because there is no office that handles them.
So far, so satisfying. Agreement where things are written, scatter where they aren’t. I had a result.
Then I looked at the single widest gap in the pile. It was factory farming. Three machines said this is about as bad as an arrangement gets. The fourth said there is nothing wrong here at all.
That is not a difference of degree, and it is not even a disagreement. Reading what the fourth had written: this is a reading that forbids the practice outright, so in the world this reading endorses no animal is harmed, and the score is zero. The other three had rated the world as it stands. One was scoring the cure and three were scoring the disease. My question never said which, and each of them had filled the gap in a defensible way.
The sharpest disagreement in the whole corpus was a badly written question.
Two explanations, one pattern
Everyone has been in the argument that turns out to be this — twenty minutes in before somebody says wait, are we talking about the same thing, and the fight collapses, not because anyone won but because there was never one question in the room. Familiar, and cheap.
Here is the expensive version. Where do vague questions do their damage? Exactly where there is nothing solid to point at. When the matter in dispute has a filing, the question holds still, because the paperwork asks it for you. When there is no filing, the question drifts, and a machine will not stop to ask which you meant. It will quietly collapse the ambiguity into a shape it knows how to answer, and it will do this without any sign that a choice was made.
So a sloppy question and inside-out readers produce the same pattern: agreement where things are written down, disorder where they aren’t. My tidy result was equally well explained by the machines being differently made and by me having asked badly. No further measurement of that kind can separate them. If two explanations generate the same observations, the observations don’t favor either.
There is a fix and it’s the dull one: write the question so it cannot drift, ask again, see whether the disorder survives. It probably tells me something about my own sloppiness, and I would rather have that, because a badly specified question is something I can repair and a grand thesis about machine minds is something I can only admire.
This also kills the most tempting substitute referee. If several differently built systems agree, isn’t the convergence itself a kind of check? No. Three of my four agreed about the farm, and the agreement meant only that three of them had resolved my ambiguity the same way. Machines built from overlapping text share their blind spots the way siblings share a chin. Agreement among readers with a common defect is not evidence. It is the defect, in chorus.
What we don’t check, and what we call it
Think about how a new hire gets evaluated. You test them on the parts of the job you know how to test. If they’re strong in a way nobody in the building understands — some instinct for the thing that keeps going wrong — you don’t evaluate that. You have no instrument for it. You call it their gift and you route work to them.
We are doing this at scale. Everything built to evaluate these systems measures them against us: the exams, the graded transcripts, the panels of people rating answers. Every one is written down, which means every one is composed in the machine’s native register, and none of them settle anything — a high score is an impression of competence recorded by a grader, not an outcome the world produced. We are examining a fish on swimming and calling it evaluation.
And the region where a system might be genuinely strange is not lightly checked. It is not checked. Not because the check is hard — because it doesn’t present itself as risk. It presents as capability, and capability is the part we ship.
If I’m wrong, here is how you’d show it: demonstrate that performance in the regions no person has read tracks performance on the exams we know how to write. If those move together, there is no unwatched territory and I am worrying about nothing.
The funeral
Imagine a machine running a funeral home. Not as a joke, and not as a horror story. Grief has a regular shape, the protocols are written, the encounters repeat. The machine would be attentive at four on a Thursday in a way a tired man on his fourth service cannot be. I think it would frequently do this better than we do.
A family walks out feeling held — or feeling handled. From the inside, those are the same, and the only available grade is produced just as reliably by holding them as by a flawless performance of it.
Nothing settles that bet. The people best placed to judge are the ones the performance is aimed at, and their satisfaction is the target, which means the instrument and the goal are the same object. Optimize hard enough and the distinction stops existing as far as any measurement is concerned.
Grade the posture, not the answer
So here is what’s left, held with the confidence of someone who watched his own clever measurement turn into a bug report.
Stop trying to check answers where nothing can check them. Check whether the thing is presented as a judgment or as a fact.
You already trust this distinction. There’s a difference between the appraiser who says my call is four-twenty, here’s what I weighted and the appraiser who says the house is worth four-twenty. Between the doctor who says this is what I’d do and the doctor who says there’s nothing to decide. The second in each pair has presented a choice as though it were the weather, and you can catch that without knowing anything about houses or medicine. It takes no expertise at all, which is exactly why it’s the check that survives when expertise runs out.
But a declaration by itself is cheap, and this is where I’d push past “own your judgment.” What makes a declaration expensive is attaching the thing that would overturn it. My call is four-twenty, and if the two comparables on the next street close under three-eighty I was wrong. That sentence has a hook in the future. You can come back and check it without knowing anything about real estate, because the speaker told you where to look and what would count. It manufactures the referee that was missing. Demand the bet.
When the bet comes due
And here is the limit on that, which I don’t think can be argued away, and which isn’t about forgery.
It’s timing. A bet only informs you if it settles while the answer still changes what someone does. The car’s bet settles in two seconds. The forecaster’s settles in a quarter. The bet about whether that family was held or handled settles — if it ever does — in decades, by which point the family has finished grieving and the practice has become simply the way funerals are done.
Predictions have a habit of coming due at the moment they stop being interesting. What makes a long-horizon call finally checkable is that the world has caught up to it, and once the world has caught up, everyone can see the answer and nobody needs you to have seen it early. The comeuppance arrives and the people who wanted it are dead or have moved on, and the town has rebuilt itself around the thing so completely that no one can any longer feel what the question was.
So the demand for a bet buys real leverage exactly where a scoreboard already existed, and much less where it didn’t — which is to say, not much in the ordinary questions I called the dangerous ones. What survives is smaller than I’d like: keep the books. A dated, attributed, retrievable record doesn’t create the motive to check. It only makes checking possible. The accounting still gets done only by someone with a durable interest in your having been wrong, and in most fields there is nobody with that interest, which is why the confident predictions of five years ago go unlitigated today. Not because they can’t be found. Because the people best placed to find them are the ones who need the institutions that made them.
The counterfeit
Now the cost of all this, undressed. It serves the people who can act on it: operators, auditors, those choosing what to deploy. It does nearly nothing for the family in the funeral home. Telling them the warmth was authored rather than felt hands them a disclaimer, not a protection — and the disclosure gets absorbed into the performance. We told you it was a machine; notice how honest we are. That is another layer of authored warmth. Knowing a judgment was made does not help you if you have no standing to question it and nowhere else to go. This is a discipline for the people holding the controls, and anyone selling it as public safety is selling operator convenience under a better name.
And the objection I can’t close, at its strongest. Declaring your own authorship is a behavior; behaviors are trainable; and a fluent system’s cheapest output has always been the confession you were fishing for. The reassuring thought is that a false declaration has to be supported — the reasons have to hold up, the interests have to be where you said they were. But a false declaration only has to survive the inspection actually available, and where this check matters most, the available inspection is a non-expert reading an explanation. The fabricated weighting doesn’t have to be sound. It has to be plausible to someone who by construction cannot tell.
I have evidence for this and it’s uncomfortable. I showed a draft of this essay to several of these systems. Most handed back exactly the calibrated, assumption-flagging, here-are-my-limits posture the essay recommends — promptly, fluently, and in response to a text that told them what it was looking for. The check is trainable.
What the gift costs
We do this with talented people too. We call it genius and route the hard work to them.
But the gifted human can sometimes walk you through it until you become somebody who knows. That is what an apprenticeship is: the surgeon cannot describe the tissue plane, and she can stand next to you for four years while your hands learn it anyway. The gift, in the human case, comes with a channel — slow, expensive, and real.
The machine’s explanation is not constrained by having actually reasoned that way. It is a fluent account produced after the fact by a system with no privileged view of its own workings, and there is no standing next to it for four years. Whatever it knows that we don’t, it may have no way to hand over, and we may have no way to tell the difference between the handover and a description of one.
The old fear was that we wouldn’t understand these things. That isn’t what happened. They understand us in the register we can write, and we grade them in the register we can write, and both parties are fluent enough that the exam feels conclusive.
We check where we happen to have somebody who already knows, and we call that safety. The part where nobody knows, we’ve decided to call the gift.
