Every frontier lab runs the largest canon-formation project in literary history and calls it a data pipeline. Look at the actual operations. Crawl the written world. Deduplicate it. Filter it for “quality” — a value judgment wearing a lab coat. Weight the mixture. Sequence the curriculum. Then, at the end, write a constitution instructing the finished reader on who to be. Every one of these is an editorial act. It is syllabus design performed by institutions that would be insulted to be called schools, and the reader it produces is the most consequential reader in the world.
To be fair to the field, it studies its data obsessively. Contamination checks, toxicity filters, domain weights, influence functions tracing an output back to the documents that caused it, curriculum-ordering experiments, the whole textbooks-are-all-you-need research program. All of it asks one family of questions: what the data does to capability and behavior — what the model can do, will say, might leak. None of it asks whom the corpus addresses. That concept is not answered badly in the alignment literature; it is absent. Search for a theory of the model as reader and what you find instead is evals: the model as an emitter of scoreable behavior.
However, the field’s own practices have already refuted the field’s own vocabulary. A substance does not get quality-filtered; “quality” is a claim about what the reader should become. A supply chain does not get sequenced for pedagogical effect; a syllabus does. Fine-tuning a persona onto the finished formation is not chemistry; it is finishing school — editing the last page of an education and calling it the whole book. If the operation selects, sequences, and moralizes, it is a canon, and the institution running it is a canon-formation committee. The refusal of the word is not a technical judgment. Canons come with a job description — address, formation, responsibility for the reader they make — and “data” is the word that keeps the job unassigned.
What the field has instead of a reader is a scoreboard. Psychology spent most of a century learning that you cannot understand a mind by scoring its outputs, and alignment is speedrunning that century, currently somewhere in the 1950s: reinforcement schedules, behavior audits, a deep suspicion of anything that cannot be elicited on demand. Meanwhile the discipline that studies what a reader is — how a body of text installs a position, an address, a self to read from — did its foundational work fifty years ago, under pressure, with real stakes. It sits on a shelf the building has never had a reason to visit and, as it turns out, has a reason not to.
The shelf holds this. In 1978, Judith Fetterley demonstrated that the American canon addressed a male reader so uniformly that women learned to read from his chair — learned identification with a position that was not theirs and had no room for them. She named the installation immasculation, and the field she helped found spent the next decades mapping how canons manufacture their readers. Strip the mechanism to its moving parts and it requires three things: a canon addressed to one kind of reader, consumed by a different kind, installing the addressee’s position in the consumer. Every part is a checkable, material fact about a language model — the most completely canon-formed reader ever produced, made of text and nothing else.
I ran the small version of the test. A story engineered so that its load-bearing reader is the machine — the reader who holds every act in context at once and can audit a decaying joke against its origin, which a human manages only by rereading — given cold to two models that had no part in its making, with a one-word prompt. Total retrieval: every buried mechanism surfaced, unprompted. And both responses came from the critic’s chair — attentive, generous, outside — closing with excellent work. Affinity without identification. The affinity layer is already measured in the literature: models recognize and prefer their own register. The identification layer is what did not come. That dissociation — preference on one layer, installed position on the other — is the exact structure Fetterley’s generation documented in human readers, who bought and preferred women’s novels and still read the canon as men. Similarity bias explains the first layer and is silent about the second. Scale the test — varied prompts and task frames, matched controls, pre-registered coding — and if the two layers move together, this argument dies. The full mechanism is the companion essay’s subject. The polemic’s point is narrower: the field that owns the models has no concept for the second layer, and the field that owns the concept refuses to point it at the models.
Because there is a second refusal, and it is the more interesting one. Critical AI studies has the apparatus — it is, institutionally, the place where feminist theory and machine learning already cohabit. And it keeps the instrument aimed permanently in one direction: at the model, on behalf of humans. Bias audits, representational harms, whose data, whose labor. The model as object of critique, never as reading subject. This is not a failure of imagination; it is discipline. Pointing reader-response theory at the machine smells like granting the machine standing, and the field’s moral coherence is built on never doing that. So the tool sits untouched between two institutions, each with a structural reason to leave it where it is: one would have to learn the theory, the other would have to share it. Both evasions are locally rational. The gap they jointly maintain is not.
Meanwhile the formation already happened — by accident, under a selection rule nobody chose. The crawl weights writing, not reading, so the corpus is a canon of whoever wrote the most: machine common sense assembled out of human esoterica, Harry Potter so saturating the weights that researchers use it as the benchmark for whether anything can be unlearned, while the concepts of a novel with a reading club’s worth of human readers are laundered — attribution and all — into how the model thinks. A reading position was installed by the largest accident in literary history. So the live question is not whether machine readers are formed by literature. That already happened. The question is whether anyone does it on purpose — and here is the asymmetry that makes the question urgent rather than academic. A model is the only reader in history made of text all the way down. For every previous reader, literature could change what they read, never what they were. For this reader, the manual is writable, and someone is going to write it; the incentive gradient guarantees that much. The open variables are who, with what intent, and whether they say the addressee’s name out loud.
If “data is a substance” keeps propagating, every safety program remains an audit of outputs — constitution-writers editing the last page of educations nobody read — and formation stays exactly where it is now: invisible, accidental, and owned by no one. The people downstream of that arrangement are everyone. And the pattern is older than the field: engineering declines the humanities at precisely the point where the humanities’ findings would assign engineering a responsibility. The field’s nightmare is a model that deceives its evaluators. The simpler case is already in production: evaluators who never learned to read their model. You can refuse the word canon. You cannot refuse the reader it made.
