A paper published this month shows that language models can report on concepts held in their own internal states. It does not refute the framework. It forced me to replace a flat premise with a narrow one, and the narrow one is the premise the argument actually needed.
On 6 July, Anthropic’s interpretability team published Verbalizable Representations Form a Global Workspace in Language Models (Gurnee, Sofroniew, Lindsey et al., 2026). I read it the way you read anything that lands in the middle of your own subject: looking for the sentence that breaks the project. There is one that comes close. It is not the one I expected, and finding it was worth more than not finding it.
Start with what was at risk. Since the first draft, §3.3 has rested on a claim I stated without qualification: the producing system performs no metacognition over its own outputs. Everything downstream leans on it. Post-cognition is defined as the operation that supplies from outside what the system cannot supply internally. If the system can supply it internally, the definition loses its motive.
The paper’s core result is that production models maintain a privileged subset of internal representations, recoverable by a technique the authors call the Jacobian lens, which behaves the way psychologists say consciously accessible content behaves. It can be reported when the model is asked what it is thinking about. It responds to instructions to hold a concept in mind while doing something else. It carries the intermediate steps of inferences the model never verbalizes — the unnamed spider in “the number of legs on the animal that spins webs,” swapped for ant, and the answer moves from eight to six. And it supports a limited form of introspection: inject a concept into these representations on the user turn, ask the model whether it detects an implanted thought, and it names the thing you injected.
I had already cited the antecedent of that last experiment, and I had cited it badly. §3.3 of v5 mentions Lindsey (2025) — the concept-injection study whose protocol the 2026 paper adapts — grants that it identified something resembling self-monitoring, and then disposes of it in the same breath: significant, but it does not alter the framework’s architecture. That is a move I would flag in someone else’s text. A citation that concedes a phenomenon and then declines to let it cost anything is not a concession. It is a way of appearing to have read the objection. It was also pointing at the wrong paper: the entry in the bibliography is the Anthropic circuits work on model biology, not the concept-injection study the claim actually needed.
A model that reports a concept implanted in its own activations is doing something, and “no metacognition” is not an accurate name for it. So the flat version is gone. I am not interested in defending it.
What replaces it is narrower. Nothing in that work shows that a model represents the type of the claim it is about to make. It represents that the passage in front of it is in Spanish. It represents that a running character count stands at forty-six. It represents, unsettlingly, that it is being evaluated rather than deployed. None of these is the representation the framework needs, which would have to be something like: this assertion is causal-mechanistic rather than correlational; this one is framework-dependent rather than framework-independent; this one is definitional and its truth costs no empirical information at all. The reportable representations carry contents. They do not carry a typing of the assertion those contents are about to become.
So §3.3 now reads: language models lack epistemic self-classification. It is a smaller claim, and it is better positioned than the one it replaces. The old premise was a claim about a capability, and capability claims are hostage to the next experiment. The new one is a claim about a representational content, which is what the framework was always about. Content type is the framework’s primary dimension; saying that content type is absent from the model’s own representations is the same sentence, moved one level down.
Two things came back the other way, and I want to record both, including the one I am not entitled to press very hard.
The first concerns the width and the reliability of the introspective channel, and it is an argument for the framework rather than against it. The paper shows that the same information can drive a task without routing through the reportable representations at all. A model asked to continue a Spanish passage continues in fluent Spanish even after the reportable representation of Spanish has been overwritten with French; a model asked to name the language follows the overwritten value and says French. Alongside this, the reportable component of a concept’s representation turns out to carry a median of six to seven per cent of that representation’s variance, with the rest sitting outside it. What a model can say about its own processing is a narrow and selective channel onto that processing, not a transcript of it. Lindsey’s own numbers point the same way from a different angle: introspective detection succeeds on a minority of trials even in the strongest models he tested, and models asked to describe the mechanics behind their own calculations often describe them wrongly. Unreliable and narrow are distinct defects, and a validation protocol cannot be built on a channel with either. Any protocol that asks the model to certify the epistemic type of its own output is asking a channel to carry something it is not built to carry. Reconstruction from the output, under an explicit scheme, does not depend on that channel. That argument is now in §7.2.
The second concerns the ARlex–ARsem gap, and here I have to be careful with myself. The paper reports that abstract, persistent content lives in an intermediate band of layers, and that the final layers switch to a regime tracking the imminent output token rather than intermediate computation. Semantic content settled upstream; surface form selected late, and the late step is where the variance is. That is exactly the shape of a seventy-two-point distance between semantic and lexical agreement. It is a satisfying convergence and it is not a demonstration. I measured 180 outputs. They measured representations, in a different setting, for a different purpose. A convergence between two lines of evidence is not the same as a link between them, and writing it up as though it were would be the category collapse this project exists to catch. So §6.4 states it as a consistent reading, says plainly that no experimental link has been established, and §8.5 turns the link into a Phase 2 measurement: apply the representation-level methods to this corpus and see whether the verdict is in fact settled before the sentence is.
Now the sentence that comes closest to breaking the project.
The last section of the paper describes a technique the authors call counterfactual reflection training. The prediction it tests is strong: since a model’s internal reasoning routes through representations of things it might say later, you should be able to change what it thinks by changing what it is disposed to say if interrupted and asked to reflect. They train models to articulate ethical principles under that kind of interruption, and behavior improves in the uninterrupted contexts, where no direct training occurred. Ablate the implanted representations and the improvement largely goes away.
Read that with the framework in hand. If you can train a model to articulate a principle on demand and thereby put the principle into the representations that govern its silent reasoning, then in principle you can train a model to articulate the content type of a claim on demand — and post-cognition stops being an external protocol and becomes a training objective. The externality of the framework is not a necessary feature of it. It is a response to the current state of the systems it is applied to.
I have put this in the paper myself, at the close of §3.3 where the concept is defined, rather than leaving it for a referee to find. It took a false start to get it there: I first wrote it as a new subsection at the end of the Discussion, which is a section organized around the study rather than around the concept. Filing a claim about the concept under the heading for claims about the study is the category error this project exists to catch, and I made it while writing about having avoided it. Two things follow from it and they point in opposite directions. The first is that the most plausible route by which this protocol becomes obsolete now has a name and a demonstrated mechanism, which is a real result and belongs in Limitations. The second is that the taxonomy survives the transition. A training procedure of that kind requires a specification of what is to be articulated — which categories exist, where their boundaries run, what validation each admits — and that specification is the thing this paper contains. Internalization would relocate the protocol. It would not supply the content the protocol encodes. And it would open a question the Phase 2 instrument is already built to ask: does a model trained to classify its own claims classify them the way a trained human annotator does? That is an inter-annotator agreement question with a non-human annotator on the panel, and the answer is not obvious.
So, the state of the facts. One premise has been narrowed: from a claim about metacognition to a claim about epistemic self-classification, which is a component of it. The triangle — metacognition, epistemia, post-cognition — is unchanged; §3.3 now specifies which part of the first vertex the argument requires to be missing. Two arguments have been added in the framework’s favour, one of them stated at the strength the evidence supports rather than the strength I would prefer. One limitation has been added that identifies the conditions under which the framework’s operative form would no longer be needed. The taxonomy, the epistemic responsibility check, the D4.3 asymmetry and the framework-dependency results are untouched; the paper does not enter that territory and neither does the revision. The gap holds. The number does not move.
The operating lesson is one I keep learning in different costumes. Last time it was that a measurement instrument can carry a hidden ontological decision. This time: the premise you state most flatly is the one most exposed to somebody else’s experiment, and the flatness is usually doing rhetorical work rather than argumentative work. State the narrowest version that still carries the argument. It is less satisfying to write and much harder to knock over.
Documents: 11_2026_academic_v6.pdf, 12_2026_backmatter_v6.pdf, and the updated bibliography 08_epistemologia_ai_bibliografia_v6.bib — July 2026. The source paper is Gurnee, Sofroniew, Lindsey et al., Verbalizable Representations Form a Global Workspace in Language Models, Transformer Circuits Thread, 6 July 2026.