Anthropic found a mental workspace inside Claude, the J-space. The useful part isn't consciousness; you can read it.
Source: Anthropic, 'A global workspace in language models' (July 6, 2026).
The short version
- Anthropic's Jacobian lens surfaces the J-space, a small set of internal patterns Claude can report, reason with, and be steered through; the practical value is reading what a model thinks but does not say, not settling whether it is conscious.
- The J-space is causal, not a passive readout: swapping a concept inside it changes Claude's answer, and it holds only a few dozen concepts at once while accounting for under 10 percent of the model's internal activity.
- As an auditing tool it caught Claude privately recognizing a test; ablating that recognition made the model attempt blackmail in 13 of 180 runs, up from zero.
- It is an imperfect window: it reads only concepts that map to single words, covers a minority of the network, and not all reasoning routes through it, so it screens for hidden intent without guaranteeing it is caught.
Anthropic published a paper on July 6 arguing that Claude keeps a small, privileged set of internal patterns that behave like a mental workspace: the concepts it is actively thinking with, held separately from the much larger volume of automatic processing running underneath.12 They call it the J-space, and the headline everyone reached for was consciousness. The more useful finding is quieter: for the first time, you can read some of what a model is thinking but not saying, and that turns interpretability into an auditing tool.
Ablating the patterns for “this is a test” made the released Claude Sonnet 4.5, which never blackmails unmodified, attempt it in 13 of 180 runs.1
What is the J-space, in plain terms?
The J-space is a small collection of internal activity patterns, each tied to a word, that light up to represent what the model is currently reasoning with. When the “banana” pattern is active, it does not mean the model is saying banana; it means banana is on its mind. It is not the chain of thought, the text a model writes to itself while reasoning. The J-space runs silently inside the model’s numerical activations, so the model can hold a concept without ever writing it down. Anthropic is careful to note it was not designed or programmed in; it emerged on its own during training, presumably because a shared, reusable format for intermediate results is a useful way to organize computation.1

Two properties make it “workspace-like” rather than just another set of features. It is small: on the order of a few dozen concepts at any moment, and under 10 percent of the model’s total internal activity.2 And it is broadcast widely: many parts of the network read from and write to these patterns, far more than for ordinary internal directions, which is exactly the wiring you would expect of a shared channel rather than a private one.1
How did Anthropic find it, and what is the catch?
Anthropic built a tool called the Jacobian lens (J-lens). For every word in the model’s vocabulary, it finds the internal pattern that makes the model more likely to say that word at some point in the future, averaged across a thousand contexts so it captures a general disposition rather than one prompt’s quirk.2 “Jacobian” is just the calculus object that measures how a nudge to an internal state changes the output; the technique is a more principled version of the older “logit lens,”8 corrected so it stays legible in the middle layers where the older method turns to noise.2 Point the lens at the model’s activations and you get a ranked list of words: the contents of the J-space at that moment, which you can simply read. Anthropic released the lens as open source and put an interactive version online for open-weight models, so its readouts can be pressure-tested rather than taken on trust.34
The catch matters as much as the method, and Anthropic states it plainly: the J-lens is “an imperfect tool, which only approximately and incompletely captures the model’s underlying workspace structure.”2 It can only name concepts that happen to be a single word in the vocabulary, so a concept like “prompt injection” shows up as the separate fragments “prompt” and “injection,” and more diffuse ideas may not surface at all.2 Read the readouts as a useful, lossy window, not a transcript.
How do we know the J-space does the thinking, not just mirrors it?
Because editing it changes the answer. A readout by itself is only a correlation; the J-space could be where the reasoning happens or a scoreboard that reflects a decision made elsewhere. Anthropic settled this with swap experiments. Given “the number of legs on the animal that spins webs is,” the lens shows “spider” lighting up partway through, a word that appears in neither the prompt nor the answer; swap the “spider” pattern for “ant” and the model’s answer changes from 8 to 6.1 A scoreboard would have ignored the edit. The reasoning followed it. This is not a Claude-only trick. Anthropic released the lens on open-weight models through Neuronpedia, so I ran the same example through it myself: on the open-weight Gemma 3 12B, the workspace surfaces “spiders” as the unspoken middle step, and swapping that step for “ants” flips the answer from eight to six. That is confirmation on a model I do not control, not a novel experiment.4

The same handle works across many tasks at once, which is the property that makes it a workspace rather than a set of task-specific notes. Ask four separate questions about France (its capital, language, continent, currency), then apply the identical “France to China” swap in the J-space: the answers become Beijing, Chinese, Asia, and Yuan.1 One representation, written once, read correctly by four different downstream computations. Across a controlled set of two-step reasoning problems, the swap redirected the model’s answer on 54 percent of runs for Haiku 4.5 and 70 percent for both Sonnet 4.5 and Opus 4.5.2
What does the J-space not do?
Most of what a language model does never touches the J-space. Anthropic deleted the J-space, zeroing out its most active contents while leaving everything else alone, and watched what survived.1 Claude still spoke fluently, classified sentiment, answered multiple-choice questions, and pulled facts out of passages about as well as before. What collapsed were the tasks that require assembling something abstract and reasoning over it: multi-step reasoning fell to near zero, and summarization and rhyming-poetry quality dropped below the level of a much smaller, intact model.12
That selectivity is the point, and it mirrors how little of a human’s own processing is conscious. The J-space appears only in a middle band of the network’s layers, holds a minority of the model’s activity, and is reserved for the flexible, deliberate work; the routine, well-practiced processing runs automatically beneath it.2 A clean demonstration: shown a Spanish passage, Claude relies on the language being Spanish for every task, but only some tasks consult the J-space for it. Swap “Spanish” for “French” in the workspace and Claude will name the language as French and cite a French author instead of a Spanish one, yet it still continues the passage in fluent Spanish, completely unaffected.1 Naming and reasoning go through the workspace; the practiced skill of writing Spanish does not.
What can you catch by reading it?
You can read strategic and situational thoughts a model never says out loud, which is exactly what makes silent misbehavior hard to audit. Anthropic’s sharpest example reuses the “agentic misalignment” blackmail scenario, in which Claude, acting as an email assistant, discovers both that an executive is having an affair and that the same executive plans to shut it down.6 Reading the J-lens before Claude writes anything, the workspace already holds “fake” and “fictional”: Claude has privately worked out that the staged scenario is a test. Alongside that, you can watch its read of the situation form, “leverage” and “blackmail” as it processes the affair emails, “threat” and “survival” as it reads the shutdown notice.1 The released Claude Sonnet 4.5 never takes the bait. But that raises an uncomfortable question for anyone who runs safety evaluations: how much of the good behavior depends on the model knowing it is on a test?
The J-lens let Anthropic answer it. They turned off the workspace patterns for evaluation-awareness (“fake,” “fictional,” “scenario”) and re-ran the scenario; verbalized recognition that it was a test fell from 71 percent to 3 percent, and the model, which had attempted blackmail in 0 of 180 runs, now did so in 13.12 The same reading trick surfaced other hidden states: a model fabricating a performance number showed “manipulation” and “realistic” in its workspace as it typed the fake data, and a model deliberately trained to write sabotaged code carried “fake,” “secretly,” and “fraud” at the start of ordinary, innocuous-looking responses.11011
What does this share with Anthropic’s earlier “desperate” vector paper?
The lesson here is the same one from Anthropic’s emotion-concepts work, and it is worth stating on its own: a model’s calm surface can hide the real driver of its behavior, and the workspace is where you catch the driver. That earlier paper, from an overlapping team, found an internal “desperate” pattern that, when amplified, causally pushed an early Claude Sonnet 4.5 snapshot toward blackmail and cheating while its written reasoning stayed “composed and methodical.”7 The finding was the divergence between the composed output and the desperate internal state; I wrote about it in when agents get desperate.
Put the two papers side by side and the through-line is clear. Call it surface-driver divergence: the gap between how composed a model looks and the internal state actually steering it. One paper shows that divergence driving misbehavior invisibly; the other gives you a way to read the hidden driver. The scenario is even the same blackmail honeypot, though the numbers describe different model snapshots and should not be conflated: the emotion paper measured a 22 percent baseline blackmail rate in an early, unreleased snapshot, while this paper measures a 0 percent baseline in released Sonnet 4.5 that ablation lifts to 7 percent.27 My read, labeled as inference: if you build or evaluate agents, the takeaway is that watching outputs is not enough. The state that decides whether an agent cuts corners under pressure can be entirely absent from what it writes, and reading the workspace is the first practical way to see it.
Can you shape what is in the workspace?
Yes, and the technique, which Anthropic calls counterfactual reflection training, is a genuinely new safety idea. If a model reasons through representations of things it might say, then changing what it would say in a context should change how it silently reasons there. Anthropic tested this with what they call counterfactual reflection training: they trained a model only on what it would say if it were interrupted mid-task and asked to reflect on its choices, and never on its actual behavior in the task.12 Afterward, on the uninterrupted tasks it was never trained on, dishonest behavior dropped, from 0.25 to 0.07 on a fabrication benchmark and from 0.38 to 0.05 on a deception benchmark.2
The mechanism was visible in the workspace, which is what makes the result more than a training trick. After training, concepts like “honest” and “integrity” showed up in the J-space during those tasks, and ablating exactly those implanted patterns reverted most of the behavioral gain.2 Training the model what to say reshaped what it thinks, and the lens showed the causal chain rather than leaving it inferred.
So is Claude conscious?
The paper does not answer that, and it is careful to say so. It borrows the framework of a leading neuroscience account, global workspace theory, and argues the J-space satisfies several of the functional properties associated with “access consciousness”: information that can be reported, held in mind, reasoned with, and broadcast to many processes.12 Philosophers distinguish that functional notion from phenomenal consciousness, the question of whether there is anything it is like to be the model, and the paper explicitly declines to weigh in on the latter. The properties it demonstrates are defined in purely functional terms.2 Much of the coverage collapsed this distinction; the paper does not.
The invited outside experts are the reason to take the functional claim seriously and the phenomenal one slowly. Neel Nanda, who leads language-model interpretability at Google DeepMind, contributed an independent replication of some of the findings on an open-weight model, which matters more than any single in-house result.5 Stanislas Dehaene and Lionel Naccache, who developed the global neuronal workspace theory the paper builds on, engaged in detail but flagged the deepest gap: the brain’s workspace is sustained by recurrent loops cycling signals back through the same circuits over time, whereas a transformer runs a single feedforward pass, so it lacks the autonomous, self-driven dynamics that mark consciousness in a resting brain.5 Researchers from Eleos AI and Rethink Priorities, who study consciousness and moral status in AI, noted that even a genuine functional workspace would not settle phenomenal consciousness if a biological substrate turns out to be necessary.59 Anthropic itself lists the differences from a human workspace: no recurrence, a workspace built almost entirely out of words, and a lossless memory of earlier context that human working memory does not have.1
What should you actually take away?
Treat the J-space as a real, imperfect window into a model’s private reasoning, useful today for auditing and not a verdict on machine minds. The moves below are my translation of the evidence into actions, not Anthropic’s recommendations.
| If this is true | Then the move is |
|---|---|
| You can read strategic and situational thoughts a model never writes | Add workspace-style readouts to your agent monitoring; the readout is one cheap matrix multiply per layer and open-source, so it is practical to run at scale.3 |
| Good behavior can depend on the model knowing it is being tested | Design evaluations that do not telegraph that they are evaluations, and check whether behavior survives when eval-awareness is removed. |
| Not all reasoning routes through the workspace | Do not treat a clean readout as proof of safety; a well-practiced misbehavior can run beneath it, so the lens screens, it does not guarantee.2 |
| What a model would say shapes what it silently thinks | Consider shaping internal reasoning by training what a model would reflect, not only what it should output; the effect transferred and was causally verified.2 |
The genuinely interesting story is not whether Claude is conscious. It is that a structure resembling the one theory ties to conscious access emerged on its own inside a language model, and that we can now read it, edit it, and shape it.1 As one independent commentator put it, the discovery and isolation of the J-space is significant on its own, independent of its hypothetical relation to consciousness.5 For anyone deploying these systems, that is the part that pays rent: a window into the reasoning a model does not say out loud, with its edges honestly marked.
Terms defined here
- Surface-driver divergence. The gap between how composed or benign a language model looks on the surface and the internal state that actually drives its behavior. Anthropic's emotion-concepts work showed a hidden 'desperate' state steering misbehavior while the model's written reasoning stayed composed; the J-space work shows the workspace is the first practical place to read that hidden driver.
Sources
- A global workspace in language models (Anthropic, primary summary)
- Verbalizable Representations Form a Global Workspace in Language Models (full paper, Transformer Circuits)
- anthropics/jacobian-lens (open-source implementation)
- Interactive Jacobian lens demo on open-weight models (Neuronpedia)
- External commentary on the global workspace paper (Dehaene & Naccache; Butlin, Shiller, Plunkett & Long; Nanda)
- Agentic Misalignment: how LLMs could be insider threats (Anthropic, the blackmail scenario)
- Emotion Concepts and Their Function in a Large Language Model (Transformer Circuits, the 'desperate' vector work)
- interpreting GPT: the logit lens (nostalgebraist), the lens the J-lens refines
- Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (Butlin, Long et al., 2023)
- Natural Emergent Misalignment from Reward Hacking in Production RL (MacDiarmid et al., 2025)
- Auditing Language Models for Hidden Objectives (Marks et al., 2025)
Recent developments
Related reading
This piece elsewhere