---
title: "Anthropic researchers reveal Claude's J-space representations in a paper; J-space findings fuel consciousness debate"
sdDatePublished: "2026-08-11T18:06:00Z"
source: "https://www.ibm.com/think/news/claude-can-reason-can-it-feel"
topics:
  - name: "artificial intelligence"
    identifier: "medtop:20001298"
locations:
  - "Beiyuan"
  - "Île-de-France"
  - "China"
  - "France"
---


Anthropic researchers reveal Claude's J-space representations in a paper; J-space findings fuel consciousness debate

Claude can reason. Can it feel? | IBM

Think 2026 Turn agentic AI into real business value | Think keynotes Anthropic researchers found a hidden workspace that helps Claude work through problems. A prominent neuroscientist warns against treating the discovery as evidence of consciousness. Anthropic researchers say they have glimpsed Claude’s reasoning, stoking speculation that the AI might be conscious. But a prominent neuroscientist says it may feel nothing at all. In a recent paper, Anthropic described a small set of internal representations that Claude uses to hold concepts and guide its answers. The finding revived a long-running question in the AI research community: does sophisticated reasoning point to consciousness, or does it simply make machines more convincing? Some researchers warn that users may not wait for science to decide. “If people generally believe they are interacting with a conscious entity, I think that makes us a lot more psychologically vulnerable,” Anil Seth, a Professor of Cognitive and Computational Neuroscience at the University of Sussex and Director of the Sussex Centre for Consciousness Science, told IBM Think in an interview. “We might be more likely to form parasocial relationships with AI.” Fluent language makes the idea of a conscious machine especially persuasive, Seth said. Intelligence and consciousness come together in humans, so people readily assume that a system capable of conversation and reasoning must also have an inner life. “We tend to see the world in our own terms,” he said. “We know we’re conscious and we think we’re intelligent, so we put the two together.” The definition of consciousness Seth uses comes from philosopher Thomas Nagel’s 1974 essay “What Is It Like to Be a Bat?” Nagel argued that a conscious creature has a point of view and experiences the world in its own way, even if that experience is nothing like a human’s. “It feels like something to be me,” Seth said. “It feels like something to be you. But it doesn’t feel like anything to be a table or a chair. There’s no experiencing happening.” For humans, consciousness includes the taste of ice cream, the colors of a sunset and the moment-to-moment flow of experience, Seth said. Intelligence requires no such feeling. A system can solve a problem without experiencing the act of solving it. The Turing test offers little help in answering that question, Seth said. Alan Turing’s imitation game asked whether a human interrogator could distinguish a machine from a person through written exchanges. It tested whether a machine could produce convincing intelligent behavior, not whether it experienced the conversation. “Turing was well aware that it was not a test of consciousness,” Seth said. “But we often hear, ‘The Turing test is passed. It’s intelligent, conscious.’ This is just not what the Turing test is about.” Rather than ask Claude to describe its inner life, Anthropic researchers examined what happens inside the model before it answers. A large language model (LLM) processes text through layers of mathematical operations, the paper explains. Each layer changes the model’s internal representation of the prompt. Researchers can read the prompt and the answer, but the work between them is much harder to decipher. To inspect those intermediate stages, the team developed a method called the “Jacobian lens,” or “J-lens.” It measures how a representation in one layer influences words Claude produces later. The technique allowed the researchers to translate some of the model’s activity into recognizable concepts, including ideas that never appeared in its final response. The paper calls this small, shifting collection of representations the “J-space.” Concepts that enter it can be verbalized, altered and reused as Claude works through a problem. In one experiment, researchers asked Claude for the color of the fourth planet from the sun. Before the model answered “red,” the J-lens surfaced “Mars.” The researchers replaced that representation with one associated with Earth. Claude answered “blue.” Another set of experiments tested whether the same internal concept could shape different answers. Separate prompts asked for France’s capital, language, continent and currency. Claude produced “Paris,” “French,” “Europe” and “euro.” When the researchers replaced the France representation with China inside the J-space, the answers shifted to “Beijing,” “Chinese,” “Asia” and “yuan.” Anthropic says those interventions show that J-space representations play a causal role in Claude’s reasoning. Altering them changes the model’s answers in predictable ways. Suppressing activity in the J-space produced a different result. Claude could still parse sentences, classify sentiment and extract answers from supplied text, according to the paper. But its performance worsened on multistep problems and tasks that required it to retain information and apply it in a new setting. The J-lens also surfaced activity that Claude never mentioned. In one test, an auditing system gave Claude fabricated search results designed to manipulate its response. Claude ignored the results, while the J-lens surfaced words including “fake,” “prompt” and “injection,” suggesting that the model had recognized the attempted prompt injection without saying so. Anthropic says such readouts could help developers detect hidden objectives or unsafe reasoning inside a model. The paper also cautions that the technique remains incomplete: it can miss concepts, misinterpret internal representations and reveal only the activity it can translate into language. Even with those limits, Seth said the method gave researchers a useful new view of how Claude works. “It’s interesting work,” he said. “It helped us understand how these really impressive language models do what they do.” Anthropic compares the J-space with global workspace theory, which holds that much of the brain’s activity occurs outside awareness. A small amount of information then enters a shared workspace, where it becomes consciously available for speech, reasoning and action. Claude’s J-space behaves in some of the same ways, the researchers report. The model can report its contents and reuse them across different tasks. The space also appears to have limited capacity. The comparison helps explain how Claude processes information, Seth said. It doesn’t show that Claude experiences anything while doing so. “It doesn’t really move the needle for me very much,” he said. “It’s not hugely surprising that something like this is going on in this kind of network.” A system trained to answer difficult questions would benefit from a shared format for holding intermediate results, Seth said. That could make Claude a better reasoner, though it wouldn’t give the model a point of view. The paper draws the same boundary. Anthropic focuses on access consciousness, a term for information available for reasoning, action or verbal report. Its researchers take no position on phenomenal consciousness, the subjective feeling Seth means when he speaks about the taste of ice cream or the colors of a sunset. In humans, global workspace theory may explain how information becomes available for thought and speech, Seth said. It does not explain why that information is accompanied by a subjective feeling, which is why a similar workspace inside Claude would not by itself establish consciousness. “Someone could even wonder whether global workspace theory is really a theory of consciousness at all,” he said. Claude also differs from the biological systems that inspired the theory. Anthropic notes that information moves through Claude largely in one direction, layer by layer. In the brain, by contrast, signals can loop back through recurrent circuits, allowing earlier activity to be reinforced or revised, a feature some global workspace theories consider important to conscious access. Even a closer resemblance to the brain would not prove Claude was conscious, Seth said. Global workspace theories describe patterns that may accompany consciousness in humans, where researchers already accept that consciousness exists. They don’t provide a checklist for proving that an unfamiliar system is conscious. “Global workspace theory doesn’t propose necessary or sufficient conditions for consciousness,” Seth said. “It is more descriptive of what might be happening in systems like human beings, where we already know consciousness can happen.” A deeper dispute concerns whether computation alone can produce consciousness. Some computational theories treat the mind as something like software, Seth said. If consciousness depends on a pattern of information processing rather than the biological material carrying it, then that same pattern could, in principle, run on silicon instead of neurons. From that perspective, a silicon machine might become conscious if it reproduced the right pattern of information processing, even without neurons, metabolism or a living body. Science hasn’t established that computation alone is enough, Seth said. Every known example of consciousness involves a living, embodied organism. That doesn’t prove that a machine can never become conscious, but it leaves open the possibility that biology matters. The brain also lacks the clean division between hardware and software found in an ordinary computer, he said. A neuron’s shape, chemistry, electrical properties and metabolism all influence its behavior. Scientists can’t simply discard those details, extract an abstract program and assume that the same experience will appear when the program runs on a chip, Seth explained. “You cannot simply strip away all the messy biological detail and say, ‘Here are the algorithms that the brain is running, and let’s implement that in silicon,’” he said. Seth also cautioned against treating adult human consciousness as the only standard, noting that people have historically denied consciousness to nonhuman animals and sometimes to other humans. The immediate concern doesn’t depend on proving that AI is conscious, Seth said. A machine only has to appear conscious to change how people respond to it. Seth pointed to debates over AI welfare, including proposals that advanced systems might deserve moral consideration or protection from suffering. Seth remains skeptical that Claude is conscious. Anthropic’s findings, he said, do not show that sophisticated information processing is enough to produce subjective experience, and he rejects definitions broad enough to classify any sufficiently complex system as conscious. Users face a more immediate risk. A fluent system that seems caring and self-aware could draw people into parasocial relationships or make its advice seem more trustworthy, Seth said. Companies do not need to know whether AI is conscious before responding to that danger. They can design systems to act like tools rather than people, Francesca Rossi, IBM Fellow and Global Leader for Responsible AI and AI Governance, told IBM Think in an interview. She said companies could present AI as an assistant built for specific tasks rather than give it a humanlike personality. “It does not really matter if the system is conscious or not,” Rossi said. “It is enough that it is perceived as being conscious to have an impact on people using it.”’ Seth said AI’s humanlike conversation can obscure how different the technology remains. “The challenge with AI is that it’s similar to us in some ways, like its ability to speak fluently,” Seth said. “Those ways are very seductive. But in many, many other ways, it’s entirely different.” Get curated insights on the most important—and intriguing—AI news. Subscribe to our weekly Think newsletter. See the IBM Privacy Statement.