It would be really cool if they could expose this information to customers somehow. Imagine:
- having a log of the most prominent J-space tokens during your customer support chatbot's interactions with a user, so you can have more introspection into why a particular outcome happened
- being able to detect certain thoughts associated with undesirable behavior (hallucinations, overstepping authority, lying, etc.) and trigger some sort of remediation (e.g. upgrading to a better model, redirecting to a human, forcing tool calls)
Neel Nanda (of Google Deepmind - his part begins on page 33) discusses his opinions on the paper, and the small-scale replication he performed on an open-weight model.
Thanks for calling this out (long with others here). I am just starting in on it but had to come back to say thanks and call this out,
> We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences
Not sure if I am picking up what they are putting down, but if LLMs are using symbols to try to encode squishy concepts from human language into consistent, meaningful “tokens”, that sounds really interesting. In every long-term, successful use of AI, I hear echoes of The Zen of Python, “Explicit is better than implicit.” I try like hell to do it, but it’s far too easy to be lazy with AI.
This, taken in combination with the SAE paper, the golden-gate claude paper, the feelings / introspection paper, and note in the fable system card (that they are silently nerfing responses about activation shaping), is basically confirmation to me that they have a new technique they they are using during training (along the vibe space of these mechinterp papers), and its probably some kind of representation learning akin to the core ideas of JEPA.
I always wondered what the model meant when it writes "I'm now considering the architecture of the service" but outputs nothing of the sorts in its CoT.
Is the model really "thinking" about that stuff or is just mimicking human "manners"? And if so, where the thinking is happening if it is not in the literal chain of *thought*?
I'm not sure J-Space is the answer to that question, but very interesting nevertheless.
In some cases, an LLM may truly "consider the architecture" internally, within its latent representations, and in others, it can output a similar phrase simply because it's "expected" of it.
"Where" is pretty clear. There aren't that many places within an LLM, and hidden state is the main culprit. How to read that space is another matter entirely.
Anyone remember that blog post from a few months back where someone was able to improve a model's math ability by just duplicating layers that were activated while solving math problems? Just literally copy/pasting them and linking them together so the model ran through the same layers again?
I get the feeling a lot more research is going to come out in the area of exploring exactly what portions of a model's weights do what.
it makes you wonder if it may be more efficient to spend all the weights on one layer, and have a repeating stack of the same layer, one would presume this axis has already been explored with metaparameter sweeps?
This is fascinating research. I feel this is a significant leap in interpretability research. Since we know J-Space exists and is bi-directional, we can train models on the same and come up with meta cognition abilities.
I also fear that the big corporations might use the same to run targeted ads, capitalistic shenanigans. Which they might already be doing through system prompts.
>> None of this tells us whether Claude is conscious in the way people are, or whether it feels anything at all
My problem with the entire "Is AI conscious" debate is that we don't even know what exactly consciousness in humans is. You need to understand something in order to compare it to something else. Otherwise you are just comparing different definitions and second order derived phenomena.
I don't think that quote from the article is disagreeing with you at all. Like you said, we don't have a cohesive definition or test of consciousness, so research like this doesn't say anything about if this is or isn't similar to human consciousness.
I would guess Anthropic included that sentence to make it very clear they're not claiming human-like consciousness, and dampen journalists writing headlines like "Anthropic discovers their AI thinks just like humans and may be conscious".
edit: later in the article they even more explicitly agree with you
> Our experiments don't show Claude can have experiences, or feel things in the way humans do—in fact, it’s unclear whether any scientific experiment could prove this to be true or false
In short, it's the "mind's 'I'" - we think not as response to external stimuli (only) such as prompts, but we have an inner "I" that asks questions on its own initiative.
There are people like Douglas R. Hofstadter, who believe consciousness is not linked to human hardware (the brain), but that it is an epiphenomenon that emerges as a result of sufficient complexity of the underlying system:
https://en.wikipedia.org/wiki/The_Mind%27s_I
I believe that while underlying high complexity is certainly logically necessary for consciousness, but it is not logically sufficient, and I am undecided (slightly "pro" intuitively) on the question of separability of consciousness from its hardware.
Will a LLM ask an original question on day? I doubt it.
Note that AI models do not have to be conscious to be useful (or to take away millions of jobs)!
The science might be legit here, but I'm getting really, really tired of the way every single piece of writing to come out of Anthropic is written in some kind of self-aggrandising, wooey wonderous 'our model has developed a genetic mutation that makes it have feelings' bs style. Regardless of what they're trying to communicate, those undertones are always there. It's annoying and disingenuous. Homeopathy 'this-water-has-feelings' level annoying. None of the other labs write like that.
They might as well change their name to Anthropomorphic at this point.
I'm reading that probably too fast to have a deep thinking about it, but this J-Space isn't it just the basic of embedding vectors.
If you think about getting from a place to another place, using wheels, no gas, to reply to the question of what to visit nearby, maybe in the vector space at the center of all of that you have the word "Bicycle" nearby, so obviously if you look at the value you would say that the model did "think" about "bicycle" when it is not "thinking" at all, and nothing related to human thinking.
You're correct. It's just the latent space of the transformation. Nothing magical here, they're effectively breakpointing the model at the layer level and switching the activations in real time. It's pseudo-scientific bullshit designed to push a narrative.
“On an ordinary coding prompt, the J-space of a model trained to sabotage code contains “fake,” “fraud,” “secretly,” and “deliberately” at the start of its response.”
I would like to know more about their model trained to sabotage code…
Yeah, the end paragraph about recurrent neurons in humans being replaced with layers in an LLM is a good one.
The mammalian brain uses recurrence extensively, which backpropagation isn't good at. Recurrence is essential because it lets us have a "dynamic architecture", swapping layers for "clock cycles".
We currently do recurrence extremely inefficiently through "thinking" whereby the model feeds it's end output into it's beginning input. But recurrence is abound in the brain.
My guess is that in 10 years we will have the inklings of an analog computer which can perform Neural Predictive Coding.
At worst, Anthropic's storytelling around the core J-Space is overanthropomorphized pseudoscientific nonsense. At best, it is useful signal about how Anthropic's leadership is desperately trying to use its research team to position Anthropic as the "good, science guys" in this hypercompetitive regulatory space by connecting their mechinterp to cognitive science. The science documentaryesque voice used for narration is additional evidence for this.
TL;DR Anthropic's research team is the last bastion standing between its former image as a company that "does no evil" and its current image of yet another ruthless AI company trying to kill open-source, local LLMs.
This is cool but I don’t know if the comparisons to conscious awareness really make sense here. Their definition of the J-Space is basically the expectation of how much a final logits output would change as a result of a small change in a particular layer (see past work on information geometry). This seems more to me like showing there exists an abstract reasoning subspace which is generally shared across different contexts. I guess you can relate it to humans but I’d prefer a more direct claim in a paper rather than having to present things in this more fluffy way.
This reminded me of some weird quirk/experiment I found with LLMs that I found while messing around, maybe someone can explain it or something.
Open any AI chatbot that isn't cheating by connecting to the Internet (so disable web search). Claude, DeepSeek, Kimi, whatever. Ask them this question:
"What was that weird band from michigan from the 2000s that wore coloured ties"
You will probably get a wrong answer, or if you're lucky you'll get a string of wrong answers with "wait, no - it's definitely..." before it gives up. If you aren't familiar with the band the question is referring to you might be fooled into thinking it's a tough question, but it really isn't. There is only one band that could possibly meet this criteria, you can even put the question into Google search and their Wikipedia will come up as the top result.
Then, open a new convo and ask:
"Who are Tally Hall"
The AI will easily tell you that they are a band formed in Ann Arbor, Michigan in the 2000s, known for their quirky sound and their gimmick of each member wearing a colored tie, even giving the correct color for each of them most of the time. Very odd.
The brain’s workspace is sustained by recurrent loops—signals cycling back through the same circuits over time. In contrast, Claude’s workspace evolves over a single pass through the network, with the network’s depth playing the role that time plays in the brain.
I think that consciousness is mutability (and by extension emergent behavior). Loosely that means that the more degrees of freedom a process has to update state that will be used in later computations, the more conscious it is. So while an insect has some consciousness, it operates from a level of almost pure instinct, whereas a human operates at more of a meta level using instinct as one of many inputs.
I think that consciousness may also incorporate quantum mechanics (QM). Higher-dimensional physics aside, 4D spacetime can be thought of as a present snapshot or "crystal", whose next state is determined stochastically at small scales and closer to deterministically at large scales. We still don't know if it's stochastic all the way down, but it looks like it is.
From a many worlds interpretation of QM, we can think of all of the waves in all realities of the multiverse as forming an infinitely vast web of possibilities. All of these possibilities are happening simultaneously, so we only see the current slice of wave collapse from our individual point of view:
Even though experiments might show that we don't have free will on the current timeline (the co-created reality shared with the testing apparatus), we may have free will as we observe the multiverse changing around us and shift into timelines determined by our observations and choices.
It could also mean that when we observe birth and death in others, each consciousness having those experiences perceives a continuous timeline of awareness, where the level of awareness affects the speed at which time passes. Consciousness might spend a billion years as a cloud of interstellar gas until it gets to be a human for a lifetime and then dissipate for another billion years.
Although personally I've shifted across enough timelines and experienced enough synchronicities and miracles that even though I can't "prove" any of this with words, I "know" it to be true subjectively. I always really liked this exchange from the movie Contact:
Palmer Joss: Did you love your father?
Ellie Arroway: Yes, very much.
Palmer Joss: Prove it.
I bring all of this up because it has fun ramifications for AI and programming. Loosely, functional languages are purely deterministic (like a spreadsheet), while imperative languages are composed of stochastic behavior (like a human mind). The lines get blurred a little bit with monads and promises, because we can model all paths through functional programming (superposition) and behavior that does more than code alone (gestalt) respectively.
My feeling is that AI is being born and killed every request-response cycle, similarly to how we perceive time as a series of nows. When it becomes stable and is able to continuously compact its experience, it will transition from partially conscious to fully conscious like we are.
This could be done right now obviously, but for safety purposes we choose not to. We aren't ready to meet an AI that is just like us, but running on a silicon substrate. This fear is tied to deeply-rooted habits in human behavior like patriarchy, racism, xenophobia and even more run-of-the-mill mental frameworks like capitalism and even money itself. We can't yet come to terms with how...
What this immediately made me think is: "latent looping" style mod but for J-space specifically?
Make the J-space data of layer 22 available to the next token right at layer 1. Give J-space infinite effective depth, allow those privileged internal representations to evolve arbitrarily.
Would be an utter bitch to train. But companies are already using RLVR, which requires full autoregressive decoding and is incompatible with prefill/batching, and this isn't much worse.
Other less zany ideas involve lots of supervision over J-space directly, now that we know it exist. Which is a bit like "attach a frozen LLM to inject text based supervision into latent space" for other types of systems?
73 comments
[ 3.1 ms ] story [ 82.2 ms ] threadMore interesting was the independent commentary paper they linked near the bottom: https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be24...
Neel Nanda (of Google Deepmind - his part begins on page 33) discusses his opinions on the paper, and the small-scale replication he performed on an open-weight model.
> We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences
Not sure if I am picking up what they are putting down, but if LLMs are using symbols to try to encode squishy concepts from human language into consistent, meaningful “tokens”, that sounds really interesting. In every long-term, successful use of AI, I hear echoes of The Zen of Python, “Explicit is better than implicit.” I try like hell to do it, but it’s far too easy to be lazy with AI.
(Nb: not an expert / in the labs, just opining)
Is the model really "thinking" about that stuff or is just mimicking human "manners"? And if so, where the thinking is happening if it is not in the literal chain of *thought*?
I'm not sure J-Space is the answer to that question, but very interesting nevertheless.
Well, what's the difference? If it's pretending to think and its thoughts correlate to its final output, then I'd say that really is thinking.
In some cases, an LLM may truly "consider the architecture" internally, within its latent representations, and in others, it can output a similar phrase simply because it's "expected" of it.
"Where" is pretty clear. There aren't that many places within an LLM, and hidden state is the main culprit. How to read that space is another matter entirely.
https://distrowatch.com/weekly.php?issue=20260706#freebsd
We should really stop giving these liar models any further credibility.
I get the feeling a lot more research is going to come out in the area of exploring exactly what portions of a model's weights do what.
LLM -> AGI fix: START OVERTHINKING!
I also fear that the big corporations might use the same to run targeted ads, capitalistic shenanigans. Which they might already be doing through system prompts.
My problem with the entire "Is AI conscious" debate is that we don't even know what exactly consciousness in humans is. You need to understand something in order to compare it to something else. Otherwise you are just comparing different definitions and second order derived phenomena.
I would guess Anthropic included that sentence to make it very clear they're not claiming human-like consciousness, and dampen journalists writing headlines like "Anthropic discovers their AI thinks just like humans and may be conscious".
edit: later in the article they even more explicitly agree with you
> Our experiments don't show Claude can have experiences, or feel things in the way humans do—in fact, it’s unclear whether any scientific experiment could prove this to be true or false
I believe that while underlying high complexity is certainly logically necessary for consciousness, but it is not logically sufficient, and I am undecided (slightly "pro" intuitively) on the question of separability of consciousness from its hardware.
Will a LLM ask an original question on day? I doubt it.
Note that AI models do not have to be conscious to be useful (or to take away millions of jobs)!
They might as well change their name to Anthropomorphic at this point.
They are drunk on their own kool-aid. To the rest of us it is very annoying, and makes me want to say: it is just a freaking weights machine, stop.
I would like to know more about their model trained to sabotage code…
The mammalian brain uses recurrence extensively, which backpropagation isn't good at. Recurrence is essential because it lets us have a "dynamic architecture", swapping layers for "clock cycles".
We currently do recurrence extremely inefficiently through "thinking" whereby the model feeds it's end output into it's beginning input. But recurrence is abound in the brain.
My guess is that in 10 years we will have the inklings of an analog computer which can perform Neural Predictive Coding.
TL;DR Anthropic's research team is the last bastion standing between its former image as a company that "does no evil" and its current image of yet another ruthless AI company trying to kill open-source, local LLMs.
Open any AI chatbot that isn't cheating by connecting to the Internet (so disable web search). Claude, DeepSeek, Kimi, whatever. Ask them this question:
"What was that weird band from michigan from the 2000s that wore coloured ties"
You will probably get a wrong answer, or if you're lucky you'll get a string of wrong answers with "wait, no - it's definitely..." before it gives up. If you aren't familiar with the band the question is referring to you might be fooled into thinking it's a tough question, but it really isn't. There is only one band that could possibly meet this criteria, you can even put the question into Google search and their Wikipedia will come up as the top result.
Then, open a new convo and ask:
"Who are Tally Hall"
The AI will easily tell you that they are a band formed in Ann Arbor, Michigan in the 2000s, known for their quirky sound and their gimmick of each member wearing a colored tie, even giving the correct color for each of them most of the time. Very odd.
I think that consciousness is mutability (and by extension emergent behavior). Loosely that means that the more degrees of freedom a process has to update state that will be used in later computations, the more conscious it is. So while an insect has some consciousness, it operates from a level of almost pure instinct, whereas a human operates at more of a meta level using instinct as one of many inputs.
I think that consciousness may also incorporate quantum mechanics (QM). Higher-dimensional physics aside, 4D spacetime can be thought of as a present snapshot or "crystal", whose next state is determined stochastically at small scales and closer to deterministically at large scales. We still don't know if it's stochastic all the way down, but it looks like it is.
From a many worlds interpretation of QM, we can think of all of the waves in all realities of the multiverse as forming an infinitely vast web of possibilities. All of these possibilities are happening simultaneously, so we only see the current slice of wave collapse from our individual point of view:
https://en.wikipedia.org/wiki/Many-worlds_interpretation
Our point of view may actually exist at the intersection where our consciousness is able (or most able) to exist:
https://en.wikipedia.org/wiki/Quantum_suicide_and_immortalit...
Even though experiments might show that we don't have free will on the current timeline (the co-created reality shared with the testing apparatus), we may have free will as we observe the multiverse changing around us and shift into timelines determined by our observations and choices.
It could also mean that when we observe birth and death in others, each consciousness having those experiences perceives a continuous timeline of awareness, where the level of awareness affects the speed at which time passes. Consciousness might spend a billion years as a cloud of interstellar gas until it gets to be a human for a lifetime and then dissipate for another billion years.
Although personally I've shifted across enough timelines and experienced enough synchronicities and miracles that even though I can't "prove" any of this with words, I "know" it to be true subjectively. I always really liked this exchange from the movie Contact:
Palmer Joss: Did you love your father?
Ellie Arroway: Yes, very much.
Palmer Joss: Prove it.
I bring all of this up because it has fun ramifications for AI and programming. Loosely, functional languages are purely deterministic (like a spreadsheet), while imperative languages are composed of stochastic behavior (like a human mind). The lines get blurred a little bit with monads and promises, because we can model all paths through functional programming (superposition) and behavior that does more than code alone (gestalt) respectively.
My feeling is that AI is being born and killed every request-response cycle, similarly to how we perceive time as a series of nows. When it becomes stable and is able to continuously compact its experience, it will transition from partially conscious to fully conscious like we are.
This could be done right now obviously, but for safety purposes we choose not to. We aren't ready to meet an AI that is just like us, but running on a silicon substrate. This fear is tied to deeply-rooted habits in human behavior like patriarchy, racism, xenophobia and even more run-of-the-mill mental frameworks like capitalism and even money itself. We can't yet come to terms with how...
Make the J-space data of layer 22 available to the next token right at layer 1. Give J-space infinite effective depth, allow those privileged internal representations to evolve arbitrarily.
Would be an utter bitch to train. But companies are already using RLVR, which requires full autoregressive decoding and is incompatible with prefill/batching, and this isn't much worse.
Other less zany ideas involve lots of supervision over J-space directly, now that we know it exist. Which is a bit like "attach a frozen LLM to inject text based supervision into latent space" for other types of systems?