When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results
An intervention study in Qwen 3.6–27B exploring the conditions for verbalized awareness of concept injections in its J-Space.
Read the resultsA Caltech physics-trained independent researcher studying how language models represent, transform, and report on information inside their activations.
Currently focused on mechanistic interpretability and intervention methods—especially what they can, and cannot, tell us about model cognition.
Selected work
An intervention study in Qwen 3.6–27B exploring the conditions for verbalized awareness of concept injections in its J-Space.
Read the resultsWorking notes
If you’ve read Anthropic’s findings on discovering an LLM’s “J-Space” and how it represents its “global workspace” and “internal thoughts,” they may seem somewhat...