A new research paper from Anthropic identifies specific features within Claude that represent complex concepts. The team used dictionary learning to decode how the model processes information internally. This transparency helps researchers mitigate harmful outputs.
The Signal
It provides a concrete method for steering model behavior without relying solely on trial-and-error prompt engineering.