Natural Language Autoencoders Explain LLM Activations | dailyai.report
23 stories from today
Safety
113d ago
Natural Language Autoencoders Explain LLM Activations
Two LLM modules, an activation verbalizer and reconstructor, now map internal model activations to natural language descriptions. Researchers trained this system using reinforcement learning to reconstruct residual stream activations. This method helped diagnose safety-relevant behaviors during a pre-deployment audit of Claude Opus 4.6.
The Signal
NLAs provide a concrete tool for auditing model internals without supervised labels.