New Method Decomposes Language Model Parameters | dailyai.report
23 stories from today
Safety
114d ago
New Method Decomposes Language Model Parameters
The adVersarial Parameter Decomposition method now allows researchers to decompose attention layers, a known hurdle for transcoders and SAEs. This technique improves upon previous stochastic and attribution-based approaches. It enables the creation of attribution graphs using causally important subcomponents.
The Signal
Practitioners can now apply parameter decomposition at scale to larger, more complex models.