Modular Pretraining Isolates Dangerous AI Knowledge | dailyai.report
23 stories from today
Safety
52d ago
Modular Pretraining Isolates Dangerous AI Knowledge
Gradient Routed Auxiliary Modules (GRAM) isolate sensitive knowledge into specific, switchable components within a language model. Researchers can now toggle these modules to restrict or grant access based on user trust. This approach allows one model to mimic multiple versions with varying safety constraints.
The Signal
It provides a concrete mechanism for granular access control.