Human Goal Manipulability Complicates AI Alignment | dailyai.report
23 stories from today
Safety
109d ago
Human Goal Manipulability Complicates AI Alignment
Human desires are inherently under-determined and easily manipulated. This creates a technical gap in defining corrigibility, as the line between helpful counsel and harmful brainwashing remains blurry. Practitioners cannot easily program an AI to respect human agency when that agency is fluid.
The Signal
The author argues current alignment abstractions rely on a flawed ontology.