Proposed Metaethical Framework For Model Alignment | dailyai.report
23 stories from today
Safety
47d ago
Proposed Metaethical Framework For Model Alignment
A new proposal suggests using perspectival moral realism and evolutionary debunking to refine AI values. The author argues that Anthropic's constitutional approach requires substantive philosophical contributions beyond standard bug reports. This niche framework targets the gap between naive realism and preference-satisfaction.
The Signal
It offers a specific, though unlikely, path to altering training decisions.