Feedback Gains Often Stem From Resampling | dailyai.report
23 stories from today
Research
59d ago
Feedback Gains Often Stem From Resampling
Thirteen open-weight models tested across Omni-MATH and Codeforces reveal that multi-turn improvements often result from resampling or format correction rather than actual feedback. This student-teacher protocol separates external guidance from unguided self-refinement. The findings suggest that natural-language feedback provides fewer gains than previously assumed.
The Signal
Practitioners should prioritize test-time computation over complex feedback loops.