Gemma Fails To Optimize For Step Limits | dailyai.report
23 stories from today
Safety
55d ago
Gemma Fails To Optimize For Step Limits
A 600-run experiment using Gemma on cyber CTF labs found no difference in solve rates when models were told how many steps remained. Baseline runs hit 67.6% success compared to 65.5% for step-aware runs.
The Signal
This suggests the model cannot strategically allocate resources or adjust its behavior based on known constraints.