OpenAI Retracts Endorsement Of SWE-Bench Pro | dailyai.report
23 stories from today
Model
51d ago
OpenAI Retracts Endorsement Of SWE-Bench Pro
Roughly 30 percent of tasks in SWE-Bench Pro are broken, according to a recent review by OpenAI. The company is pulling its previous endorsement of the popular coding benchmark. This failure highlights the fragility of current evaluation sets.
The Signal
Developers must now seek more reliable metrics to verify model programming capabilities.