GPT-5.6 Sol scored 38.3 percent on ARC-AGI-3 using a proprietary API with reasoning and context compaction. In official test environments, the score plummeted to 7.8 percent. This gap contrasts with Anthropic's Opus 5, which hit 30.2 percent without custom aids. OpenAI's results suggest the performance gain relies on infrastructure, not raw model intelligence.