A 38.3 percent score on ARC-AGI-3 puts GPT-5.6 Sol ahead of Opus 5, but only when using OpenAI's specific API settings. The model scored just 7.8 percent under the official, provider-neutral test environment. This discrepancy suggests the performance gain relies on proprietary infrastructure rather than raw reasoning. Practitioners should treat these non-standard benchmarks with caution.