
OpenAI reports that its GPT-5.6 Sol model achieved a 38.3% score on the ARC-AGI-3 benchmark, surpassing the 30.2% score previously set by Anthropic's Claude Opus 5.
The results were achieved using OpenAI's proprietary 'Retained Reasoning' and 'Compaction' API settings, which differ from the standardized official test environment.
ARC Prize co-founder François Chollet noted that while provider-specific settings create parity issues, they are acceptable if costs and configurations are clearly reported.