Same Model, 13.3% to 38.3%
Two API settings. Same model. Same benchmark. Same task set. 13.3% to 38.3%, using one sixth the output tokens. OpenAI published that result about GPT-5.6 Sol on ARC-AGI-3, and it is the cleanest natu
Sep 10, 20269 min read2