October 8, 2026 · 345 runs · three models, five effort levels
Claude Sonnet 5.5 vs Claude Opus 5.5 vs GPT-6.1 Sol
The three models side by side on the same 23 tasks at every effort level, judged by the same neutral panel (Muse Spark 1.3 and Grok 4.6): combined score with ±1 SE and judged n, Code quality, tasks solved, cost and minutes per task, with every caveat and every data link in one place.
Tasks solved Low to Max: Sonnet 5.5 15, 17, 20, 22, 23; Opus 5.5 17, 23, 22, 22, 21; GPT-6.1 Sol 20, 22, 23, 23, 23. Cost per task: $2.39 to $5.52, $1.70 to $8.75 and $0.31 to $0.46. Sonnet 5.5 has the highest best-level score of the three (92.02 ± 0.35 at Max), edging Opus 5.5 (91.11 ± 0.57 at High) by about 0.9 points, not a clear win, and clearly ahead of GPT-6.1 Sol (88.78 ± 0.40 at Extra-high).