Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
GPT-5.6 Sol edges Kimi K3 on single-shot quality (72.7% vs 68.5% pass@1), but Kimi K3 wins on pass@k with k>1 (89.4% vs 85.8% pass@4) and costs 64% less per rollout ($4.65 vs $8.37). The models diverge in strengths (correlation 0.46), enabling a cascade strategy that covers 95.6% of DeepSWE tasks at lower cost than either model alone.
From the source
GPT-5.6 Sol edges Kimi K3 on single-shot quality, but Kimi wins on pass@k with k > 1 and costs 64% less per completed task. The two models succeed and fail in different ways, which makes routing between them the strongest play on the benchmark.
together.ai