Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The post compares DeepSeek V4 Pro 0813 and GPT-5.6 Sol on the DeepSWE benchmark. GPT-5.6 Sol has higher single-shot accuracy (72.7% pass@1 vs 62.8%) but is 35x more expensive ($8.37 per rollout vs $0.24). The cheaper model DeepSeek V4 Pro 0813 matches or exceeds Sol's accuracy with multiple attempts (88.5% pass@4 vs 85.8%). A cascade strategy—using DeepSeek V4 Pro 0813 first and escalating to GPT-5.6 Sol on failure—solves 83.0% of tasks at $3.35 each, outperforming either model alone in cost-efficiency.
From the source
Don't pick one. Run DeepSeek V4 Pro 0813 first, escalate to GPT-5.6 Sol when the tests fail. That cascade solves 83.0% of DeepSWE tasks at $3.35 each.
together.ai