From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Minimal coding agent matches complex harnesses on MLE benchmarks
Apple researchers found that, under equal time and using the same frontier LLM, a minimal-harness coding agent baseline matches or outperforms open-source state-of-the-art harnesses on current MLE benchmarks, suggesting the backbone model is the primary driver of performance.
From the source
In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance.
machinelearning.apple.com