Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Ask your AI
Top stories
Models & availability
Latest
Microsoft AI developed a sequential diagnosis benchmark (SD Bench) using 304 NEJM cases and an orchestrator (MAI-DxO) that combines multiple language models. The best-performing combination of MAI-DxO with OpenAI's o3 achieved 85.5% diagnostic accuracy, outperforming the 20% mean accuracy of experienced physicians.
From the source
The best performing setup was MAI-DxO paired with OpenAI’s o3, which correctly solved 85.5% of the NEJM benchmark cases.
microsoft.ai