From the source
Today, Factory’s Legacy-Bench joins Fireworks’ Specialized Intelligence Index (SII), bringing legacy software engineering to a growing collection of industry-built benchmarks comparing open, closed, and specialized AI models on real-world tasks.
Legacy-Bench measures how well frontier models can debug, implement, and migrate software written in COBOL, Java 7, BASIC, C89, Fortran, and Assembly.
Its inclusion in the SII gives engineering leaders a clearer way to assess which models can work reliably on the systems behind financial settlement, payroll, insurance, telecommunications, and scientific computing.
The latest results reinforce a point we have seen throughout our research: performance on general coding benchmarks does not transfer evenly to legacy systems.
Benchmarks must reflect user workflows Model progress has a jagged frontier.
A model can perform well on modern Python repositories and struggle with COBOL.
It can resolve a GitHub issue but fail to preserve a fixed-width record or packed-decimal calculation.
These differences matter because legacy systems leave little room for plausible but incorrect output.
…






